Skip to Content

2020

Evidence-based Software Engineering

This book discusses what is currently known about software engineering, based on an analysis of all the publicly available data. This aim is not as ambitious as it sounds, because there is not a great deal of data publicly available. The intent is to provide material that is useful to professional developers working in industry; until recently researchers in software engineering have been more interested in vanity work, promoted by ego and bluster. The material is organized in two parts, the first covering software engineering and the second the statistics likely to be needed for the analysis of software engineering data.

Software development has progressed in to the age of the ecosystem successfully building a software system is dependent on a team capable of effectively selecting the libraries and packages providing the algorithms that are good enough to get the job done, write code when necessary, and to interface to a myriad of services and other software systems,

It’s true: even in large projects that putatively written in-house, any given project has to interface with a large suite of other projects.

All the of data and analysis within this book is recorded in this repo of R code: https://github.com/Derek-Jones/ESEUR-code-data

The craft approach has survived because building software systems has been a sellers market, customers have paid what it takes because the potential benefits have been so much greater than the costs. In a competitive market for development work and staff, paying people to learn from mistakes that have already been made by many others is an unaffordable luxury.

This is an interesting pitch — and one that’s aimed squarely above the people who write code. One could also argue that the reasons for learning from others mistakes is an important precondition to maintaining and growing a rigorous field, no matter the endeavor.

Lots of focus here on the economics of writing software systems, which seems very reasonable.

The demand for people with the cognitive firepower needed to implement complex soft-ware systems has drawn talent away from research. Consequently, little progress has been made towards creating practical theories capable of supporting an engineering/scientific approach to software development

… huh. I guess that’s true, tho I would also suggest that our systems of sustaining research are inaccessible to those who would wish to pursue it.

Those who develop software systems are not motivated to invest in change when customers are willing to continue paying for systems developed using craft practices.

… huh. I guess this is a very good point. Like who gives a fuck, as long as I still command a high price for my work, how efficient that work is.

There is little or no evidence for many existing theories of software engineering

Hahah haha ha ha ha.

Some of the basic framing for this book is a little weird, but should not be too weird downstream.

Provided software development projects looked like they would deliver something that was good enough, those involved knew that the customer would wait and pay; complaints received a token response. Software vendors learned that one way to survive in a rapidly evolving market was to get products to market quickly, before things moved on

I mean … yikes but also yeah. Software vendors don’t really give a shit.

After framing some axioms, Jones gives quick tour through the history and economic impacts of software.

Researchers with a talent for software engineering either moving on to other research areas or to working in industry, leaving the field to those with talents in the less employable areas of mathematical theory, literary criticism (of source code) or folklore.

Haha depressing. But ties in to the above point — this field is under-researched because no one wants to know.

A lack of evidence has not prevented researchers expounding plausible sounding theories that, in some cases, have become widely regarded as true.

This is … bad! What the hell! And obviously true — this is the era of the “thought leader”.

The dearth of experimental evidence has left a vacuum that has been filled by folklore

This seems accurate as well. So much of ones early career is learning these stories and legends, sometimes recounted as literal fire-side tales from grizzled old developers. The morals tend to be hazy, but they are still al folklore, and we build our image of the industry out of them.

Jones is basically creating a strong argument that programming today is not rigorous.

Although academics continue to work in a feudal based system of patronage and reputation, they are incentivised by the motto “publish or perish”,1453 with science perhaps advancing one funeral at a time.

Damn dang what the hell. Thats harsh.

I sort of skipped chapter two — very interesting stuff about how humans think about stuff, but I’m willing to let my fingertips-feeling carry the bulk of this reasoning for now. I might backtrack over it later if I feel some specific assertions need to be supported from this realm.

Chapter three gets us into “cognitive capitalism” which is an interesting chapter head.

Is it worthwhile investing resources, to implement software that provides some desired functionality?

This is the key question that we look to answer with return projections and cost modeling. A decent equation is given for figuring ROI, which requires some making up of uncertainty numbers. Another equation is given for figuring discount rates. Jones gets into the nitty-gritty actual-math here that governs investments strategies; something that I have heard talked about remarkably little by anyone in the field. As programmers we’ve managed to isolate ourselves from this world, for reasons hinted at in the introduction. It would be very interesting to operationalize these equations and allow for a plug-and-play approach to figuring ROI outcomes.

Jones ends section 3.2.1 by addressing the basal rate of capital investments — US government bonds.

…specified rates, varied between 0.9% over a three-year period and 2.7% over 30 years.

This is the low-water mark for the gravity of capital. Capital will always try and flow downhill from this point following a pressure gradient of higher returns, so this is the benchmark above which you’re not burning money. Which seems you know … pretty achievable. The rest of section 3.2 build up more and more robust financial quant shit, which is very interesting, but beyond use as a base for computational rhetoric, seems difficult to get anyone who isn’t a financial quant to give a shit about.

Section 3.3 comes out the gate song:

Organizations whose income is predominantly derived from the output of cognitariate are social factories.

Damn. 3.3.1 and 3.3.2 continue to discuss the ecosystems around the way the industry works. Its very interesting and illuminating, perhaps more so for people who have not been entrenched in it for a decade plus.

Passion is a double-edged sword; passionate employees are more easily exploited,1000 but they may exploit the company in pursuit of their personal ideals (e.g., spending time reworking code that they consider to be ugly, despite knowing the code has no future use).

I’ve been saying this! Passion in an engineer can be a very bad thing.

The mass market for those seeking to promote female equality is incompetence; a company cannot be considered to be gender-neutral until incompetent women are equally likely to be offered a job as incompetent men.

Haha I love this.

Software developers are not professional programmers any more than they are professional typists; reading and writing source code is one of the skills required to build a software system. Effort also has to be invested in acquiring application domain knowledge and skills, for the markets targeted by the software system.

Thats an interesting assertion. Chapter 3 continues to vacillate weirdly between very capitalist definitions of good, and critique of current systems, ie going from praising Taylor to

Technological progress continues to be written about as a form of religious transcendence.

Chapter 3 continues to plod along, examining how and why people work and learn and learn about work. Kinda dry! I like this gem tho in 3.4.6 Information Asymmetry:

Vendors bidding to win the contract to implement a new system can claim knowledge on similar systems, but cannot claim any knowledge of a system that does not yet exist.

Some interesting theory gestured at in 3.4.7 Moral Hazard — very brief discussion, but maybe worth thinking about.

Okay, just a rough smattering of ideas through chapter 3 so far, now lets think for a second about 3.5 — Company Economics!

Nevermind it’s kinda boring. Chapter 3 wraps up with section 3.6 which is more interesting in a SaaS model or other sort of “selling units” structure, which at this time is not what Im thinking about. Maybe I’ll think about that again later.

On to Chapter 4, Ecosystems!

Software ecosystems have been rapidly evolving for many decades, and change is talked about as-if it were a defining characteristic, rather than a phase that ecosystems go through on the path to relative stasis. Change is so common that it has become blasé, the mantra has now become that existing ways of doing things must be disrupted. The real purpose of disruption is redirection of profits, from incumbents to those financing the development of systems intended to cause disruption. The fear of being disrupted by others is an incentive for incumbents to disrupt their own activities; only the paranoid survive748 is a mantra for those striving to get ahead in a rapidly changing ecosystem.

… huh. Yes. What does this imply? What the shape of a stable ecosystem? How does the flow of capital and profit determine this turbulence, and is it desirable to settle it down? What would be the political ramifications of stability?

The first release of a commercial software project implements the clients’ view of the world, from an earlier time. In a changing world, unimportant software systems have to adapted, if they are to continue to be used; important software systems force the world to adapt to be in a position to make use of them.

The world changes, or the information about the world changes, and the softwares relationship to reality

Left untouched, software remains unchanged from the day it starts life; use does not cause it to wear out or break. However, the world in which software operates changes, and it is this changing world that reduces the utility of unchanged software. Software systems that have had minimal adaptation, to a substantially changed world.

The hardware update cycle drives a Red Queen133 treadmill, where businesses have to work to maintain their position, from fear that competitors will entice away their customers.

This is a good explanation for the rapidly changing ecosystem, but it does indicate that the recent narrowing of hardware advancement might settle things down. Skimming the rest of chapter 4 gives an overview of the ecology of the industry, which again, while interesting, is not really capturing my interest. Something that would be worth returning to, certainly.

Building a software system is a creative endeavour; however, it differs from artistic creative endeavours in that the externally visible narrative has to interface with customer reality.

ITS TRUE. This similarity is addressed at length by Daglian in Conceptual Labor.

While the reasons of wanting a cost estimate cannot be disputed, an analysis of the number of unknowns involved in a project, and the experience of those involved can lead to the conclusion that it is unreasonable to expect accurate resource estimates.

More hard truths!

People cling to the illusion that it’s reasonable to ask for accurate estimates to be made about the future.

Estimates are fake!

While the reasons of wanting a cost estimate cannot be disputed, an analysis of the number of unknowns involved in a project, and the experience of those involved can lead to the conclusion that it is unreasonable to expect accurate resource estimates.

Accurate resource estimates don’t exist!

This is additionally slippery when then definition of what a software project is is so slippery. The same spec probably could take 2 months or 20 months, with different degrees of quality or non-specced behavior.

On estimation modeling,

The problem with fitting equations to data, is that the resulting models are only likely to perform well when estimating projects having the same characteristics as the projects from which the data was obtained.

And even then it’s no guarantee! If we think projects share characteristics, that’s part of us rather than the work. If we’re wrong, and need to perform conceptual labor at any point during the project — that is, engage in work where the work, contexts, and agents are all able to modify each other — the fat-tail nature reverts itself and our estimate model is now deceptive.

Apparently we used to try and use a Rayleigh curve to model development time, which is deeply foolish. If the base assumption of Completing a project involves solving a set of problems is true at all, than that means you must know all the members of the set at the start. This is often deeply untrue, with problems nested inside of each other in unpredictable ways.

Building a simulation model requires understanding the behavior of all the important factors involved, and as the analysis in this book shows, we are a long way from having this understanding.

This might be too hard to be worthwhile!

Breaking a problem down into its constituent parts may enable a good enough estimate to be made (based on the idea that smaller tasks are easier to understand, compared to larger, more complicated tasks), assuming there are no major interactions between components that might affect an estimate

Haha, wishful thinking! Christopher Alexander showed that these design problems can be expressed as a graph, and graphs can be decomposed via a min-cut algorithm, but doing this is very hard and involved. The odds of decomposing your tasks into minimal-interacting components is very small.

The charts in chapter 5.20 demonstrate that this line of estimation is basically a shit show.

Creating a usable software system involves discovering a huge number of interconnected details; these details are discovered during: requirements gathering to find out what the client wants (and ongoing client usability issues during subsequent phases), formulating a design for a system having the desired behavior, implementing the details and their interconnected relationships in code, and testing the emergent behavior to ensure an acceptable level of confidence in its behavior.

This is the crux of the issue! Software worth writing is emergent and discovered through the act of writing it.

…it is possible that daily contact enabled the contractor to convince the client to be willing to accept a deliverable that did not support all the functionality originally agreed.

This is a known thing that good client-touch people can wrangle well. Basically everything at the top of the project is bullshit, and where you land on the matrix of money/time/thing is basically a function of trying to ride it out and steer the process as best as you can.

system development methodologies have been found to provide management support for necessary fictions, e.g., a means to creating an image of control to the client, or others outside the development group

It’s all bullshit! All of it! It’s stories we tell to control relationships not have any real impact on the work.

The Waterfall model1270 continues to haunt software project management, despite repeated exorcisms over several decades

No notes. Perfect sentence.

It seems like a lot of the work that goes in to project estimation are for the benefits of contract and breach negotiation rather than as a tool for developers.

Bespoke software development is not a service that many clients regularly fund, and they are likely to have an expectation of agreeing costs and delivery dates for agreed functionality, before signing a contract.

My sweet, summer children.

What is a cost effective way of discovering requirements, and their relative priority?

A key question! No real answer provided beyond a cursory examination of finding salient steak holders.

Chapter 5 close without any really substantive insights beyond “there are lots of parts of a project and you can fuck up at any of them.”

Chapter 6 is about reliability!

Proposals that programming should strive to be more like mathematics are based on the misconception that the process of creating proofs in mathematics is less error prone than creating software.

Errors happen! Mistakes, oversights, and unexpected emergent behaviors! Woo hoo!

The classification of program behavior as a fault, or a feature, can depend on the person doing the classification, e.g., user or developer

Bugs vs Features! When using features as a tool to do product estimating we get get blindsided here – features are things we meant to do, bugs are things we did not mean to do, but either way they can be expressed in terms of behavior.

Intel’s Pentium processor was introduced in 1993, as the latest member of the x86 family of processors. Internally the processor contains a 66-bit hardwired value for π. A double precision floating-point value is represented using a 53-bit mantissa, which means internal operations involving values close to π (e.g., using one of the instructions that calculate a trigonometric function) may have only 13-bits of accuracy, i.e., 66 − 53. To ensure that the behavior of new x86 family processors are consistent with existing processors, subsequent processors have continued to represented π internally using 66-bits (rather than the 128-bits needed to achieve an accuracy of better than 1.5 ULP)

What the fuccccck. I spend my time so far from the metal that I never really think about stuff like this, but it is interesting. I’ve hit limits with float maximums for similar reasons, but it is insane to think that our machines are rough-grained sieves that need to reduce the resolution on math to be effective. Weird.

Chapter 6 has some interesting summaries of testing, but nothing jumps out after a cursory reading. Worth going back to again when specifically thinking about testing and fault tolerance.

Chapter 7! Source code!

Building a software system involves arranging source code in a way that causes the desired narrative to emerge, as the program is executed, when processing user inputs.

More focus on emergent behaviors!

Creating a program requires explaining, as code, everything that is needed.

Code is just very, very precise documentation!

Source code is the outcome of choices made in response to implementation requirements, the culturally derived beliefs and experiences of the developers writing the code (coupled with their desire for short-term gratification), available resources, and the functionality provided by the development environment.

Source code is aliiiiiive!

Developers have a collection of beliefs, and mental models, about the semantics of the programming languages they use, as well as a collection of techniques they have become practiced at, and accustomed to using.

Vernaculars! Patterns! Approaches! Yeah! Source code has style because there are many variants that produce the same behaviors and outcomes! See also 100 days of fizz buzz.

Acquiring an understanding of the behavior of a program, by reading its source code, is not an end in itself; one reason for making an investment to acquire this understanding, is to be able to predict a program’s behavior sufficiently well to be able to change it. By reading source code, developers acquire beliefs about it, which are a means to an end; understanding a program is a continuum, not a yes/no state.

Again, much like literature. Also, architecture. Code has a topology, and therefor shares similarities with structures that have geometry.

The paper in which McCabe1229 proposed what he called complexity measures contains a theoretical analysis of various graph-theoretic complexity measures, and some wishful thinking that these might apply to software; no empirical evidence is given.

Seems like we have no reliable way to measure source code complexity as a metric. That is interesting, but aligns with my assumptions about quant metrics in this regard. Source code is a fundamentally qualitative endeavor. Jones also supports the idea that applying quant metrics can be hazardous:

This metric also suffers from the problem of being very easy to manipulate, i.e., is susceptible to software accounting fraud.

When metrics become goals, they cease being good metrics, see https://en.wikipedia.org/wiki/Goodhart%27s_law

What characteristics are desirable in source code?

Let’s talk about that good good qual! And moral values!

Use of the terms maintainability, readability and testability are often moulded around the research idea, or functionality, being promoted by an individual researcher, i.e., they are essentially a form of marketing.

Woah.

Stylistically, guideline documents are often more akin to literary criticism than engineering principles, i.e., they express personal opinions that are not derived from evidence.

It true! Harsh, but true! Jones elaborates further than any particular opinion regarding how to write code – structure and patterns to use or avoid – is a product of qualitative and experiential opinions, and has no basis in observation. Which is fine and to be expected, but it means that there is no ground here for representing it as anything but opinion, taste, and style. The core question we’re trying to answer becomes:

How might source code be organized to minimise the expenditure of cognitive effort per amount of code produced?

This is the goal! We want to write code from which our desired behavior emerges, and we want to minimize the effort it takes to understand, conceptualize, and alter that code to address new desired or extant undesired emergent behaviors.

Jones quickly identifies the process of writing modules of software with clearly defined I/O surface areas as a core method of reducing the amount of code one needs to read in order to understand the system. If a module called add has a function signature of 2 integers, and returns a single integer, and says the return is the result of adding the two, we can be done. We don’t need to care about implementation details at that point – we do need to keep a pin in it though if it appears to not be behaving as we expect, in which case the implementation needs to be read and understood.

Jones indentifies that modularity commonly occurs in

biological systems where connections between components have a cost (e.g., they cause delays in a signalling pathway), and modularity is an organizational method that reduces the number of connections needed;374 modularity as a characteristic that makes it easier to adapt more quickly when the environment changes may be an additional benefit, rather than the primary driver towards modularity. Simulations1263 have found that, for non-trivial systems, a hierarchical organization reduces the number of connections needed (for a viable implementation).

Which is super interesting, as that’s what we also tend to use module for in software systems.

Pivoting to code, and the use of modules in a dependency structure, Jones states that:

Dependencies between units of code can be used to uncover possible clusters of related functionality. One dependency is function/method calls between larger units of code, with units of code making many of the same calls likely to have something in common.

This is a min-cut analysis of a dependency graph. Decomposing a complete dependency graph into clusters of closely related modules efficiently would be a very interesting exercise indeed.

Beyond this level of mathematical analysis — which wont happen at a human scale beyond rough pattern recognition — Jones suggests that narrative structure can be a guide to code structure, since it functions in very similar ways.

Developers do not understand programs, as such, they acquire beliefs about program behavior; a continuous process involving the creation of new beliefs and the modification of existing ones, with no well-defined ending.

Haha damn, tho that’s true. Full comprehension tends to be elusive, and instead we tend to generate a shibboleth and work off that. This work is … conceptual labor.

So, to make our question from above actionable:

What can be done to reduce the cognitive effort that needs to be invested to obtain a good-enough interpretation of the behavior of code?

There are interesting hints that language is an overlay for the same complicated graph structure mathematics that dependency analysis can reveal. Jones demonstrates some subject/predicate/object examples where a set of statements gets comprehended as three semantic units. The semantics of the graph remain with the reader, even as the precise statement syntax does not.

Jones, as always seeking evidence to discuss, moves to examining reading comprehension studies that focus on order, dependency closeness, and maps those concepts onto code structures. This is a good impulse, as we’ve discussed already that code is a practice in literature and conceptual labor, but is complicated by the fact that code is non-linear and reading order is not actually related to line order. It’s a relatively simple process to skip lines or read code backwards, searching through a codebase to create topological relatedness in the code. So measuring code shape in that way is perhaps less compelling, tho there are some useful parallels, since “Putting too much information in one sentence has costs”. This is the root basis for function decomposition and separation of concerns in code structures.

Jones also mentions here that the visual shape of the code, combined with the expectations of the reader, can make a difference in the amount of time to takes to achieve comprehension. An example is a study where people read inverted, mirrored, or inverted and mirrored text. After being exposed to it, they adapter their expectations of the text and got much better quickly. This mirrors my own experience setting lead type for letterpress — after an initial exposure and constant practice, I became quite literate in reading reverse & mirrored text. Jones also identified patterns and norms of code as guiding principles for writing. Reproducing what is already written, basically, allows for pattern recognition to hold quicker when attempting to comprehend a codebase. Identifier and variable names also clearly cary a load of semantics that a reader to use to apply assumptions (hopefully productively).

On the subject of good identifiers, Jones continues for some length. To sun up, the goal with an identifier is simply appropriate semantic meaning to the task, and reducing omission and substitution errors across identifiers. For instance, when identifying a coordinate pair, lat and long are more appropriate than lang and long — lang is at once carrying the wrong semantic meaning (usually used to indicate language rather than lattitude) and has an easy substitution error (a for o) that won’t cause any errors, but will cause unexpected behaviors.

In chapter 7.3, Jones takes a looks at patterns of usage. It’s clear, and demonstrated in the text, that patterns are an important tool when writing code. Patterns emerge from contexts and uses, and have powerful uses. But, as we also know from Alexander, patterns are not enough on their own to provide a positive value — patterns are largely value-agnostic, and can be used to great significant positive or negative impacts on a process. Some of the more egregious forms of this are termed Dark Patterns or Anti-patterns, but even without being explicitly negative or a bad implementation, patterns can freight a hegemony into a codebase that does not serve the end projects values. Jones than moves on to a brief mention of the Sapir-Whorf hypothesis, which I think captures my above sentiments on patterns well.

Lots of Chapter 7.3 examines things like control flow, typing, variable naming, basic sort of building blocks of codebases. Quite a bit of it is “this is how these things work in code” and less value judgement around effectiveness. There are definitely some gestures towards good practices that could be extracted and discussed at length, but also a good amount of backing for saying “there’s no evidence that this or that makes any difference”, ie:

To summarise: when a language typing/feature effect has been found, its contribution to overall developer performance has been small.

Or:

While studies have found that identifiers are sometimes declared with greater visibility than necessary (given their existing use in code), there has not been any analysis of the cost/benefit of supporting potential future unintended/intended uses.

The project of extraction these insights into guidance for code reviews is for another day, since Im mostly reading in search of metrics, but chapter 7.3 is a great resource place for that.

Lets move on to 7.4 and talk about how code grows:

Factors influencing the rate of evolution of source code characteristics include: • customer limited: insufficient customer demand (as measured by willingness to pay) for it to be economically worthwhile updating existing functionality (to support changes in the world), or adding new functionality, e.g., new hardware requiring device drivers, • developer limited: bottlenecks in the development process that restrict the quantity of change per unit time. For instance, a limited number of people with the necessary skills, change requests requiring sign-off by a handful of senior managers, or increasing developer resources required to support a growing system leading to diminishing returns from adding more developers, • competition from other applications: source code may cease to evolve because its host, the application, is out-competed, e.g., customers stop using the application and/or it looses developer mindshare

Elsewhere

software-development
data-analysis

© 2026