ShowcaseDaniel Coakley
A Building Energy Model Can Match the Bills and Still Be Wrong
Building Energy SimulationDigital TwinsResearchData Analytics

A Building Energy Model Can Match the Bills and Still Be Wrong

Why evidence, uncertainty and hourly patterns matter more than a single 'calibrated' result

2014
By Daniel Coakley

A calibrated building model can look reassuringly precise. But if it reaches the right energy total for the wrong operational reasons, the retrofit decisions built on it may still be unreliable.

Building_Energy_Models

Imagine a model being used to justify a major investment in heating, controls or fabric upgrades. It reproduces the monthly utility bill, so it is declared calibrated. Confidence rises. Yet behind the total, the model may be using the wrong occupancy schedule, the wrong weather data, an inaccurate HVAC sequence or assumptions copied from a design document that no longer reflects the building.

That is the central problem explored in my doctoral research on the calibration of detailed building energy simulation models. The challenge is not simply to make a simulated number resemble a measured number. It is to build a traceable explanation of how the building actually behaves, acknowledge what remains uncertain and carry that uncertainty into the decision.

The core idea

Calibration is not curve fitting. It is an evidence-led investigation into the causes of building performance.

The uncomfortable truth: many models can fit the same bill

Detailed building simulation is an over-parameterised and underdetermined problem. A model contains large numbers of interacting inputs — construction properties, weather, occupancy, lighting, equipment, controls, ventilation and plant performance — while the available measured outputs are comparatively limited.

The result is non-uniqueness. Several different combinations of inputs can produce a similar monthly energy total. A statistically acceptable fit is therefore necessary, but it is not proof that the model is physically correct. A model can be right in aggregate and wrong in its internal explanation.

Two profiles, one monthly total

Daily energy profile (kWh) — both sum to ≈ 360 kWh

00020406081012141618202206121824
  • Design assumptions
  • Actual operation
Figure 1. Illustrative example: matching a daily or monthly total does not guarantee that a model reproduces the building's operating pattern.

Treat every model input as a claim that needs evidence

The proposed methodology starts by treating model inputs as evidence-backed claims. Each input should identify its source: a sensor, a spot measurement, a physical inspection, an as-built drawing, an O&M manual, a design document, a standard or, where nothing better exists, a default assumption.

This matters because evidence quality determines how tightly an input can be trusted. A live sensor reading should normally justify a narrower uncertainty range than a generic reference value. The thesis proposed the following indicative hierarchy as an initial way to translate source confidence into ranges of variation. These percentages were explicitly presented as preliminary heuristic estimates, not universal tolerances.

Source evidenceClassIndicative range
BMS or sensor data1±2%
Spot-measured or physically verified2±5%
As-built drawings, O&M or commissioning documents3±10%
Design documents4±15%
Guides and standards5±30%
Reference manual or default values6±40%
No available information7±50%

Table 1. Indicative evidence hierarchy and ranges of variation proposed in the thesis. The values are starting assumptions to be refined, not generic accuracy guarantees.

The practical benefit is accountability. If a result changes, the analyst can trace the change back to a revised assumption and the evidence supporting it. The model becomes an auditable chain of reasoning rather than an opaque file that merely produces plausible outputs.

Calibration should be a controlled investigation

The method combines seven stages: define the modelling purpose and acceptance criteria; collect and classify data; build an initial evidence-based model; compare it with measured performance; investigate uncertain and influential inputs; search the plausible parameter space; and communicate the range of valid outcomes.

1

Define purpose & criteria

Establish modelling objectives and statistical acceptance thresholds

2

Collect & classify data

Gather evidence from sensors, documents and audits; rate source quality

3

Build evidence-based model

Construct the initial model from verified, source-attributed inputs

4

Compare with measured data

Evaluate fit against utility meters and sub-meter data at multiple time resolutions

5

Investigate uncertain inputs

Use sensitivity analysis to identify parameters that are both uncertain and influential

6

Search parameter space

Sample plausible ranges with Latin Hypercube Monte Carlo to find acceptable model families

7

Communicate outcomes

Report ranges of valid results, remaining uncertainty and the evidence behind each assumption

Figure 2. A practical evidence-led calibration workflow adapted from the thesis methodology.

Version control is a surprisingly important part of this process. Each model revision should record what changed, why it changed and which source justified the change. This allows performance improvements to be linked to specific corrections and prevents calibration from becoming a sequence of undocumented adjustments.

The case study: the building told a clearer story than the drawings

The method was demonstrated on the Nursing Library at NUI Galway, a naturally ventilated university building. The case study assembled information from building documentation, audits, spot measurements, computer login records, occupancy observations, weather data, the building management system and energy meters.

Several revisions illustrate why this breadth of evidence matters. A generic weather file was replaced with local measured weather because heating demand was strongly temperature-dependent. Default constructions were updated using more reliable as-built and manufacturer information. Occupancy assumptions were revised using audits and computer-use data. A discrepancy in heating profiles was eventually traced to an incorrect HVAC schedule and corrected using interviews and BMS evidence.

Outside air temperature vs. daily heating demand

Nursing Library, NUI Galway · R² = 0.955

071425Outside air temperature (°C)050100150200Heating demand (kWh)
Figure 3. The case study showed a strong inverse relationship between outside air temperature and daily heating demand (R² = 0.955). Source: thesis Figure 4-25.

What changed the model. Better evidence — local weather, observed occupancy, verified constructions and actual control schedules — mattered more than indiscriminate parameter tuning.

Hourly patterns reveal what monthly totals conceal

Monthly statistics are useful for a high-level check, but they can hide timing errors that are critical to retrofit analysis. Comparing profiles by hour of day, day of week and month makes it easier to identify scheduling faults, incorrect baseloads, holiday effects, anomalous sensor data and weather-related discrepancies.

In the case study, a carpet plot turned a year of electrical data into an immediately readable operating signature. It exposed weekday and weekend schedules, the lower-intensity summer period and anomalous intervals. The inferred main equipment schedules were approximately 07:00–23:00 on weekdays, 08:00–18:00 on Saturdays and 10:00–18:00 on Sundays. These patterns would be almost invisible in a monthly total.

Annual electrical energy carpet plot

365 days × 24 hours · colour intensity = energy use

00061218JanFebMarAprMayJunJulAugSepOctNovDec
Low
High
Figure 4. A carpet plot makes operating schedules, seasonal shifts and anomalies visible at a glance. Source: thesis Figure 4-30.

Prioritise the inputs that are both uncertain and influential

Not every unknown deserves the same investigative effort. The methodology uses sensitivity analysis to estimate which inputs have the greatest effect on model outputs, including their interactions with other variables. The most valuable targets for further measurement are therefore the parameters that combine high uncertainty with high influence.

High uncertaintyLow uncertainty
High influenceMeasure, audit or verify firstRetain and monitor; the evidence is already strong
Low influenceDocument the limitation, but avoid disproportionate effortFreeze the parameter and move on

This is a practical optimisation of the calibration process itself. It directs limited survey, metering and analytical resources toward the assumptions most capable of changing the decision.

One 'best' model creates false precision

Once plausible ranges have been assigned to uncertain inputs, the research proposes sampling many combinations within those bounds using a Latin Hypercube Monte Carlo approach. Each simulation is ranked against measured data using goodness-of-fit indicators such as normalised mean bias error and the coefficient of variation of root mean square error.

The point is not to discover a single magical parameter set. It is to identify a family of models that are both evidence-plausible and statistically acceptable. Those models can then produce a distribution of predicted retrofit outcomes rather than a single deterministic number.

A better decision output. Replace 'the measure will save 18%' with a transparent range, its confidence and the assumptions that drive the spread.

Five questions to ask before trusting a calibrated model

For asset owners, energy managers and retrofit decision-makers, the methodology translates into five simple due-diligence questions.

  1. What evidence supports the critical inputs? Ask for the source of occupancy, fabric, plant, control and weather assumptions — not only the final values.
  2. Was the model checked at the right time resolution? A monthly fit may be inadequate where schedules, controls or peak demand determine the investment case.
  3. Which assumptions dominate the result? Sensitivity should guide both investigation and the way risks are communicated.
  4. Can each revision be traced? The audit trail should connect a model change to an observation, measurement, document or explicit judgement.
  5. Are savings shown as a range? A decision-ready model should reveal plausible outcomes and remaining uncertainty, not only a point estimate.

From simulation file to defensible decision

The value of a calibrated energy model does not come from complexity alone. It comes from the quality of the evidence, the transparency of the assumptions and the model's ability to reproduce the building's behaviour at the resolution relevant to the decision.

A good model is therefore not simply a prediction engine. It is a structured argument: this is what we observed, this is what we inferred, this is what remains uncertain and this is how that uncertainty affects the decision.

Building performance analysis should not aim only to be more detailed. It should aim to be more auditable, more honest about uncertainty and more useful for making decisions under risk.


Adapted from Daniel Coakley's doctoral thesis, Calibration of Detailed Building Energy Simulation Models using an Analytical Optimisation Approach (September 2012). Case-study graphics are reproduced from the supplied thesis; the workflow and illustrative comparison graphic were created for this article.

Discuss this research

Want to explore how these findings apply to your portfolio or organisation?

Get in touch