Battery Management: Estimating State of Charge
A battery management system knows two things about a cell: how much current is flowing, and what voltage the terminals are at. Neither is the state of charge. The current only gives you the charge if the current sensor is telling the truth, and it never tells you whether it is. The voltage says something, but only through a curve, and the curve is not straight.
This case study builds the estimator that a real BMS runs: an extended Kalman filter that watches both signals, learns the offset of the one it cannot trust, and reports the charge together with its own uncertainty.
The Cell
A 18650 lithium-ion cell, as a Thévenin circuit: an ohmic resistance, one polarisation branch, and an open-circuit voltage that depends on the charge.
The four numbers below are a worked example rather than a datasheet. Panasonic's own sheet for the NCR18650B gives 3.25 Ah minimum and bounds the internal resistance from above at 100 mΩ, so neither the 3.20 Ah nor the 0.05 Ω is quoted from it. They are chosen to suit the sampling rather than to suit a cell: one pole at 0.05 rad/s, and twenty samples per time constant.
| Parameter | Meaning | Value |
|---|---|---|
| Capacity | ||
| Ohmic resistance | ||
| Polarisation resistance | ||
| Polarisation capacitance | ||
| Polarisation time constant |
Two states. One of them — the polarisation voltage — is easy, because it is just the difference between the terminal voltage and what the charge says the cell ought to read. The other is the problem. The charge is not on a gauge.
Why Counting Is Not Enough
Integrate the current, divide by the capacity, subtract from the start. No model, no curve, nothing to tune, and for an hour it looks perfect.
The catch is that the meter is not the cell. An error in the current sensor becomes charge that never comes back:
The rate follows the absolute offset in amperes, and the capacity. It does not follow the load.
| Load | Sensor offset | Drift | |
|---|---|---|---|
| 2 A | 2 % | 0.04 A | 1.25 % per hour |
| 10 A | 2 % | 0.20 A | 6.25 % per hour |
| 10 A | 5 % | 0.50 A | 15.6 % per hour |
| 2 A | 10 % | 0.20 A | 6.25 % per hour |
| 10 A | 10 % | 1.00 A | 31.3 % per hour |
The last two rows are the point, and the first two are the trap. A percentage offset of a bigger load is a bigger absolute offset, so the table reads as though the drift scaled with the load. Hold the absolute offset at 0.20 A and the rate is 6.25 % per hour whether the load is 2 A or 10 A. A shunt is specified as a percentage, which is why this is easy to get wrong.
Two things follow. The drift is a ramp, so it is invisible on a 0…100 % charge axis — a chart of absolute charge looks perfect the entire time, and that is the failure. And no load makes the number smaller, so the honest answer is not a gentler current. It is a second opinion about the charge.
What the Voltage Reveals
Remove the current and the cell settles. The polarisation voltage decays, and what is left on the terminals is the open-circuit voltage — a function of the charge and nothing else.
This is the way in. But the curve is not straight, and how straight it is varies enormously:
| SOC | against the plateau | |
|---|---|---|
| 0–10 % | 5.00 V per unit charge | 8.3× |
| 10–20 % | 1.50 | 2.5× |
| 20–90 % | 0.60 – 0.90 | the plateau |
| 90–100 % | 0.60 | 1.0× |
The same 10 % of charge is worth half a volt at the bottom of the curve and six millivolts in the middle. That factor of 8.3 is not a curiosity: it decides the observability of the model directly. With the measurement row and a diagonal state matrix, the observability determinant is
so the charge is observable exactly as long as the curve is not flat, and how well observable it is scales with that slope.
A Linear Observer, and Where It Fails
The classic estimator takes the measurement, subtracts what the model expects, and corrects the state with a gain derived from the two noise levels. It is the right machinery, and it has one decision that is yours to make: the charge-to-voltage relation is curved, so the estimator linearises it.
Once. The model becomes affine:
The slope is one choice. The intercept in brackets is another, and it is the one that decides where the filter settles — because a linear observer converges to the point where its own measurement model reproduces the measured voltage, and that is not where the cell is.
Linearise at one operating point, sit at rest, and the result is a systematic bias, not a delay:
| Tangent taken at | True charge | Settled bias |
|---|---|---|
| 0.45 | 0.05 | −53.3 % |
| 0.45 | 0.85 | +7.5 % |
| 0.05 | 0.45 | −28.8 % |
| 0.05 | 0.05 | −0.0 % |
With the tangent at 45 % and the true charge at 5 %, the estimate comes to rest at −0.48: a charge no cell can have. Only a tangent drawn near the true operating point produces a small bias, which means one choice of working point is a decision you pay for over the whole run.
There is a second, counter-intuitive detail. The damage comes from a tangent that is too steep, not too shallow. At the wrong slope the filter believes a noisy measurement far more than it should, so its gain collapses and the estimate crawls — and where it stops is 63 % from the truth, which is the last row of the table above. A too-gentle tangent is comparatively harmless: the innovation is large enough to pull the estimate to the truth anyway.
The Extended Kalman Filter
Nothing about the machinery changes. One line does — instead of freezing the slope at the start, look it up where the estimate currently is:
The curve is not linearised once and then trusted. It is linearised again, at the point the filter has reached, so the sensitivity it uses is the sensitivity that is actually there. The EKF is a Kalman filter that admits it is approximating.
| Tangent at | True charge | Linear bias | EKF bias |
|---|---|---|---|
| 0.45 | 0.05 | −53.3 % | 0.0 % |
| 0.45 | 0.85 | +7.5 % | −0.0 % |
| 0.05 | 0.45 | −28.8 % | −0.0 % |
| 0.05 | 0.85 | −63.1 % | +0.0 % |
One set of and , both regions, no bias anywhere. And the EKF is not uniformly faster: in the steep part it takes 71 s against the linear filter's 2 s. The re-linearisation is cautious exactly where the measurement is genuinely informative. That trade is the subject of the tuning step, not a flaw to be hidden.
Starting From the Wrong Answer
The interesting case is not the filter agreeing with a patient experiment. It is the filter finding the charge from nothing: estimate starting at 0.50, true charge 0.90, cell at rest, nothing tuned to make it easy.
It converges in about three seconds. Not "eventually" — the settling curve against time is the whole point of the step, and the two traces sit on top of each other afterwards.
The EKF has one real weakness, and it is worth naming. The filter starts on the flat plateau, where the curve is uninformative, so the first correction is computed with a small slope and overshoots into the steep tail. It lands where the slope is 8.3 times larger — but the covariance it has already collapsed there was computed for the wrong slope, and a collapsed covariance is a small gain. It recovers, slowly. That is why the settling time is about 3 s at 0.45, 71 s at 0.05, and 71 s at 0.15.
Tuning
is what the model admits it does not know. is what the sensor admits it does not know. The gain is the argument between them, and there is no setting that is right — only one that is right for the situation the cell is in.
Raise and the filter leans on its model: smooth, and slow to notice anything. Lower it and the filter leans on the voltage: responsive, and it believes the noise. moves the other side of the same balance, and it is subtler. The charge state is close to a pure integrator, so its dominant error is the current sensor's gain, which is constant and multiplicative rather than white — an additive is not sensor noise at all, it is admitted model error. It is a density per second, so it accumulates:
| σ after one hour | What it represents | |
|---|---|---|
| 1e-6 | 6.0 % charge | more spurious uncertainty than the quantity being measured |
| 1e-8 | 0.6 % charge | unmodelled drift: temperature, ageing |
The Sampling Question
Both filters run at 1 Hz, which is where a BMS lives. That is comfortably inside the requirement rather than marginal: the fastest pole in the plant is , giving 20 samples per time constant and 63× of margin against the Nyquist limit. The only pole in the discrete model sits at 0.9512, so the filter is unconditionally stable and no step-size limit applies.
The plant and the filter are discretised separately and exactly, over the same interval:
Mixing an exact with a forward-Euler is a real error, not untidiness: the second term is 2.52 % above the exact value (the exact value is 2.46 % below it) — and either way that difference surfaces in the observer's innovation, where the filter has no way to tell it from an error in the model. The block recomputes the discretisation — the noise included, by Van Loan — from the parameters, so changing the cell's capacitance or the sample time changes the estimator's own model with it.
Further Reading
The model this case study is built on is not described here from nothing. The
thevenin package from the US National Renewable Energy Laboratory states the
same equations this guide does —
$$\frac{d\,SOC}{dt} = \frac{-I\,\eta}{3600\,Q_{\max}} \qquad
\frac{dV_j}{dt} = -\frac{V_j}{R_jC_j} + \frac{I}{C_j} \qquad
V_{\text{cell}} = V_{\text{OCV}}(SOC) - \sum_j V_j - I R_0$$
— with the hysteresis term set to zero, which is what the two-branch fit in
"Further Reading" below is for.
- Randall, C. R. thevenin: Equivalent circuit models in Python (NREL, SWR-24-132).
BSD-3-Clause.
— the model description, with the equations above and the cell's own parameters
as inputs. The four numbers in this guide are not from it; they are a worked
example.
- Thevenin equivalent-circuit modelling of lithium cells; the two-branch extension and
the second RC time constant that a single-branch model leaves out
- Extended and iterated extended Kalman filtering for battery management, and the
role of the OCV curve's slope in the achievable observability
- The normalised innovation squared, which is the number that tells you whether the
assumed and are the ones the cell actually has
Trying It Yourself
The last step of the interactive case study is a playground, and the two things
worth dragging are the ones this guide says are the argument:
- Measurement noise variance and process noise , the two
quantities in the table above. Both are on decade sliders, because four decades
of on a linear track puts every setting worth having in the first twentieth
of the travel.
- The true charge of the cell. Put it at 15 % and the estimate takes 71 s to
the same single point that 90 % reaches in 28 s, for the reason in the plateau
paragraph above.
The load, the sensor offset and the sensor's own noise are there too, and they are
the three that step 2 of the lesson is about.