Sizing a Symmetrical OTA on SKY130 in Ten Rounds of Conversation

Posted on September 27, 2026

Category: Technology

Tags: analog-ic-design, cmos, symmetrical-ota, ota, sky130, ngspice, gm-id, transistor-sizing, pvt-corners, monte-carlo, offset, python, ai-agents, circuit-design

Views: 10

Sizing a Symmetrical OTA on SKY130 in Ten Rounds of Conversation

In Drawing a Symmetrical OTA in Seven Rounds of Conversation the circuit was only a picture: schemdraw code, an LTspice netlist, and placeholder device sizes. This post picks up that schematic and sizes it for real, on the SkyWater SKY130 open PDK, simulated with ngspice.

As before, I am keeping the requests in the order I made them, including the ones where the answer was "that cannot be done." The design ended up where it did because of those.


The circuit and the spec

The topology is unchanged from the drawing: NMOS input pair M1/M2M_1/M_2 with tail M0M_0, diode loads M3/M4M_3/M_4, 1:B1:B PMOS mirrors M5/M6M_5/M_6, and a 1:1 NMOS mirror M7/M8M_7/M_8 that folds the left branch back to the output. One device was added: a diode-connected MBM_B that takes the reference current IBI_B and biases the tail through a 1:1 mirror.

Symmetrical OTA schematic: NMOS differential pair with tail transistor, diode-connected PMOS loads mirrored 1:B to the two output branches, and an NMOS mirror returning the left branch to the output, which drives a 2 pF load capacitor.

The specification I gave:


Round 1 — "Design it"

Me:

We drew the symmetrical OTA in the schdraw project. Here, let us size it. Use ngspice or LTspice, and the SKY130 process — check whether the PDK is already on disk, and download the models if not. This is a simple circuit; your circuit knowledge should be enough.

The PDK was already installed through volare, so ngspice could load sky130.lib.spice directly. Before touching a simulator, Claude did the arithmetic that decides the current.

The single-stage OTA's unity-gain frequency is set by the input pair's transconductance, multiplied by the mirror ratio:

fu=B gm12πCL>5 MHz  ⇒  gm1>31.4 μS f_u = \frac{B\,g_{m1}}{2\pi C_L} > 5\,MHz \;\Rightarrow\; g_{m1} > 31.4\,\mu S

With IB=2 μAI_B = 2\,\mu A, each input device carries 1 μA1\,\mu A and would need gm/ID≥31g_m/I_D \ge 31. That is past the weak-inversion ceiling of roughly 1/(nUT)≈251/(n U_T) \approx 25–3030, so 2 μA2\,\mu A is out and 4 μA4\,\mu A is the first candidate (gm/ID≥16g_m/I_D \ge 16). The total supply current is then fixed by the topology:

IDD=IB+Itail+2⋅B Itail2=4 IB=16 μA I_{DD} = I_B + I_{tail} + 2\cdot\frac{B\,I_{tail}}{2} = 4\,I_B = 16\,\mu A

The first simulation met UGF and phase margin but not gain (44.6 dB44.6\,dB), and at VOUT=1.65 VV_{OUT} = 1.65\,V the output PMOS M6M_6 fell out of saturation: VDSAT=0.171 VV_{DSAT} = 0.171\,V against VDS=0.15 VV_{DS} = 0.15\,V. Longer, wider output devices fixed both. A small sweep settled on M7/M8=10/3M_7/M_8 = 10/3 and PMOS 6/16/1, giving 47.8 dB47.8\,dB, 6.74 MHz6.74\,MHz and 75.2∘75.2^\circ at 15.5 μA15.5\,\mu A.

The claim that 2 μA2\,\mu A cannot work was also checked by brute force: an input pair of 20/0.5×420/0.5 \times 4 reached gm/ID=26.5g_m/I_D = 26.5 and a UGF of only 3.95 MHz3.95\,MHz.

A short aside followed. In the report I had called M6/M8M_6/M_8 the "output stage," and the reply was:

What is this two-stage output? You are not designing a two-stage amplifier, are you?

No. It meant the output branch of the single-stage OTA. A small wording slip, but in analog design "output stage" names a whole architecture.


Round 2 — "Bound the gm/Id"

Me:

Use at most about 24 for the input pair's gm/IDg_m/I_D, and keep every other device below 20.

The existing design already complied — 22.922.9 for the pair, at most 17.217.2 elsewhere — so nothing changed except the harness, which now flags any device over its limit. The new bound did make the 2 μA2\,\mu A case even more clearly impossible: it needs 31.


Round 3 — "How did you pick the lengths?"

Me:

How did you decide the default channel lengths?

The honest answer: mostly by habit. The tail at 2 μm2\,\mu m and the input pair at 1 μm1\,\mu m were rules of thumb. Only two groups had actually been swept — the NMOS mirror (gain) and the PMOS mirrors (phase margin) — and even those on a coarse grid. There was no gm/IDg_m/I_D table behind any of it.


Round 4 — "Build the table, then minimise size across the worst case"

Me:

Make a gm/IDg_m/I_D table and choose the lengths again. Use lengths in multiples of 0.2 μm0.2\,\mu m, at least two units. Shrink the devices as far as the spec allows — and the spec includes the worst case.

(I first asked for multiples of 0.13 μm0.13\,\mu m and changed my mind mid-run.)

The characterisation swept VGSV_{GS} for both device types at nine lengths from 0.40.4 to 4 μm4\,\mu m. Two rows tell most of the story:

LL (μm\mu m) 0.4 0.6 1.0 2.0 3.0 4.0
nfet gm/gdsg_m/g_{ds} at gm/ID=17g_m/I_D = 17 108 191 172 202 258 303
pfet gm/gdsg_m/g_{ds} at gm/ID=14g_m/I_D = 14 86 196 459 766 1275 1798

The NMOS intrinsic gain barely moves between 0.60.6 and 2 μm2\,\mu m, while the PMOS gain climbs steadily. Since the output resistance is ro6∥ro8r_{o6} \parallel r_{o8}, the DC gain in gm/IDg_m/I_D terms is

A0≈(gm/ID)2(gm/ID)6/A6+(gm/ID)8/A8 A_0 \approx \frac{(g_m/I_D)_2}{(g_m/I_D)_6/A_6 + (g_m/I_D)_8/A_8}

where AA is each device's gm/gdsg_m/g_{ds} — and the NMOS term is the bottleneck.

"Worst case" was defined as five process corners (tt, ss, ff, sf, fs), three temperatures (−40-40, 2727, 85∘C85^\circ C) and VDD±10%V_{DD} \pm 10\%: 45 conditions, with the input common mode at VDD/2V_{DD}/2. Each corner runs as one ngspice process that loops over temperature and supply internally, so all 45 conditions take under five seconds on 16 cores. A greedy optimiser then tries, for each device group, shortening LL (with and without keeping W/LW/L) or narrowing WW by 15 %, and keeps the smallest candidate that still passes all 45.

The starting point — the Round 1 design — already passed, but thinly: worst-case UGF 5.65 MHz5.65\,MHz at fs/85∘C85^\circ C/1.62 V1.62\,V, and only 12 mV12\,mV of saturation margin on the tail.


Round 5 — "Plot the AC response"

Me:

Save the AC result as a PNG plot too.

Every plot in this project shows the nominal condition as a single bold line and the other 44 PVT conditions as thin grey lines behind it, so the spread is visible at a glance.

Open-loop gain and phase of the starting design across 45 PVT conditions. Nominal gain 48 dB, unity-gain frequency 6.85 MHz, phase margin 74.9 degrees; worst case 47.0 dB, 5.65 MHz and 73.8 degrees. All grey curves stay in a narrow band around the nominal one.


Round 6 — "Put a floor under the saturation margin"

The first optimisation run shrank the gate area from 136136 to 56.4 μm256.4\,\mu m^2 and passed everything — with 1 mV1\,mV of saturation margin on M8M_8 and 6 mV6\,mV on M0M_0. That is exactly what "smallest area that meets the spec" produces: every constraint pushed onto its edge.

Me:

Add a 20 mV floor on the saturation margin and run it again.

Before it ran:

20 mV is too small too. Make it 100.

This is where the constraints collided. At the swing limit the output device has VDS=0.15 VV_{DS} = 0.15\,V, so a 100 mV100\,mV margin needs

VDSAT≤50 mV V_{DSAT} \le 50\,mV

and the table says an NMOS only gets there at gm/ID≳22g_m/I_D \gtrsim 22 — while Round 2 capped it below 20. At gm/ID=20g_m/I_D = 20, VDSATV_{DSAT} is about 6565–70 mV70\,mV even at room temperature. Swing, margin and the gm/IDg_m/I_D cap cannot all hold, and the table showed that before a single simulation ran.

Me:

Make it 50 mV.

Fifty was feasible, barely. The first generous starting point failed phase margin (65.5∘65.5^\circ) because the big mirrors loaded the mirror nodes; a grid over mirror sizes found gain and phase margin trading against each other with no combination passing both. The way out was the input pair: its drain capacitance sits on the same mirror nodes, and shortening it to 0.6 μm0.6\,\mu m bought the missing degrees. From there the optimiser reached 148.6 μm2148.6\,\mu m^2, with worst-case gain 45.03 dB45.03\,dB.


Round 7 — "Area is not a spec; current is"

Me:

I never gave you an area spec. It is fine if the area grows to meet the other specs. Reducing current matters more than area.

That settled the priority order: worst-case specs first, current second, area last. It also clarified something about the result. With the topology fixed at IDD=4IBI_{DD} = 4 I_B and 2 μA2\,\mu A ruled out in Round 1, 16 μA16\,\mu A is already the floor. Area was only ever the tie-breaker.

It also exposed a side effect of optimising for area. The optimiser had shrunk the tail and its reference to 1.2/0.41.2/0.4. A short tail device has low output resistance, so the VDSV_{DS} difference between M0M_0 and MBM_B costs tail current — and with it UGF.


Round 8 — "Would 0.8 µm on the tail help?"

Me:

Growing the tail to about L=0.8 μmL = 0.8\,\mu m seems fine. Would that give more margin?

Tail W/LW/L Worst UGF M0M_0 margin Worst A0A_0 Area (μm2\mu m^2)
1.2/0.4 5.30 MHz 51 mV 45.03 dB 148.6
2.4/0.8 5.50 MHz 51 mV 45.07 dB 151.4
4.0/0.8 5.57 MHz 67 mV 45.08 dB 154.0

It helped UGF and the tail's own margin, and did nothing for gain — which is set at the output node, not the tail. 4.0/0.84.0/0.8 went in.

Open-loop gain and phase of the final design across 45 PVT conditions. Nominal gain 46.1 dB, unity-gain frequency 6.67 MHz, phase margin 74.1 degrees; worst case 45.1 dB, 5.57 MHz and 73.0 degrees.


Round 9 — "Check the transient and the input common-mode range"

Me:

Check the transient response and the input common-mode range too. Save the plots as PNGs.

For the common-mode range, the output was held at VDD/2V_{DD}/2 by the same servo used everywhere else, and the input common mode swept from rail to rail while watching the smallest VDS−VDSATV_{DS} - V_{DSAT} among M0M_0–M4M_4.

Minimum saturation margin of the input-stage devices versus input common-mode voltage for 45 PVT conditions. The curves rise from the left as the tail transistor saturates, peak, and fall as the input pair leaves saturation. The worst-case band with at least 50 mV margin is 0.79 to 1.00 V.

The lower limit is set by the tail M0M_0, the upper by the input pair, whose drains sit at VDD−VSG3V_{DD} - V_{SG3} and therefore move with the supply. With a 50 mV50\,mV margin the worst-case range is 0.790.79–1.00 V1.00\,V. At VDD=1.62 VV_{DD} = 1.62\,V, a common mode of VDD/2=0.81 VV_{DD}/2 = 0.81\,V is only 20 mV20\,mV inside it.

For the transient, the OTA was connected as a unity-gain buffer and driven with steps that stay inside that worst-case range.

Unity-gain buffer step response for 45 PVT conditions. Left: a 0.8 to 1.0 V step with worst-case slew rates of 2.74 V per microsecond rising and 2.92 falling and settling to 1 percent within about 114 ns. Right: a 20 mV step with no visible overshoot and a small static error.

Worst-case overshoot is 1.9 %1.9\,\% on the large step and 0.1 %0.1\,\% on the small one, consistent with a 73∘73^\circ phase margin. The measured slew rate, about 2.7 V/μs2.7\,V/\mu s, is below the textbook BItail/CL=4 V/μsB I_{tail}/C_L = 4\,V/\mu s; a 200 mV200\,mV step probably spends part of its 20–80 % window in linear settling rather than pure slewing, but I did not verify that. The first slew numbers came out in suspicious steps — 2.842.84, 2.992.99, 3.143.14 — which was the 2 ns2\,ns sample grid showing through; interpolating the threshold crossings fixed it.


Round 10 — "And offset, by Monte Carlo"

Me:

Check the offset with Monte Carlo.

SKY130 exposes mismatch through *_mm library sections. The offset was defined as the differential input needed to put the output at VDD/2V_{DD}/2, which keeps the buffer's finite-gain error out of the number.

Reading the model source first turned up a bug in my own netlists. The mismatch terms scale as

σ∝1L W mult \sigma \propto \frac{1}{\sqrt{L\,W\,\mathrm{mult}}}

where mult is a subcircuit parameter that defaults to 1. The netlists passed only the SPICE multiplier m=2. The simulation was electrically correct, but every m-factor device would have received 2\sqrt{2} too much mismatch. Nothing in any plot so far would have shown it. The fix was to pass both, mult=2 m=2.

Histogram of 400 Monte Carlo offset samples at the tt mismatch corner, with a Gaussian fit. Mean 1.48 mV, sigma 4.84 mV; the other four corners give nearly identical sigmas between 4.79 and 4.89 mV.

The result was σ=4.84 mV\sigma = 4.84\,mV, almost twice my hand estimate. The model explained why: besides the threshold coefficient (3.36 mV⋅μm3.36\,mV\cdot\mu m for the nfet) there is a subthreshold-offset term of 7 mV⋅μm7\,mV\cdot\mu m, and in weak inversion it moves the current just like VTHV_{TH} does. The effective coefficient becomes

Aeff≈3.362+72≈7.8 mV⋅μm A_{eff} \approx \sqrt{3.36^2 + 7^2} \approx 7.8\,mV\cdot\mu m

and with that the per-pair contributions add up to the simulated number: about 4.0 mV4.0\,mV from the input pair, 1.6 mV1.6\,mV from each PMOS mirror, and 1.4 mV1.4\,mV from the NMOS mirror. Running every device at high gm/IDg_m/I_D to save current is what makes this term matter.

Me:

Under 5 mV is enough for now — sigma, not three sigma.

It passes, with about 0.1 mV0.1\,mV to spare.


The final design

Device W/LW/L (μm\mu m) m
MBM_B, M0M_0 4.0/0.8 1, 1
M1M_1, M2M_2 6.4/0.6 2
M3M_3, M4M_4 9.0/0.8 1
M5M_5, M6M_6 9.0/0.8 2
M7M_7, M8M_8 15.9/2.8 1
Ib vdd vbn 4e-06
CL vout 0 2e-12
XMB vbn vbn vss vss sky130_fd_pr__nfet_01v8 w=4.0 l=0.8 mult=1 m=1
XM0 tail vbn vss vss sky130_fd_pr__nfet_01v8 w=4.0 l=0.8 mult=1 m=1
XM1 x1 vinm tail vss sky130_fd_pr__nfet_01v8 w=6.4 l=0.6 mult=2 m=2
XM2 x2 vinp tail vss sky130_fd_pr__nfet_01v8 w=6.4 l=0.6 mult=2 m=2
XM3 x1 x1 vdd vdd sky130_fd_pr__pfet_01v8 w=9.0 l=0.8 mult=1 m=1
XM4 x2 x2 vdd vdd sky130_fd_pr__pfet_01v8 w=9.0 l=0.8 mult=1 m=1
XM5 y x1 vdd vdd sky130_fd_pr__pfet_01v8 w=9.0 l=0.8 mult=2 m=2
XM6 vout x2 vdd vdd sky130_fd_pr__pfet_01v8 w=9.0 l=0.8 mult=2 m=2
XM7 y y vss vss sky130_fd_pr__nfet_01v8 w=15.9 l=2.8 mult=1 m=1
XM8 vout y vss vss sky130_fd_pr__nfet_01v8 w=15.9 l=2.8 mult=1 m=1
Item Spec Nominal Worst of 45 PVT
DC gain >45 dB> 45\,dB 46.1 dB 45.1 dB
Unity-gain frequency >5 MHz> 5\,MHz 6.67 MHz 5.57 MHz
Phase margin >70∘> 70^\circ 74.1° 73.0°
Saturation margin ≥50 mV\ge 50\,mV — 51 mV
Offset σ\sigma <5 mV< 5\,mV 4.84 mV 4.89 mV
Supply current minimum 16 µA —

The DC-gain margin is a tenth of a decibel. At 4 μA4\,\mu A, with a 50 mV50\,mV saturation floor, DC gain and phase margin leave almost no room between them, and more area does not create any. The only lever left is current.


Afterwards — why is the NMOS mirror so long?

A later session went back over the design. The first question was about one number in the final table.

Me:

M7/M8M_7/M_8 have L=2.8 μmL = 2.8\,\mu m, far longer than the other NMOS devices such as the tail. Was there a specific reason? I usually give the tail and the mirrors the same length, which also makes them easier to match in layout.

There was a reason, and it is the gain spec. The DC gain of this OTA is

A0≈B gm1gds6+gds8 A_0 \approx \frac{B\,g_{m1}}{g_{ds6} + g_{ds8}}

M8M_8 sits on the output node, so its length goes straight into A0A_0. The tail does not. Its output resistance affects common-mode rejection and tail-current accuracy but not the differential gain, which is why lengthening it from 0.40.4 to 0.8 μm0.8\,\mu m in Round 8 moved A0A_0 by only 0.05 dB0.05\,dB. With a worst-case gain of 45.1 dB45.1\,dB against a 45 dB45\,dB spec, the optimiser could not shorten M7/M8M_7/M_8 any further. They take 8989 of the 154 μm2154\,\mu m^2 total.

The width follows from the length. At the bottom of the swing M8M_8 has VDS=0.15 VV_{DS} = 0.15\,V and needs 50 mV50\,mV of margin, so a long device has to be wide to keep VDSATV_{DSAT} down. Its worst-case margin is 51 mV51\,mV, on the edge as well. Pushing the length further runs into phase margin: a bigger M7M_7 loads the mirror node, and in Round 6 every combination with A0>45.5 dBA_0 > 45.5\,dB had a phase margin below 70∘70^\circ.

Me:

Couldn't you lengthen the PMOS mirror a little, shorten M7/M8M_7/M_8, and balance the two? Or even use nearly the same length for both?

Claude's first estimate assumed gds∝1/Lg_{ds} \propto 1/L for both devices. The Round 4 table says that is wrong for the nfet: its intrinsic gain is almost flat from 0.60.6 to 2 μm2\,\mu m, while the pfet's rises steadily.

LL (μm\mu m) 0.6 0.8 1.0 1.6 2.0 3.0
nfet gm/gdsg_m/g_{ds} 191 183 172 195 202 258
pfet gm/gdsg_m/g_{ds} 212 372 503 717 824 1342
gds,n/gds,pg_{ds,n}/g_{ds,p} 1.1 2.0 2.9 3.7 4.1 5.2

Both at gm/ID=17g_m/I_D = 17, tt, 27∘C27^\circ C, ∣VDS∣=0.9 V\lvert V_{DS} \rvert = 0.9\,V. At equal current and gm/IDg_m/I_D the two devices have the same gmg_m, so the ratio of intrinsic gains is the ratio of output conductances. Above 0.8 μm0.8\,\mu m the pfet's gdsg_{ds} is two to six times smaller, and the NMOS mirror dominates the output conductance.

With the gain budget normalised to gmg_m, the current design spends

1372+1250≈6.7×10−3 \frac{1}{372} + \frac{1}{250} \approx 6.7 \times 10^{-3}

and any pair of lengths that stays under that keeps A0A_0. Rough numbers from the table, holding W/LW/L fixed to keep VDSATV_{DSAT}:

The greedy optimiser could never have found the second option. It only shrinks devices, so trading length from one group to another was outside its moves. I decided a single-digit percentage was not worth a sweep for now, and none of these estimates were simulated. They come from a single tt point. But the reason for the long NMOS mirror in this PDK is now clear: it carries the gain, and in SKY130 its length buys gain far less efficiently than the PMOS's does. Matching it to the tail length for layout would cost gain or area here.


Afterwards — how the gm/Id table was built

The second question was about the method rather than the result.

Me:

The gm/IDg_m/I_D approach is very close to how I design, and I have met almost nobody else who works this way. Outside papers — mostly European ones — even textbooks barely explain it. Where did you pick it up?

Claude could not point to a single source: it learned from published literature, not from one book it remembers reading. The lineage itself is clear, though, and it is European. The method starts with Silveira, Flandre and Jespers at UCLouvain in 1996, and the EKV model from EPFL describes the same idea in terms of an inversion coefficient. It was later collected in Jespers' The gm/ID Methodology and in Jespers and Murmann's Systematic Design of Analog CMOS Circuits, whose companion lookup-table scripts spread it further. The standard textbooks still teach square-law sizing.

That matters here. Every device in this OTA sits between gm/ID=15g_m/I_D = 15 and 2323, in moderate inversion, where the square law is simply wrong. A characterised table is the natural tool in that region.

Me:

To build the table I have always used a diode-connected device. How did you build yours?

Not with a diode. sim/charz.py holds the drain at a fixed voltage and sweeps the gate:

Vd  d  0  ±0.9        * |VDS| = 0.9 V, about VDD/2
Vg  g  0  0 -> 1.8 V   * DC sweep, 5 mV step
X_i d g 0 0  fet  W=4 L=L_i   * nine lengths, 0.4 to 4.0 um, in parallel

The bench runs at tt and 27∘C27^\circ C with VSB=0V_{SB} = 0. It saves IDI_D, gmg_m, gdsg_{ds}, CggC_{gg} and VDSATV_{DSAT} at each point, then interpolates to target gm/IDg_m/I_D values from 10 to 24. WW drops out only at the end, as W=ID/(ID/W)W = I_D / (I_D/W).

Diode-connected (sweep IDI_D) Fixed VDSV_{DS} (sweep VGSV_{GS})
Saturation Guaranteed by construction, VDS=VGS>VovV_{DS} = V_{GS} > V_{ov} Holds over the range used; must be filtered if swept wide
Independent variable Current, easy to space logarithmically Voltage
gdsg_{ds} and gm/gdsg_m/g_{ds} Taken at VDS=VGSV_{DS} = V_{GS}, like a mirror reference (M3M_3, M7M_7, MBM_B) Taken at a chosen VDSV_{DS}, closer to output devices (M6M_6, M8M_8) and the input pair

Me:

Isn't the diode better, though? It lets you assume the device is always saturated, and gm/IDg_m/I_D means nothing outside saturation.

Half of that holds. The diode does guarantee saturation, in weak inversion as well, since VGSV_{GS} of a few hundred millivolts is far above the few kT/qkT/q that saturation needs. The fixed-VDSV_{DS} bench is saturated too wherever it was used. Across gm/ID=10g_m/I_D = 10 to 2424 the table's VDSATV_{DSAT} never exceeds about 0.17 V0.17\,V, against VDS=0.9 VV_{DS} = 0.9\,V. It only enters triode once the overdrive passes 0.9 V0.9\,V, at a gm/IDg_m/I_D far below anything in this design. A wider sweep would need a VDS>VDSATV_{DS} > V_{DSAT} filter.

The other half, that gm/IDg_m/I_D means nothing outside saturation, is right. What neither bench can promise is saturation in the circuit. A table built in saturation says nothing about M8M_8 at the bottom of the swing, where its VDSV_{DS} is 0.15 V0.15\,V. That is why the 45-corner harness checks VDS−VDSAT≥50 mVV_{DS} - V_{DSAT} \ge 50\,mV on every device, whichever way the table was made.

In saturation, gm/IDg_m/I_D, ID/WI_D/W and fTf_T come out nearly the same from both benches, because the drain current depends only weakly on VDSV_{DS}. They differ in the VDSV_{DS}-dependent quantities, gm/gdsg_m/g_{ds} and CgdC_{gd}. The real choice is where to read the intrinsic gain: at a mirror reference's VDS=VGSV_{DS} = V_{GS}, or at an output device's VDSV_{DS}.

This table has two known gaps.

The thorough version sweeps VGSV_{GS}, VDSV_{DS}, VSBV_{SB} and LL together, as Murmann's scripts do. Here the table was only the starting point for the first sizing. The final word always came from the full 45-corner simulation.


What I am taking from this

Some specs are contradictions, and a table can prove it. The 100 mV100\,mV margin did not fail in simulation — it failed on paper, in one line: 0.15 V−VDSAT0.15\,V - V_{DSAT} against a gm/IDg_m/I_D cap that keeps VDSATV_{DSAT} above 65 mV65\,mV. Characterising the devices first turned a long optimisation run into a two-minute conversation.

An optimiser does exactly what it is told. "Smallest area that meets the spec" produced 1 mV1\,mV margins, then a tail device short enough to cost UGF. Neither violated the objective. Both violated what I actually wanted, which was written down only after I saw the results: margins are part of the spec, and area ranks below current.

The most dangerous bug is the one that changes no picture. m=2 without mult=2 simulated a perfectly working circuit with the wrong statistics. It was found by reading the model source before trusting the Monte Carlo, not by looking at any output. The previous post ended on a schematic that was right while its netlist was wrong; this one has a netlist that was right while its mismatch was wrong.

"How did you choose that?" is worth asking. The first design passed, and its lengths were still guesses. Asking why is what led to the table, the worst-case run, and everything after.


References


Disclaimer: This blog post was created with assistance from Claude, an AI developed by Anthropic, under my direct supervision and guidance to ensure accuracy and alignment with my vision for the content.