Making a performance turn, without steering
LAION voice-acting model · 360 scored conditions · 30 August 2026 · second round
The question, unchanged. A person who is grieving and then finds something bitterly funny does not switch. The grief is still under the amusement at the end, and the amusement was already forming before the words arrived. What the product does today is a step: prompt sentence one for sadness, sentence two for something lighter, generate, join. The seam is audible because there is one.
What changed. The first round built the fade out of steering vectors — nudging the model's internal state towards one feeling and away from another. Those measure well and, a listener reported, sound wrong: more emotional, with strange artefacts, and the timbre off. The demo server has rolled back to adapters only. So this round throws steering out entirely and rebuilds the fade from the three levers that are still trusted: the adapters, the brief, and guidance.
1. Read this first: can the instrument see anything?
Every number in a study like this is produced by another model that listens to the audio and scores it. The way to find out whether that scorer can see the thing we are asking about is to give it the two ends of the journey and check that it can tell them apart: the same sentences, same voice, same random seed, performed entirely as feeling A and entirely as feeling B. Those are the references at the top of every group below.
Anchor separation, this run — the instrument works. At the strongest guidance tested (g = 3), with the emotion carried by the brief alone, the primary reading separates the two references by 0.67 standard deviations of their own within-condition scatter, and 17 of 60 texts clear a full unit. The first round measured 0.12 on 0 of 30 — and this run reproduces that number at the first round's own settings (nan). Nothing was wrong with the first round's arithmetic; it was reading three-second slices of an unguided contrast.
Three separate things produced the difference, and they can be told apart. The window: three seconds of a thirty-second zero-padded scorer input is 10 % signal. Guidance: the same reading goes nan at g = 0 to 0.67 at g = 3, rising at every step with no reversal anywhere in the grid. The adapters: switching them on lowers the same reading to 0.67 (17/60 texts) — see the second table.
Three different ways of reading the scorer are shown below, on three different lengths of audio.
own is the two emotion scores the pair is named after — what the first round
used. av is arousal and valence with the direction fixed in advance from the emotion
names. loo uses all 97 numbers the scorer produces, along a direction learned from the
other texts of the same pair, so it is not marking its own homework. w3 is
three-second windows, w6 six-second, sent the two sentences.
Brief only — the emotion carried by the words
| reading | segment | g=3 |
|---|---|---|
| own | w3 | – |
| own | w6 | – |
| own | sent | – |
| own | full | – |
| av | w3 | – |
| av | w6 | – |
| av | sent | – |
| av | full | – |
| loo | w3 | – |
| loo | w6 | – |
| loo | sent | – |
| loo | full | – |
Brief and that emotion's own adapter
| reading | segment | g=3 |
|---|---|---|
| own | w3 | 0.33 (4/60) |
| own | w6 | 0.43 (10/60) |
| own | sent | 0.42 (11/60) |
| own | full | 1.44 (29/60) |
| av | w3 | 0.49 (15/60) |
| av | w6 | 0.68 (20/60) |
| av | sent | 0.64 (19/60) |
| av | full | 1.22 (25/60) |
| loo | w3 | 0.51 (14/60) |
| loo | w6 | 0.70 (20/60) |
| loo | sent | 0.67 (17/60) |
| loo | full | 1.49 (31/60) |
Each cell is the mean separation in standard deviations, with the number of texts out of the total that reach a full standard deviation in brackets. Higher is better; below 1.0 the two ends of the journey are not reliably distinguishable and every ranking that follows is correspondingly weak. Bold marks a cell at or above 1.0.
The adapters make the two ends of the journey closer together.
Compare the two tables: at g = 3 and g = 4, all twelve readings are
smaller with the adapters on. (At g = 2 the four own readings still favour
the adapter, which is the last trace of the effect the first round saw; guidance erases it.)
It is not that adapter clips are noisier — the within-reference scatter is about the same
either way (2.48 against
2.48). It is that the two references are
1.0× further
apart without the adapter (1.61 against
1.61). This matters for everything below, because the
adapter-crossfade conditions are being asked to travel a distance the adapters themselves
shorten.
2. What the conditions are
Every condition performs the same words with the same voice and the same random seed. Only the method changes, so any difference you hear is the method.
- reference: the first feeling, all the way through — The same sentences performed with only the first emotion, from beginning to end. Both the written brief and that emotion's adapter. This is what 'fully A' sounds like, and the study's whole scale is built on it.
- reference: the second feeling, all the way through — The same again with only the second emotion.
- today's behaviour -- a step at the sentence break — Sentence one generated with the first emotion, sentence two with the second, and the two takes joined at the written pause. This is what the product does now and it is the thing to beat.
- crossfade, but the first feeling never fully leaves — As above, except the first emotion is held at a quarter weight for the rest of the clip. The grief stays under the amusement.
- the new feeling arrives fast, the old one leaves slowly — As above, but the two do not move at the same rate: the second emotion comes in quickly around the pause while the first fades out later and more gradually. This is the shape a real turn is supposed to have.
- CONTROL: the same fading, in a scrambled order — Exactly the same adapter weights, in blocks, but shuffled so there is no story from the first feeling to the second. If this sounds as though it turns, then what you are hearing in the other clips is the model being disturbed, not a transition.
Guidance is a separate dial that applies to all of them. The model is run twice per moment, once told what to do and once told nothing, and the difference between the two is exaggerated by a factor g. It makes the instruction bite harder. It also costs about twice the computing time and cannot start playing until the whole take is finished, which is why it matters exactly how much it buys. Field use puts the usable ceiling at about 4.0; above that it falls apart.
3. What it costs
| condition | g | n | turn (median) | turn (mean) | first round's reading | word error | genuineness | ms/frame |
|---|---|---|---|---|---|---|---|---|
| ANC_A | 3 | 60 | 0.10 | 0.08 | 0.92 | 0.040 | 1.27 | 98 |
| ANC_B | 3 | 60 | 0.20 | 0.09 | 0.64 | 0.053 | 1.22 | 98 |
| STEP | 3 | 60 | 1.38 | 1.25 | 1.25 | 0.027 | 1.47 | 75 |
| XF_FLOOR | 3 | 60 | 0.08 | -0.12 | 0.92 | 0.029 | 1.30 | 125 |
| XF_ASYM | 3 | 60 | 0.12 | 0.01 | 1.01 | 0.034 | 1.25 | 123 |
| C_SHUF | 3 | 60 | 0.24 | -0.08 | 0.91 | 0.037 | 1.29 | 122 |
"turn" is how far the performance travelled from the first feeling to the second, in standard deviations of the scorer's own wobble on a clip that is not changing. The floor — the end-to-start swing measured on a reference clip, which by construction does not turn at all — is 0.85. Anything below that did not measurably turn. ms/frame is the real generation cost per 80 ms of audio.
4. Listen
Pick a pair and a guidance level. The two references come first, then today's behaviour, then the methods, then the control. Rate anything you have an opinion about — 👍 if the turn sounds like one performance, 👎 if it sounds like two clips glued together or like nothing happens. Your ratings stay in this browser; the button at the bottom copies them all out as text.
5. What we would most like to know
- Can you hear the difference between the two references at all? If your ears say yes and the table says no, the instrument is the problem, not the model — and that changes what we do next more than any other answer.
- Does guidance help or hurt? Same condition, four guidance levels. It is the one lever here that is known to sound good, and 4.0 is meant to be the edge.
- Does the step actually sound broken? It is what ships today.
- Does the crossfade sound like one performance where the step sounds like two?
- Does the scrambled control sound like it turns? If it does, none of this means anything.