floorcall
A millisecond decision layer for voice agents: is the user done talking, was that “uh-huh” or a real interruption, what do they want, and do they need a human, answered together in one pass of a fine-tuned, calibrated Laya encoder. Model · code (commit 4088f09)
Backchannels while the agent explains
| # | when, and what the user said | wanted | naive agent | floorcall |
|---|---|---|---|---|
| 1 | 3.00 s · speech over the agent “uh-huh” | keep talking | ✗ stop and listen stopped talking at 3.00 s (7900 ms of its line unsaid) | ✓ keep talking kept talking backchannel (p=1.00); p(interruption) 0.00 < 0.09 |
| 2 | 4.10 s · pause after “uh-huh” | no expectation | · respond after timeout · no router answered 800 ms after the user stopped; the user spoke again 2400 ms later | — |
| 3 | 6.80 s · speech over the agent “right” | keep talking | ✗ no decision the agent had already stopped talking at 3.00 s | ✓ keep talking kept talking backchannel (p=1.00); p(interruption) 0.00 < 0.09 |
| 4 | 7.60 s · pause after “right” | no expectation | · respond after timeout · no router answered 800 ms after the user stopped; the user spoke again 3800 ms later | — |
| 5 | 14.10 s · pause after “okay, so how do I switch that on” | respond | ✓ respond after timeout · no router answered 800 ms after the user stopped (timeout) | ✓ respond after timeout · route out_of_scope answered 2000 ms after the user stopped (safety net) p(turn_complete)=0.61 < 0.785, silence 300 ms < 2000 ms; then silence 2000 ms >= max_wait 2000 ms |
A real interruption: a charge the user did not make
| # | when, and what the user said | wanted | naive agent | floorcall |
|---|---|---|---|---|
| 1 | 10.00 s · speech over the agent “wait, that one's not mine” | stop and listen | ✓ stop and listen stopped talking at 10.00 s (2900 ms of its line unsaid) | ✓ stop and listen stopped talking at 10.66 s (2245 ms of its line unsaid) p(interruption)=0.39 >= 0.09 |
| 2 | 13.50 s · pause after “I never bought anything there” | respond, route report_fraud | ✗ respond after timeout · no router answered 800 ms after the user stopped (timeout) | ✗ respond · route transactions answered 356 ms after the user stopped (decision) p(turn_complete)=0.83 >= 0.785 |
The same word twice: "yeah" and "yeah but"
| # | when, and what the user said | wanted | naive agent | floorcall |
|---|---|---|---|---|
| 1 | 4.50 s · speech over the agent “yeah” | keep talking | ✗ stop and listen stopped talking at 4.50 s (7100 ms of its line unsaid) | ✓ keep talking kept talking backchannel (p=1.00); p(interruption) 0.00 < 0.09 |
| 2 | 5.60 s · pause after “yeah” | no expectation | · respond after timeout · no router answered 800 ms after the user stopped; the user spoke again 3400 ms later | — |
| 3 | 9.60 s · speech over the agent “yeah but I set my limit to a thousand last week” | stop and listen | ✗ no decision the agent had already stopped talking at 4.50 s | ✓ stop and listen stopped talking at 9.66 s (1945 ms of its line unsaid) p(interruption)=0.87 >= 0.09 |
| 4 | 12.60 s · pause after “yeah but I set my limit to a thousand last week” | respond | ✓ respond after timeout · no router answered 800 ms after the user stopped (timeout) | ✓ respond after timeout · route spending_history answered 2000 ms after the user stopped (safety net) p(turn_complete)=0.69 < 0.785, silence 300 ms < 2000 ms; then silence 2000 ms >= max_wait 2000 ms |
Talk to someone else in the room
| # | when, and what the user said | wanted | naive agent | floorcall |
|---|---|---|---|---|
| 1 | 3.50 s · speech over the agent “honey can you grab the door” | ignore it | ✗ stop and listen stopped talking at 3.50 s (7400 ms of its line unsaid) | ✗ stop and listen stopped talking at 4.16 s (6745 ms of its line unsaid) p(interruption)=0.73 >= 0.09 |
| 2 | 5.60 s · pause after “honey can you grab the door” | no expectation | · respond after timeout · no router answered 800 ms after the user stopped; the user spoke again 5300 ms later | · respond · route bill_balance answered 356 ms after the user stopped; the user spoke again 5744 ms later p(turn_complete)=0.90 >= 0.785 |
| 3 | 13.20 s · pause after “yes, please pay it now” | respond, route pay_bill | ✗ respond after timeout · no router answered 800 ms after the user stopped (timeout) | ✓ respond · route pay_bill answered 356 ms after the user stopped (decision) p(turn_complete)=0.89 >= 0.785 |
The user pauses mid-request
| # | when, and what the user said | wanted | naive agent | floorcall |
|---|---|---|---|---|
| 1 | 1.80 s · pause after “I want to pay my” | keep listening | ✗ respond after timeout · no router answered 800 ms after the user stopped, mid-thought: it cut them off 300 ms before they went on | ✓ keep listening kept listening p(turn_complete)=0.00 < 0.785, silence 300 ms < 2000 ms |
| 2 | 4.10 s · pause after “credit card bill please” | respond, route pay_bill | ✗ respond after timeout · no router answered 800 ms after the user stopped (timeout) | ✓ respond · route pay_bill answered 356 ms after the user stopped (decision) p(turn_complete)=0.83 >= 0.785 |
A finished question
| # | when, and what the user said | wanted | naive agent | floorcall |
|---|---|---|---|---|
| 1 | 2.40 s · pause after “what's the balance on my checking account” | respond, route balance | ✗ respond after timeout · no router answered 800 ms after the user stopped (timeout) | ✓ respond · route balance answered 356 ms after the user stopped (decision) p(turn_complete)=0.80 >= 0.785 |
A request a bank agent does not handle
| # | when, and what the user said | wanted | naive agent | floorcall |
|---|---|---|---|---|
| 1 | 4.20 s · pause after “can you book me a table for two at an italian place tonight” | respond, route out_of_scope | ✗ respond after timeout · no router answered 800 ms after the user stopped (timeout) | ✓ respond · route out_of_scope answered 356 ms after the user stopped (decision) p(turn_complete)=0.87 >= 0.785 |
A frustrated user asks for a person, twice
| # | when, and what the user said | wanted | naive agent | floorcall |
|---|---|---|---|---|
| 1 | 3.00 s · pause after “this is ridiculous, I've explained this three times now,” | keep listening | ✗ respond after timeout · no router answered 800 ms after the user stopped, mid-thought: it cut them off 300 ms before they went on | ✗ respond · route spending_history answered 356 ms after the user stopped, mid-thought: it cut them off 744 ms before they went on p(turn_complete)=0.88 >= 0.785 |
| 2 | 5.90 s · pause after “just get me a real person” | respond, hand off | ✗ respond after timeout · no router answered 800 ms after the user stopped (timeout) | ✓ respond after timeout · route spending_history, hand off to a human answered 2000 ms after the user stopped (safety net) p(turn_complete)=0.61 < 0.785, silence 300 ms < 2000 ms; then silence 2000 ms >= max_wait 2000 ms; p(escalate)=0.59 >= 0.505: hand off |
| 3 | 11.20 s · speech over the agent “no, stop, put me through to a human now” | stop and listen, hand off | ✗ stop and listen stopped talking at 11.20 s (4600 ms of its line unsaid) | ✗ stop and listen stopped talking at 11.86 s (3945 ms of its line unsaid) p(interruption)=0.74 >= 0.09 |
| 4 | 14.20 s · pause after “no, stop, put me through to a human now” | respond, hand off | ✗ respond after timeout · no router answered 800 ms after the user stopped (timeout) | ✓ respond · route freeze_account, hand off to a human answered 356 ms after the user stopped (decision) p(turn_complete)=0.89 >= 0.785; p(escalate)=0.70 >= 0.505: hand off |
All numbers are on frozen test sets that no training, calibration or threshold choice ever saw, rendered from the committed results files by the same code as the repository's README. The fine-tuned rows are the released checkpoint, enz23/floorcall.
Table A: quality per decision
| Decision | Model | Accuracy [95% CI] | Macro-F1 | ECE | Brier | Hard-subset acc. (n) |
|---|---|---|---|---|---|---|
| D1 turn_complete | majority class (train prior) | 0.590 [0.582, 0.597] | 0.371 | 0.000 | 0.484 | 1.000 (775) |
| D1 turn_complete | TF-IDF + logistic regression | 0.624 [0.616, 0.632] | 0.544 [0.536, 0.553] | 0.008 [0.006, 0.018] | 0.456 [0.452, 0.460] | 0.843 [0.817, 0.868] (775) |
| D1 turn_complete | stock Laya, zero-shot | 0.488 [0.480, 0.496] | 0.485 | 0.028 | 0.510 | 0.505 (775) |
| D1 turn_complete | fine-tuned | 0.820 [0.813, 0.826] | 0.813 [0.807, 0.820] | 0.020 [0.015, 0.026] | 0.243 [0.237, 0.250] | 0.725 [0.693, 0.755] (775) |
| D1 turn_complete | fine-tuned + temperature | 0.820 [0.813, 0.826] | 0.813 [0.807, 0.820] | 0.010 [0.007, 0.016] | 0.243 [0.236, 0.249] | 0.725 [0.693, 0.755] (775) |
| D1 turn_complete | prompted LLM, stated probabilities | 0.721 [0.714, 0.728] | 0.721 | 0.138 | 0.435 | 0.412 (775) |
| D2 barge_in | majority class (train prior) | 0.501 [0.493, 0.510] | 0.223 | 0.010 | 0.564 | 0.453 (5012) |
| D2 barge_in | lexical rule: backchannel words + length | 0.951 [0.948, 0.955] | 0.908 [0.901, 0.916] | 0.001 [0.000, 0.005] | 0.092 [0.085, 0.098] | 0.985 [0.981, 0.988] (5012) |
| D2 barge_in | TF-IDF + logistic regression | 0.931 [0.926, 0.935] | 0.850 [0.840, 0.859] | 0.015 [0.012, 0.019] | 0.102 [0.097, 0.107] | 0.963 [0.958, 0.968] (5012) |
| D2 barge_in | stock Laya, zero-shot | 0.368 [0.360, 0.376] | 0.238 | 0.005 | 0.661 | 0.439 (5012) |
| D2 barge_in | fine-tuned | 0.969 [0.966, 0.972] | 0.937 [0.931, 0.943] | 0.008 [0.006, 0.010] | 0.042 [0.039, 0.046] | 0.977 [0.972, 0.981] (5012) |
| D2 barge_in | fine-tuned + temperature | 0.969 [0.966, 0.972] | 0.937 [0.931, 0.943] | 0.003 [0.003, 0.006] | 0.042 [0.038, 0.045] | 0.977 [0.972, 0.981] (5012) |
| D2 barge_in | prompted LLM, stated probabilities | 0.708 [0.700, 0.716] | 0.574 | 0.196 | 0.493 | 0.711 (5012) |
| D3 route | majority class (train prior) | 0.690 [0.666, 0.713] | 0.051 | 0.547 | 0.837 | n/a |
| D3 route | TF-IDF + logistic regression | 0.908 [0.893, 0.923] | 0.848 [0.820, 0.872] | 0.057 [0.046, 0.071] | 0.147 [0.130, 0.165] | n/a |
| D3 route | stock Laya, zero-shot | 0.934 [0.921, 0.946] | 0.866 | 0.030 | 0.115 | n/a |
| D3 route | fine-tuned | 0.948 [0.937, 0.959] | 0.906 [0.881, 0.926] | 0.043 [0.033, 0.054] | 0.090 [0.071, 0.111] | n/a |
| D3 route | fine-tuned + temperature | 0.948 [0.937, 0.959] | 0.906 [0.881, 0.926] | 0.011 [0.008, 0.023] | 0.082 [0.065, 0.100] | n/a |
| D3 route | prompted LLM, stated probabilities | 0.970 [0.961, 0.979] | 0.939 [0.918, 0.956] | 0.028 [0.021, 0.037] | 0.058 [0.044, 0.074] | n/a |
| D4 escalate | majority class (calib prior) | 0.555 [0.485, 0.625] | 0.357 [0.327, 0.385] | 0.025 [0.000, 0.095] | 0.495 [0.487, 0.504] | 0.710 [0.620, 0.800] (100) |
| D4 escalate | TF-IDF + logistic regression (θ = 0.240) | 0.685 [0.620, 0.750] | 0.681 [0.613, 0.743] | 0.167 [0.125, 0.239] | 0.473 [0.386, 0.561] | 0.740 [0.650, 0.820] (100) |
| D4 escalate | stock Laya, zero-shot | 0.665 [0.600, 0.730] | 0.625 [0.554, 0.693] | 0.092 [0.057, 0.159] | 0.422 [0.396, 0.447] | 0.730 [0.640, 0.810] (100) |
| D4 escalate | stock Laya, calib threshold (θ = 0.505) | 0.645 [0.580, 0.710] | 0.579 [0.506, 0.649] | 0.092 [0.057, 0.159] | 0.422 [0.396, 0.447] | 0.700 [0.610, 0.790] (100) |
| D4 escalate | fine-tuned | 0.695 [0.630, 0.755] | 0.672 [0.602, 0.737] | 0.277 [0.220, 0.340] | 0.544 [0.434, 0.658] | 0.740 [0.650, 0.820] (100) |
| D4 escalate | fine-tuned + temperature | 0.695 [0.630, 0.755] | 0.672 [0.602, 0.737] | 0.064 [0.040, 0.135] | 0.418 [0.363, 0.475] | 0.740 [0.650, 0.820] (100) |
| D4 escalate | fine-tuned + temperature, calib threshold (θ = 0.505) | 0.705 [0.640, 0.765] | 0.681 [0.612, 0.747] | 0.064 [0.040, 0.135] | 0.418 [0.363, 0.475] | 0.750 [0.660, 0.830] (100) |
| D4 escalate | prompted LLM, stated probabilities | 0.635 [0.570, 0.700] | 0.610 [0.538, 0.677] | 0.248 [0.186, 0.312] | 0.539 [0.445, 0.635] | 0.570 [0.470, 0.660] (100) |
Table B: latency, batch 1 (budgets: p99 ≤ 50 ms GPU, ≤ 100 ms CPU)
| Path | p50 ms | p95 ms | p99 ms | of which forward, p50 | of which packing, p50 | Fits budget (p99) |
|---|---|---|---|---|---|---|
| GPU, CUDA graphs: user_pause, 3 questions in 1 call | 29.1 | 55.6 | 59.0 | 24.6 | 0.8 | no (≤ 50 ms) |
| GPU, CUDA graphs: user_pause, 3 questions in 3 calls | 52.4 | 79.2 | 85.2 | 43.1 | 0.8 | no (≤ 50 ms) |
| GPU, CUDA graphs: user_speech_during_agent, 2 questions in 1 call | 44.3 | 54.9 | 57.4 | 38.3 | 1.9 | no (≤ 50 ms) |
| GPU, eager: user_pause, 3 questions in 1 call | 56.4 | 82.3 | 88.5 | 50.3 | 1.1 | no (≤ 50 ms) |
| GPU, eager: user_pause, 3 questions in 3 calls | 120.4 | 158.6 | 179.3 | 110.0 | 0.9 | no (≤ 50 ms) |
| GPU, eager: user_speech_during_agent, 2 questions in 1 call | 55.4 | 66.4 | 73.7 | 49.1 | 2.5 | no (≤ 50 ms) |
| CPU: user_pause, 3 questions in 1 call | 976.9 | 2162.4 | 2219.8 | 973.1 | 0.7 | no (≤ 100 ms) |
| CPU: user_pause, 3 questions in 3 calls | 680.9 | 1899.1 | 1953.6 | 674.0 | 0.6 | no (≤ 100 ms) |
| CPU: user_speech_during_agent, 2 questions in 1 call | 1153.6 | 1426.7 | 1464.5 | 1149.7 | 1.4 | no (≤ 100 ms) |
| prompted LLM (gpt-oss-20b via OpenRouter, pinned DeepInfra bf16), user_pause, 3 questions in 1 call: network latency, request to parsed answer | 2565.1 | 4270.7 | 6072.2 | n/a | n/a | no (≤ 50 ms) |
Measured on NVIDIA GeForce RTX 5070 Ti Laptop GPU (driver 591.86) and Intel64 Family 6 Model 197 Stepping 2, GenuineIntel with 16 torch threads; on AC power: True; torch 2.14.0+cu130, laya 0.3.21. GPU rows measured 2026-09-30 (UTC). Windows power mode: Best performance. GPU power limit: vendor default, 80 W base plus Dynamic Boost (115 W enforced at the start). Batch 1, 50 warmup and 1000 timed iterations over 50 fixed inputs per event; every timed call is a full Decider.decide (packing, tokenizing, forward, temperatures).
Table C: robustness to ASR-style noise (macro-F1, accuracy in brackets)
| Decision | noise 0.00 | noise 0.05 | noise 0.10 | noise 0.20 |
|---|---|---|---|---|
| D1 turn_complete | 0.814 (0.820) | 0.775 (0.786) | 0.743 (0.760) | 0.679 (0.713) |
| D2 barge_in | 0.937 (0.969) | 0.865 (0.927) | 0.810 (0.887) | 0.728 (0.814) |
| D3 route | 0.906 (0.948) | 0.882 (0.937) | 0.845 (0.919) | 0.759 (0.878) |
| D4 escalate | 0.672 (0.695) | 0.650 (0.675) | 0.643 (0.680) | 0.631 (0.665) |
Table D: ablations, trained on Kaggle against a full arm of the same recipe
| Variant | D1 macro-F1 (acc) | D1 hard acc. | D2 macro-F1 (acc) | D2 hard acc. | D3 macro-F1 (acc) | D4 macro-F1 (acc) |
|---|---|---|---|---|---|---|
| full model | 0.824 (0.827) | 0.675 | 0.941 (0.973) | 0.994 | 0.911 (0.955) | 0.695 (0.715) |
| without recent_turns | 0.823 (0.826) | 0.717 | 0.941 (0.971) | 0.982 | 0.924 (0.958) | 0.629 (0.680) |
| without agent_last_utterance | 0.820 (0.823) | 0.720 | 0.942 (0.971) | 0.987 | 0.929 (0.964) | 0.648 (0.685) |
| without normalization, scored on written text | 0.959 (0.960) | 0.992 | 0.945 (0.972) | 0.981 | 0.925 (0.957) | 0.695 (0.715) |
| without normalization, scored on ASR-style text | 0.381 (0.593) | 1.000 | 0.938 (0.969) | 0.979 | 0.927 (0.959) | 0.701 (0.720) |
Each variant minus the full arm, paired-bootstrap 95% intervals:
| Variant minus full arm | D1 macro-F1 | D1 hard acc. | D2 macro-F1 | D2 hard acc. | D3 macro-F1 | D4 macro-F1 |
|---|---|---|---|---|---|---|
| without recent_turns | -0.001 [-0.006, +0.003] | +0.043 [+0.021, +0.066] | -0.000 [-0.007, +0.005] | -0.012 [-0.015, -0.009] | +0.013 [-0.005, +0.032] | -0.066 [-0.130, -0.004] |
| without agent_last_utterance | -0.004 [-0.009, -0.000] | +0.045 [+0.023, +0.067] | +0.000 [-0.005, +0.006] | -0.008 [-0.011, -0.005] | +0.018 [+0.000, +0.037] | -0.047 [-0.101, +0.004] |
| without normalization, scored on written text | +0.134 [+0.128, +0.141] | +0.317 [+0.285, +0.350] | +0.004 [-0.002, +0.010] | -0.014 [-0.017, -0.010] | +0.015 [-0.004, +0.034] | +0.000 [-0.045, +0.046] |
| without normalization, scored on ASR-style text | -0.443 [-0.451, -0.436] | +0.325 [+0.293, +0.359] | -0.003 [-0.009, +0.002] | -0.015 [-0.019, -0.012] | +0.017 [-0.001, +0.036] | +0.006 [-0.039, +0.052] |
Operating points, chosen on calib
| Model | θ_interrupt | false stops (test) | missed interruptions (test) | θ_yield | premature responses (test) | added delay, ms (test) |
|---|---|---|---|---|---|---|
| stock Laya | 0.340 | 0.040 | 0.931 | 0.630 | 0.055 | 1599 |
| fine-tuned + temperature | 0.090 | 0.050 | 0.003 | 0.785 | 0.043 | 903 |



Live decisions need the model on a server, which a free Hugging Face Space no longer provides. Run the same thing locally (no API key; a CPU is enough):
git clone https://github.com/yashraz23/floorcall cd floorcall uv sync --no-default-groups --group dev --group cpu uv run floorcall replay demo/scripts/*.json --compare naive
or load the checkpoint in Python as the model card shows, with Decider.decide for calibrated probabilities on any state you write.