Full audio evidence — Contractor+ / Jobber
2026-09-10. DEGRADED: 2-of-3 audio evidence layers complete — full fresh speech recognition and numerical acoustic analysis completed; direct listening remains unavailable. This report does not certify vocal performance, music, sound effects, audible splice quality, or a finished audiovisual review.
The material is one nearly 30-second Jobber reference, one 28.5-second failed Contractor+ final, a 96.5-second real customer testimonial, and two earlier Contractor+ exports. The task is to understand their complete sound and spoken claims before another production decision. Audio has no fold geometry; source video geometries are retained in the probes. No media was generated or source changed.
Name the discipline that owns this failure class and its standard solution: audio post-production owns sound continuity, intelligibility, voice performance and mix review; its standard solution is continuous listening to the delivered mix, supported by a timed transcript, source stems and loudness measurements. Documentary/advertising editorial owns quote and claim provenance; its solution is a source-linked quote ledger and an explicit distinction between customer speech and authored narration. The best answer may be: do not build this — say so. Do not build another waveform classifier to make a listening claim. Use the recovered media and a capable listening surface to close the named perception gap.
What actually ran
All five complete tracks were decoded locally and freshly transcribed by the already cached Systran/faster-whisper-small, using faster-whisper 1.2.1, CPU, requested int8 and observed int8_float32, English, beam size 5, word timestamps, and no VAD trimming. Network/downloads were disabled. Full transcript text, every segment, every estimated word time and input hashes are saved in whole-understanding/audio. Whisper has no reasoning-effort control; beam search is not reasoning effort. This reviewing agent was requested as Astra/ultra; exact runtime model/effort are unobserved, not asserted.
I attempted to forward the complete Jobber MP3 through the native functions.audio helper. The actual result was “audio content omitted because you do not support audio input.” Therefore this agent did not hear it. The exact attempt and limits are in execution-receipt.json.
A fresh metadata-only probe of the six existing tailnet Ollama endpoints found no advertised audio-understanding model on the five reachable nodes; delta-amd timed out. Installed local listen.py/Whisper.cpp is another speech-to-text route, not a soundtrack listener. The current callable tool catalog exposed generation/voice-change tools but no audio-input understanding tool. PromptLab's spectrogram-to-image-model method was read and deliberately not used as listening evidence. This is a bounded availability finding, not a claim that no possible audio model exists. See local-audio-model-probe.json.
No external service upload, model download, inference to a new destination, generation job, email, service change, shared registry edit or source modification occurred. The native payload attempt used the current conversation's provided media helper and was rejected before audio became model input. A fresh SHA-256 comparison confirms that all three original videos, both alternate video members and both supplied transcript members still match the prior intake hashes: source-preservation-check.json, seven of seven matches.
Durations and full coverage
| Track | Source container | Decoded PCM | Recognized speech span | Native sound |
|---|---|---|---|---|
| Jobber | 29.985 s | 29.989819 s | 0–29.38 s | 44.1 kHz stereo |
| Failed F1 final | 28.468 s | 28.480 s | 0–27.98 s | 48 kHz mono |
| Real Rands testimonial | 96.502132 s | 96.502132 s | 1.50–85.58 s | 44.1 kHz stereo |
| Earlier F1 final 9×16 | 28.958333 s | 28.906667 s | 0–28.74 s | 96 kHz mono |
| Earlier F1 web | 28.958333 s | 28.906667 s | 0–28.74 s | 96 kHz mono |
All source streams report start time zero. Container, encoded stream and decoded-sample durations are separate facts: decoder output exceeds the final F1 container by 12 ms and ends about 51.7 ms before the alternate containers. No silence was invented to force matching lengths. Do not use an ASR endpoint as the end of the video. The full Rands file has approximately eleven seconds after its last recognized words, with substantial measured sound continuing into that tail.
The two earlier exports contain identical decoded PCM payloads, hash 5165fbf72f1206d471596e7a32a9f72d139deb8c99683e22c6b527bd82d95916. They are one audio version in two video exports. Both were nevertheless fully transcribed. Source identity, transform details and all fifteen continuous 15-second-or-shorter native-PCM listening cuts are recorded in audio-manifest.json. Rejoining each track's cuts reproduces its entire decoded PCM byte for byte; verification.json records that check and the timing bounds.
What the speech says, with original timestamps
These are fresh ASR observations, not claims of having heard the words. Times are estimates, not phoneme- or frame-accurate alignment.
| Source time | Jobber narrative beat |
|---|---|
| 0–3.46 | Introduces Danielle and 13 years in tree care. |
| 4.08–7.08 | First year with Jobber; her business crossed $1 million. |
| 7.88–12.54 | Customer calls revealed things had gone wrong; she heard one side. |
| 13.60–18.04 | Crew tells her first; photos, notes and updates live in Jobber. |
| 18.50–22.54 | She has proof; her team connects when she is absent. |
| 22.96–25.42 | Staying connected and stepping away are both possible. |
| 25.86–29.38 | Brand, “built for blue collar,” free-trial CTA and website. |
The spoken architecture is personal introduction → concrete result → old information problem → mechanism → evidence → freedom → brand/action. It includes pauses at consequential turns: 3.46–4.08 after experience, 7.08–7.88 after the million-dollar result, and 12.54–13.60 before “Now.” Measurable sound remains through those transcript gaps. This establishes the timing of the read and mixed audio, not what instruments or effects fill the spaces. Full source: jobber-asr.txt.
| Source time | Failed F1 final narrative beat |
|---|---|
| 0–0.48 | “This is Scott.” |
| 1.04–4.66 | Scott owns Rands Mechanical; his crew has over 30 years of combined experience. Raw ASR spells the brand “Rans.” |
| 5.26–9.34 | Three years with Contractor Plus; Rands is up over 350%. |
| Approximately 10.0–13.10 | Bad call reaches Scott through the homeowner rather than his tech. |
| 13.74–18.92 | The job tells him first; readings, photos and notes are logged on site before the van leaves. |
| 19.56–21.86 | He no longer defends the work; he shows it. |
| 22.38–27.98 | Brand; built by a contractor rather than a software company; start-free CTA and website. |
The ASR assigns “There” 9.34–10.08 even though the waveform is below −55 dBFS at 9.49–10.01. That is a concrete alignment warning: the recognizer can stretch a word across a pause. Use the JSON to locate content, then perform actual audible/phonetic alignment before edit decisions. The growth-number token is estimated at 8.26–9.34 in the final, compared with 6.74–8.06 in the earlier exports. Historical “350% lands at 8.06” timing belongs to that older audio, not automatically this final. Full source: failed-f1-asr.txt.
| Source time | Real Rands testimonial content |
|---|---|
| 1.50–15.40 | Scott introduces himself as owner, describes HVAC/electrical work, residential/commercial clients, and combined experience over 30 years. |
| 15.80–22.14 | Professional work should be matched by professional estimates/invoices. |
| 22.46–28.80 | Previous invoicing app worked, but did not do even a third of Contractor Plus's functions. |
| 28.80–40.74 | Estimate types, converting estimate to invoice, email/text delivery and built-in contracts. Exact connector wording near 33 s and 40 s remains transcription-uncertain. |
| 41.04–47.16 | Schedule everyone's work and communicate about each job inside the app. |
| 47.44–64.42 | Price/value; Contractor Plus works; estimate/invoice-open, payment and upcoming-job notifications; client/job information in one place. |
| 64.58–74.94 | Broad usefulness and Android/iPhone/tablet/office-computer access. |
| 75.44–78.66 | “saves us time, saves us money, and it helps us look good on paper.” |
| 78.96–85.58 | Optimism and recommendation; ends with “I'd have to agree with that.” ASR alone does not establish who delivers this closing response. |
| 85.58–96.50 | No further words recognized; sound remains through most of this interval and must be heard before labeling it music, ambience, effects or silence. |
The genuine customer story supplies professionalism, workflow, value, time and money saved, with team communication as one feature. It does not supply the failed ad's particular homeowner/technician dispute and proof narrative. Full source: rands-testimonial-asr.txt.
Exact transcript and version reconciliation
The supplied and fresh Jobber texts are lexically identical after folding punctuation/capitalization, except four instances of old “Jabber” versus fresh “Jobber.” That is a transcription spelling discrepancy, not an ad revision. Old recognition ended at 29.34; fresh recognition ends at 29.38. Neither is the container endpoint.
The supplied and fresh Rands texts differ in exactly four normalized token edits: two “RANS”→“Rands” spellings; “in the touch of a button”→“and a touch of a button”; and insertion of “your” before “DocuSign.” Retain the last two as uncertain recognition until heard. Do not silently polish them into a purported verbatim quote. Neither transcript contains a 19-year individual-career claim. Both contain combined experience over 30 years. They otherwise agree lexically, including the price/value, capabilities and closing recommendation. The supplied last segment ends 85.54; fresh ends 85.58. Exact differences and normalization rule: supplied-vs-fresh-transcript-diff.json.
| Item | Earlier 9×16 + web audio | Failed final audio | Real testimonial support |
|---|---|---|---|
| Experience | “Scott's been in HVAC 19 years,” 1.18–3.12 | “over 30 years of combined experience on his crew,” 2.68–4.66 | Combined over 30 years at 13.38–15.40. Does not prove Scott individually worked 19 years. |
| Field logging | “logged in the driveway” | “logged on site” | No matching specific event or location stated. |
| Tenure with Contractor+ | Three years | Three years | Not stated in this testimonial. |
| Growth | Up over 350% | Up over 350% | Not stated in this testimonial. The “third” quote is feature comparison, not growth. |
| Scott/homeowner/tech conflict | Authored narrated incident | Authored narrated incident | No matching incident stated. |
| Brand positioning/CTA | Contractor-built/software-company comparison and start-free URL | Same | Not a quotation from this customer's testimonial. |
“Not in this testimonial” does not establish that a claim is false or absent from every other source. Root owns the broader claim ledger. The experience correction is a real version change; misspellings and the touch/DocuSign variants are recognizer disagreements. Earlier full text: alternate-f1-final-9x16-asr.txt.
Objective soundtrack, continuity and pace evidence
Acoustic analysis used native channels with signed 16-bit PCM, non-overlapping 10 ms RMS windows, and maximum-channel level to avoid classifying a stereo cancellation as silence. A “quiet run” means below the stated digital threshold continuously for at least 120 ms. It is not perceptual silence. LUFS measurements use FFmpeg EBU R128. Full threshold intervals, per-second levels and raw loudness output are retained next to each track.
| Track | Integrated LUFS | Quiet time below −55 dBFS | What that establishes |
|---|---|---|---|
| Jobber | −13.7 | 0.18 s | Only the opening has a qualifying low-level run; the mix retains measurable energy through the read's pauses. |
| Failed final | −17.4 | 5.75 s | Repeated low-level breaks between words/phrases. The mix lacks continuous above-threshold energy across those breaks. |
| Earlier exports, each | −16.5 | 1.54 s | Different low-level continuity from the failed final; identical sound to each other. |
| Real testimonial | −20.8 | 0.60 s | Qualifying low-level runs occur at the head and tail; substantial audio remains after the last recognized words. |
Examples in the final are 0.67–1.20, 4.86–5.28, 9.49–10.01, 13.29–13.75, 14.95–15.46, 19.03–19.58 and 23.20–23.87. The Jobber reference does not have corresponding ≥120 ms drops below −55 dBFS between its recognized phrases. This is stronger evidence of a different continuity pattern than the old unsupported label “dry TTS over silence,” while still leaving sound identity unlistened. Root can compare corresponding visuals and cuts against these source times.
Rands is below −55 dBFS for 0–0.32 and 96.22–96.50. At 90–91 seconds its max-channel RMS is −22.62 dBFS, with a −5.72 dBFS sample peak, despite no recognized speech there. The tail must not be discarded as dead air on the strength of the transcript.
These measurements do not establish music genre, instruments, tempo, soundtrack source, emotional delivery, naturalness, voice identity, ducking, speech overlap, breath removal or audible splice/click defects. A final mono mix without isolated stems cannot reveal an exact ducking envelope or prove an edit algorithm. Sample/true peaks are preserved in the receipts; these encoded/captured sources are not a basis for imposing a mastering target or ranking mix quality.
ASR token counts over the first-to-last recognized-word span are Jobber 86/29.38 s = 175.6 tokens/minute; final F1 84/27.98 s = 180.1; older F1 78/28.74 s = 162.8; Rands 319/84.08 s = 227.6. Numerals, currencies, acronyms and URLs need spoken expansion, and ASR can smear pauses. These are not measured articulation rates. Final F1 and Jobber being close does not justify saying the final sounds rushed or natural.
The historical account's 1.34× compression is reported history. This pass did not recover and compare an uncompressed source performance that would verify that factor, nor prove time stretching from waveform appearance. The final's corrected opening carries additional information and moves the growth token later; that directly measured change should inform edit alignment without inventing a performance verdict.
Ready listening handoff and remaining work
Full native WAVs and fifteen timestamped cuts are local and verified. The Jobber pair covers 0–15 and 15–29.989819; failed final 0–15 and 15–28.48; Rands seven cuts cover 0–96.502132; each older version has two cuts covering its full decoded 28.906667. Filenames display rounded ends; the manifest preserves exact sample-derived ends. No normalization, fade, soundtrack replacement or speed change was applied to these cuts.
On an audio-capable surface, audition the complete source first, then use these cuts to mark music/ambience/effects, voice character and performance, audible transitions, and the final Rands tail. The current helper invocation is executable but the present session rejected its audio input; retrying the same payload here does not close the gap. Root can use the transcripts and numerical receipts now for full narrative/claim alignment. A trustworthy soundtrack verdict and audible acceptance remain the explicit successor. No new tool installation or autonomous production is required by this report.