Video evidence review for criminal defense: every camera, one timeline.
Thirteen hours of overlapping bodycam, dashcam and bystander footage is one incident. VST aligns those recordings from the cameras' own audio, so you watch every angle at once — and read every word from every camera in one searchable transcript, on a workstation you control.
Watch it work — take the tour → Ask about a small-office licence
Local processing · Annual small-office licence · Remote setup available
A routine case can arrive with more video than anyone has time to watch
A single incident can arrive as a dozen overlapping recordings — every bodycam, the dashcam, a bystander’s phone — and today that means watching each file alone and rebuilding the sequence in your head. Reviewing it all in real time consumes the hours a contracted case can’t spare, and transcription and translation services are expensive on top.
VST produces a searchable first-pass record on your own workstation, helping you find the parts that need careful human review. It works across the media sources a defense practice actually receives:
- Body-worn camera
- Dashcam
- Police interviews and interrogations
- Jail calls
- 911 and dispatch audio
- CCTV and phone video
- Multilingual recordings
One incident, every angle, one clock
The fastest way to understand what happened is to stop watching the files one at a time. Overlapping recordings of the same incident — body-worn cameras, dashcam, CCTV, a bystander’s phone — are aligned automatically from their shared audio and played side by side on a single incident timeline. Scrub once and every camera moves together; jump into the transcript from whichever camera heard it best.
Systems that only synchronise their own cameras can’t do this for the bystander’s phone or the store’s CCTV. VST syncs the footage you actually get in discovery, wherever it came from — frame-accurate where cameras share sound. Where a recording shares no usable audio, VST says so rather than guessing, and you place it by hand, clearly marked as reviewer-set.
A dozen recordings stop being a dozen files to watch and become one incident to review.
Watch it work — take the two-minute tour → · Alignment is reviewer-derived from audio, for review — not certified wall-clock time.
Then read and search every word
With the incident on one clock, the rest of the review is fast.
Step 1 — Add the case media. Import a folder containing video and audio. VST preserves the relationship between each output and its source file.
Step 2 — Transcribe and identify speakers. Convert difficult audio into a timestamped, searchable transcript. Speakers are separated and can be relabelled during review.
Step 3 — Translate where required. Detect the spoken language and produce a working translation alongside the original transcript.
Step 4 — Summarise with evidence attached. Generate a readable account of a recording, with supporting extracts that let you check what the model relied upon.
Step 5 — Search and return to the source. Search for a name, place, phrase or event and jump directly to the corresponding point in the recording.
Step 6 — Export the working record. Export transcripts, notes and structured reports for case preparation and downstream systems.
An AI answer is useful only when you can check it
General-purpose AI can produce a confident summary that is not supported by the recording. VST pairs each generated account with extracts from the underlying transcript and keeps the original recording one click away — see source-linked text and audio summarisation.
The attorney decides what matters. VST makes the evidence faster to navigate — it helps you understand the record; it does not decide what the record means legally.
The evidence does not have to leave your office
VST can run on a suitable NVIDIA workstation in the firm. Video, transcripts, translations and summaries remain on infrastructure controlled by the customer; no shared public AI service is required, and there is no mandatory upload to Rigr AI for a local deployment.
Already have a GPU workstation? Send us the model and we will confirm whether it is suitable. Need a machine? We can provide a tested specification and assist with setup. See Deployment, sovereignty & information control for how VST runs inside your environment.
A cost model built for a small office
One annual small-office licence. No per-minute transcription charge — so you can process recurring case material on the licensed workstation without watching a meter. Hardware is separate; we will help you confirm or specify it.
Who it is for
A good fit
- Solo and small criminal-defense practices
- Private assigned or appointed counsel
- Public-defender and conflict-counsel teams
- Offices receiving recurring bodycam or jail-call evidence
- Practices encountering multilingual discovery
- Teams that prefer local processing and predictable cost
Probably not a fit
- A practice receiving only a few minutes of audiovisual evidence per year
- Someone seeking a certified verbatim transcript for filing without human review
- A customer without suitable hardware who is unwilling to acquire it
- A consumer trying to transcribe a single personal recording
VST is the same evidence-review workspace Rigr AI builds for investigative teams — see VST Teams for the full platform, and Deployment for running it on one workstation in the office. A criminal-defense practice is a valid customer, not an accidental visitor.
Frequently asked questions
Can VST transcribe police body-camera footage?
Yes. VST processes audio and video into timestamped, searchable transcripts and keeps each transcript linked to the source recording.
Can it review several bodycam files from one incident?
Yes — this is one of the things VST does best. Load every recording of the incident and VST aligns them on a single timeline from their shared audio, so you watch all the cameras simultaneously and read a merged, per-camera transcript. It works for footage from any source — agency bodycams, dashcam, CCTV, bystander phones — not just cameras from a single vendor. Recordings without usable shared audio can be placed on the timeline by hand, marked as reviewer-set. You can try it in our interactive tour at rigr.ai/tours/bodycam-review/. Alignment is reviewer-derived from audio, for review — not certified wall-clock time.
Can it translate Spanish bodycam footage?
VST detects and transcribes supported source languages and can render a working translation into the reviewer's language. Human or certified translation may still be required for formal evidential use.
Does the footage have to be uploaded to Rigr AI?
No mandatory cloud upload is required for a local deployment. Processing can take place on a workstation controlled by the customer.
Can a solo attorney use it?
Yes. A small-office deployment can run on one suitable processing workstation and does not require an enterprise server.
Is the transcript certified?
No. VST produces an AI-assisted working transcript for evidence review and case preparation. It does not replace a certified transcript where one is required.
Does it charge by the minute?
The small-office licence is annual rather than per-minute, so you can process recurring case material on the licensed workstation.
What hardware is required?
Requirements depend on volume and model configuration. Send us your workstation model for a quick compatibility check, or ask us for a tested specification.