All Posts

Sample Rates in Film Production: 48kHz vs. 96kHz and When Higher Actually Matters

Post-Production14 min read
Close-up of audio waveforms on an oscilloscope representing digital sample rate and frequency analysis

The Score That Cost Double the Storage for No Audible Reason

A composer delivers 96kHz / 24-bit stems for a 90-minute documentary score. The re-recording mixer opens the files in a Pro Tools session running at 48kHz - the standard for the film's dialogue and sound effects. Pro Tools converts the 96kHz stems on import. The conversion is inaudible. The mix finishes at 48kHz / 24-bit. The broadcast deliverable is 48kHz / 24-bit. The streaming deliverable is 48kHz / 24-bit.

The composer's 96kHz session doubled every file's size, doubled the CPU load per plug-in, and produced no audible improvement in the final deliverable. It cost approximately $400 in extra storage and added processing overhead throughout the 3-week scoring process.

This isn't an argument against 96kHz in all situations. There are specific production contexts where recording above 48kHz has genuine technical value. This post explains what sample rate actually controls, when going beyond 48kHz is justified, and when it adds cost without adding quality to the deliverable your audience hears.

Standards referenced here draw from the Nyquist-Shannon sampling theorem, the Audio Engineering Society (AES) professional audio standards, the EBU Tech 3250 broadcast alignment recommendation, and SMPTE ST 302 digital audio for broadcast - all of which specify 48kHz as the standard for film and video audio delivery.

What Sample Rate Actually Controls

Sample rate is the number of times per second a digital audio system measures the amplitude of an analog signal. At 48kHz, the system takes 48,000 measurements per second. At 96kHz, it takes 96,000.

The Nyquist-Shannon sampling theorem, foundational to all digital audio, states that a digital system accurately reproduces any frequency up to exactly half the sample rate. At 48kHz, the system can represent frequencies up to 24kHz. At 96kHz, the ceiling is 48kHz. At 192kHz, 96kHz.

Human hearing extends to approximately 20kHz, and most adults over 30 hear to around 16 to 17kHz. A 48kHz system reproduces all audible frequencies with full accuracy. The 4kHz of headroom above 20kHz (up to 24kHz at 48kHz) provides margin for the anti-aliasing filter that prevents high-frequency aliasing artifacts from entering the audible range.

Why does 96kHz exist for professional audio? Three reasons hold up under scrutiny:

First, a 48kHz anti-aliasing filter must roll off sharply between 20kHz and 24kHz. Steep filter slopes can cause subtle phase shifts in the 16 to 20kHz range. A 96kHz system rolls off above 40kHz, well outside human hearing, using a gentler slope that avoids those phase shifts entirely.

Second, certain processing operations - pitch shifting, harmonic saturation plug-ins, physical modeling synthesis - generate intermodulation products above the fundamental frequency. At 48kHz, products above 24kHz alias back into the audible range as distortion. At 96kHz, they stay above the Nyquist ceiling and get filtered out cleanly.

Third, some acoustic sources produce content above 20kHz - cymbals, tape saturation, certain acoustic instruments. Whether capturing this content improves audible quality is debated. Controlled blind listening studies, including those published in the Journal of the Audio Engineering Society, have found inconsistent evidence of audible benefit from ultrasonic content in normal listening environments.

Sample Rate Reference Table for Film Audio

The file size multiplier is exact: a 96kHz file is exactly twice the size of a 48kHz file at the same bit depth and duration. CPU multipliers are approximate and vary by algorithm and plug-in design.

Sample RateFrequency CeilingFile Size vs. 48kHzCPU Load vs. 48kHzRequired ByWhen It Matters
44.1kHz22.05kHz92%92%CD, some music streamingLegacy music delivery; never use for video
48kHz24kHzBaselineBaselineAll broadcast, streaming, DCPStandard for all film and video audio
96kHz48kHz200%150-200%Optional high-res deliveryOrchestral score recording, analog transfers
192kHz96kHz400%300-400%Rarely requiredScientific archival, ultrasonic documentation

The 48kHz standard for professional video audio was established in the 1980s when the EBU and SMPTE standardized digital audio for broadcast television. The EBU Tech 3250 recommendation and SMPTE ST 302 both specify 48kHz as the mandatory sample rate for broadcast. This standard applies unchanged to theatrical DCP delivery, broadcast, and all major streaming platforms.

Three Real-World Sample Rate Decisions

Example 1: Location Sound Recording, 15-Day Narrative Feature

A production sound mixer on a 15-day narrative feature uses a Sound Devices 888. The director asks whether 96kHz would "future-proof" the location audio recordings.

The deliverables spec: 48kHz / 24-bit broadcast WAV for broadcast, 48kHz mix for the re-recording mixer, 48kHz DCP audio track for theatrical. No deliverable requires or benefits from 96kHz location audio.

Storage calculation: approximately 600 GB of 48kHz location audio over the shoot doubles to 1.2 TB at 96kHz. Added hard drive cost: $300 to $400. Added daily DIT offload time: approximately 30 minutes.

Decision: 48kHz / 24-bit. The mixer confirmed that 24-bit depth provides substantially more practical benefit than 96kHz sample rate. The 24-bit noise floor (theoretical -144 dBFS) versus 16-bit (-96 dBFS) prevents clipping on unexpected loud events without requiring conservative gain staging. Higher sample rate adds storage cost. Higher bit depth adds dynamic range headroom at no additional storage cost beyond the standard 50% increase from 16-bit to 24-bit.

Example 2: Original Score Recording, 40-Piece Chamber Orchestra

A composer records a 40-piece chamber orchestra for a mid-budget feature. Studio A at a major facility, NEVE 8078 console, three-day session at $18,000 total. The scoring budget supports either sample rate.

Live acoustic instruments - strings, brass, woodwinds - contain transient content between 20kHz and 40kHz. Some orchestral recording engineers report that 96kHz captures a slightly more open high-frequency quality, attributed to the gentler anti-aliasing filter. This is a plausible benefit in a high-quality monitoring environment, not a guaranteed audible improvement.

Storage calculation: 40 channels at 96kHz for 3 days generates approximately 1.8 TB versus 900 GB at 48kHz. Additional storage cost: approximately $200.

Decision: 96kHz / 24-bit for the orchestral recording. The 96kHz sessions go to the composer for score mixing, who keeps the session at 96kHz throughout the score's own production. The re-recording mixer receives 96kHz stems and converts to 48kHz on import using a high-quality SRC algorithm. The 96kHz sessions preserve maximum quality through multiple processing generations in the score's own workflow. The final film delivery is 48kHz regardless. Net cost of the 96kHz choice: approximately $400 in storage.

Example 3: Natural History Documentary Field Recording

A documentary sound recordist captures wildlife audio - birdsong, insect ambience - for a natural history series. Field recorder: Zoom F8n Pro.

Birds and insects produce significant ultrasonic content above 20kHz. Some bat species communicate in the 20 to 80kHz range, fully outside a 48kHz system's capture range. For standard broadcast documentary delivery, this content has no value - it's inaudible and gets discarded in post. However, if the series delivers ultrasonic recordings to a scientific broadcaster for research purposes, 96kHz has genuine, non-placebo value.

Decision: 48kHz for standard sequences; 96kHz for sequences the production specifically flags for scientific documentation. A dual-track workflow - where the F8n records simultaneously at both sample rates on separate file sets - eliminates the need to decide per setup.

When to Choose Each Sample Rate: A Decision Framework

Step 1: Identify all delivery targets. Broadcast and streaming delivery: 48kHz. DCP theatrical: 48kHz. Dolby Atmos theatrical master sessions: 48kHz or 96kHz depending on the facility. Netflix's hi-res Atmos tier: 96kHz stems accepted. If any deliverable is 48kHz, the final master will be 48kHz regardless of acquisition rate.

Step 2: Count processing generations. If the audio passes through more than 3 to 4 processing stages before final delivery - record, mix, stem mix, final mix, loudness correction - recording at 96kHz means each stage operates at higher resolution before the final conversion. The benefit is subtle: audible primarily in the 14 to 20kHz range on high-resolution monitoring.

Step 3: Check your plug-ins. Harmonic saturation processors, pitch shifters, and physical modeling synthesizers produce their cleanest output at 96kHz. Most EQs, compressors, and reverbs produce effectively identical results at both sample rates. If the session relies heavily on saturation or pitch processing, 96kHz is defensible. If it doesn't, 96kHz is overhead.

Step 4: Calculate storage. Use the Audio Bitrate Storage Calculator to enter your project's channel count, bit depth, duration, and both sample rate options. The calculator returns exact file sizes for both scenarios and the storage difference. Confirm your storage budget and drive write speeds support the higher rate before committing.

Step 5: Test CPU load. Open your standard session template, switch the sample rate to 96kHz, and run a performance test with your planned plug-in load. If buffer underruns appear at your normal buffer size (typically 256 or 512 samples), you need either a larger buffer - which adds monitoring latency for performers - or hardware upgrades. A larger buffer is acceptable for non-live mixing but not for recording sessions where performers monitor through the system.

Step 6: Match rates across all recording devices. Every recorder and camera capturing audio for the same scene must run at the same sample rate. A 96kHz field recorder delivering audio to a 48kHz camera creates a sample rate mismatch that the NLE handles by converting in real time - adding CPU load and potential quality penalties depending on the conversion algorithm quality.

Pro Tips and Common Mistakes

Pro Tip: When delivering 96kHz stems to a 48kHz final mix session, perform the sample rate conversion once at the start of the mix using a dedicated high-quality SRC tool - iZotope's 64-bit SRC, r8brain, or Reaper's resampler at "Extreme" quality - rather than letting Pro Tools or Logic convert on import. One deliberate, high-quality conversion produces cleaner results than repeated on-the-fly conversions throughout the session.

Pro Tip: 44.1kHz audio should never appear in a video production session. When 44.1kHz audio from a licensed music service or consumer recorder mixes with 48kHz video in an NLE, the NLE performs a sample rate conversion. At low-quality settings, this conversion introduces artifacts. Some NLEs handle the conversion inconsistently, creating drift over long clips. Convert all licensed music to 48kHz before importing it into any video project.

Pro Tip: For narration or podcast audio destined for YouTube or streaming, 48kHz / 16-bit is technically sufficient and dramatically reduces file sizes versus 24-bit. The 16-bit noise floor at -96 dBFS exceeds the practical dynamic range of any recording made in a real room - room noise is always above -96 dBFS. Reserve 24-bit for production recordings where unexpected loud transients could drive the signal toward clipping.

Common Mistake: Delivering a 96kHz film score to a 48kHz mix session without planning the conversion. Pro Tools, Logic, and Nuendo all sample-rate convert on import automatically. The conversion is functionally transparent with a quality SRC algorithm. The composer gains nothing audible in the final delivery from the 96kHz sessions and pays in storage and CPU overhead throughout the entire scoring process.

The fix: Decide the delivery sample rate at the start of the scoring project, not after the sessions are finished. If delivery is 48kHz, record at 48kHz unless there is a specific technical justification for higher.

Common Mistake: Not checking the default sample rate on field recorders and cameras before the shoot. Several prosumer cameras (including some Sony mirrorless bodies) and Zoom field recorders default to 44.1kHz out of the box. When these recordings enter a 48kHz NLE session, the mismatch creates conversion overhead and potential drift in some offline editing applications.

The fix: Check and set the sample rate on every audio-recording device before the first recording of the shoot. Write it on a gaffer tape strip on the recorder as a reminder.

Frequently Asked Questions

Can you actually hear the difference between 48kHz and 96kHz audio?

Under controlled conditions with high-quality monitoring, some listeners report marginal differences in the extreme high-frequency range above 14kHz. In standard studio monitoring conditions, reliably distinguishing 48kHz from 96kHz audio in a blind test is difficult and inconsistently achieved even by experienced engineers. AES research literature consistently finds that the perceptible benefit of 96kHz over 48kHz for standard dialogue and music content is minimal to non-existent at the delivery stage. The clearest measurable benefit appears in production sessions where audio undergoes heavy processing, where 96kHz provides headroom that reduces artifact accumulation through multiple generations.

Why does broadcast use 48kHz instead of 44.1kHz?

The EBU and SMPTE established 48kHz as the broadcast standard in the 1980s, partly for technical reasons - the mathematical relationship between 48kHz and common video frame rates simplifies sample rate conversion between audio-only and video contexts - and partly to establish a separate professional standard from the consumer CD format, which adopted 44.1kHz. EBU Tech 3250 and SMPTE ST 302 both specify 48kHz as mandatory for broadcast audio. This standard has been consistent since the 1980s and applies unchanged to theatrical, broadcast, and streaming delivery.

What sample rate should a film score be recorded at?

Record at 96kHz if the scoring budget, equipment, and studio support it - the anti-aliasing filter advantage and processing headroom are genuine benefits within the score's own production pipeline. Deliver to the re-recording mixer at 48kHz after a one-time high-quality SRC conversion. Never deliver at 44.1kHz to a film post environment. The sample rate mismatch creates workflow problems in the mix session that are entirely avoidable.

Does sample rate affect location audio backup file sizes?

Directly and proportionally. A 1-hour, 6-channel, 24-bit recording at 48kHz is approximately 1.24 GB. The same recording at 96kHz is approximately 2.48 GB. The Audio Bitrate Storage Calculator calculates exact file sizes for any combination of sample rate, bit depth, channels, and duration. On a 10-day documentary shoot with 8 hours of daily recording on a 6-channel recorder, the storage difference between 48kHz and 96kHz is approximately 75 GB.

Is there any film delivery format that actually requires 96kHz?

A small number of delivery contexts specify 96kHz. Dolby Atmos master sessions for theatrical exhibition can be delivered at 96kHz at Atmos-certified mixing facilities. Netflix's hi-res Atmos tier accepts 96kHz stems for immersive audio mixing. Some premium large format (PLF) theatrical digital cinema masters include 96kHz audio. For standard theatrical DCP, broadcast, and streaming delivery, 48kHz is the universal standard. If a specific platform or contract requires 96kHz, that requirement will appear explicitly in the technical delivery requirements - no interpretation is needed.

The Sample Rate Converter calculates the output parameters of any sample rate conversion - useful for planning a 96kHz to 48kHz conversion in a score-to-mix pipeline. The Audio Bitrate Storage Calculator shows exactly how much more storage a 96kHz project requires across a full production schedule versus 48kHz. For sample rate in the context of a complete audio deliverables specification, audio delivery standards for film and television covers LUFS targets, channel configurations, and bit depth alongside sample rate requirements. The LUFS, dBFS, and loudness normalization post covers the loudness compliance side of delivery that sample rate selection doesn't affect but must be resolved before any master ships.

48kHz Is the Answer Unless You Have a Specific Reason Otherwise

48kHz is the correct sample rate for all film and television audio delivery. 96kHz has a genuine, non-placebo role in production and scoring sessions where audio undergoes multiple processing stages before the 48kHz final delivery - the headroom reduces artifact accumulation through those stages. 192kHz has real value only in scientific documentation and specific archival contexts.

For any project whose final deliverable is broadcast, streaming, or theatrical DCP, recording location audio or sound effects above 48kHz adds storage cost, CPU overhead, and workflow complexity without improving what the audience hears. The decision should follow the deliverable specification, not the assumption that higher numbers are always better.

What's the highest sample rate you've used on a production - and looking back, did the final deliverable justify the workflow overhead?