Audio File Delivery Specification

Audio File Delivery Specification

Pro Tools edit window showing a stems session with three tracks: DX Stem, FX Stem and MX Stem

Revised April 14, 2024

Abstract

An example specification for formatting and delivering the audio files of a motion picture: file types, bit depth and sample rate, interleaved versus mono files, file start, pops, checksums and delivery documentation. It is offered as a point of reference rather than as a recommended practice. What to archive is set out in Preparing Modern Movie Digital Audio Files for Archive, which takes precedence where the two differ.

On this page

  1. Audio File Types: Bit Depth, Sample Rate
  2. Audio File Type: Interleaved vs. Mono
  3. Set-up Considerations: File Start
  4. Set-up Considerations: Pops
  5. Delivery: Checksums
  6. Delivery: Documentation

File naming is covered in a separate example: Audio File Naming.

Audio File Types: Bit Depth, Sample Rate

Only use uncompressed Broadcast Wave (.wav) linear PCM files, with a bit depth of 24 for all deliverable items and original recordings. For original field recordings, there are benefits to using 32-bit floating point, but final mixes should reference 24 bits, using a “zero VU” reference of -20 dBFS.

As to sample rate, should finishing work be done at 96 kHz-which is not a bad idea-a duplicate set of deliverable elements must be made to the industry-standard sample rate of 48 kHz. As a rule, one could hazard a guess that 99% of all movies only use 24/48 for all production and post-production recordings.

Note that it is also standard and recommended practice to make all transfers from analog masters (multitrack or mag film) at the 96 kHz sample rate, to ensure the highest possible resolution of audio remastering. This should be done regardless of the quality of the source material-even optical film, when it is fairly impossible to tell the difference between 48 kHz and 96 kHz transfers.

Audio File Type: Interleaved vs. Mono

Multichannel audio files can be delivered as an individual file for each channel (mono), or a single file containing all channels, known as interleaved or polyphonic. Each approach has its own advantages and disadvantages.

Interleaved/Polyphonic: In a 5.1 element, all channels are wrapped into one file.

Title_ …… 51.wav

When you import that file into a workstation, you can clearly see and access the individual channels, and separate them as necessary.

A strong benefit of interleaved is that not only are the total number of files recorded reduced, but also it is impossible to “disconnect” channels/files from its sister files when copying.

Mono/individual: In a 5.1 element, you end up with six separate files. In systems such as Pro Tools, even if you’re recording in what appears to be a 5.1 track, if you don’t have interleaved box checked in the Session Set Up window, you’ll be recording to six separate files:

Title_….. 51.L.wav

Title_….. 51.C.wav

Title_….. 51.R.wav

Title_….. 51.Ls.wav

Title_….. 51.Rs.wav

Title_….. 51.Lfe.wav

One school of thought recommends mono files, a primary benefit being that anyone handling the files downstream, are able to easily verify what source channel is going to what track on the new element. For example, when making a video file master on consumer-grade equipment, it will have a graphically clear that you’re putting the 5.1 tracks down in the desired L/R/C/Lfe/Ls/Rs sequence.

On the other hand, professional mastering equipment such Clipster, Transkoder, or Resolve have no problem with interleaved files today.

Set-up Considerations: File Start

The start time on the files sometimes depends on the timecode of the first frame of picture.

If one is working for home video, the minimum “run-up” is for the audio files to start at 00:59:52:00, leading into a head pop at 00:59:58:00, and a first frame of picture at 01:00:00:00.

For theatrical release, when work is finished by reels, it is common to have the timecodes follow the reel hour. Thus, the “Picture Start” frame in the head leader for reel 3 would be at 03:00:00:00, the head pop at 03:00:06:00, and the first frame of picture at 03:00:08:00.

This eight-second pre-roll is minimum, and options for longer times include matching pre-roll of home video masters or simply having additional time for voice slates (with channel IDs) and head tones.

Set-up Considerations: Pops

All “deliverable” sound elements must have both head and tail pops, exactly two seconds before and two seconds after the last frame of picture.

These pops must match the exact length of one frame according to the frame rate, no more, no less. For 24.0 fps, it’s 2,000 samples and for 23.976 fps it’s 2,002 samples.

Tail pop is placed 48 frames from the leading edge of the last frame of picture (LFOP). Many people confuse this spec and place the sync pop 48 frames after the last frame of picture, which is incorrect. Thus, the tail pop starts 47 frames after the last frame of picture.

Delivery: Checksums

For good digital archiving hygiene, checksums should be created for all files. They must be made by the respective facility (as in sound mixing or digital intermediate) from the master files.

They should be contained in the same folder as a given deliverable element, listed in either a .csv or Excel file.

This is not the time and place to debate the various flavors of checksums-MD5, SHA-1, SHA-512, etc. They all work and to that end we’re recommending that MD5 be the checksum of choice because of its smaller size (32 characters) and its modest CPU demands-read: speed of creation and verification.

Delivery: Documentation

Also in the delivery folder should be a PDF file that “assumes nothing” on the part of a person who will handle the material in the future.

This file should note who did the work and where, the picture reference (and its frame rate); the location of head and tail pops; location of first and last hard picture cuts (to verify picture version); sound dynamic range (whether for theatrical, nearfield, or both); contact information for anyone who has questions.