So what makes a WAV file a WAV file? In short, a WAV file is an audio container format based on Microsoft and IBM’s RIFF specification. It stores audio data-most commonly uncompressed PCM-along with structural information in a series of data blocks called chunks. The two essential chunks are fmt for the audio format and data for the audio samples. Everything else is optional, like the metadata.

How standardization is labeled

Below is a chart showing the four standardization labels. Each chunk listed in the various tables beneath it will have one of these four labels under the “STANDARDIZATION” heading.

Reading the four-character codes

WAV files use the four-character code, or FourCC, to identify audio format and data chunks, and any other chunks specific to the audio format. Spaces are allowed as padding for a FourCC with less than four characters, so a three letter FourCC would be valid, as in the case of fmt, for example. Notice the space trailing the fmt. The padding must be on the right with blank characters. And don’t overlook letter case. FourCC values are case-sensitive. So axml and AXML are not the same. RIFF conventionally uses uppercase identifiers for registered chunk types that can apply across RIFF formats, including RIFF, LIST, and JUNK, while lowercase identifiers are generally used for form-specific chunks, such as WAVE’s fmt and data. Later extensions do not always follow this convention. Lastly, when it comes to LIST / INFO and LIST / adtl, LIST is the chunk ID while INFO & adtl are list types stored at the beginning of the LIST chunk’s data.


Core RIFF structure

This section covers the structures that make up a basic WAV file: the container/header, the audio format description, the audio sample data, auxiliary format information, and large-file extensions for files larger than four gigabytes.

ChunkCarriesDefining specificationStandardization
RIFFLittle-endian WAV containerMicrosoft / IBM RIFF (MMPIDS 1.0, 1991)Formal standard
RIFXBig-endian RIFF container, usable with any RIFF formMicrosoft / IBM RIFF (MMPIDS 1.0, 1991)Formal standard
RF64Over-4GB WAV container (EBU)EBU Tech 3306Formal standard
BW64Over-4GB file identifier (ITU)ITU-R BS.2088Formal standard
fmtCodec, channel count, sample rate, bit depth; the EXTENSIBLE form adds a channel mask and a format GUIDMicrosoft / IBM RIFF and WAVE; WAVE_FORMAT_EXTENSIBLE is a later Microsoft additionFormal standard
dataThe audio samplesMicrosoft / IBM RIFF and WAVEFormal standard
factSample count and file-dependent information for non-PCM and compressed dataMicrosoft / IBM RIFF and WAVEFormal standard
ds64The 64-bit size table that an RF64 or BW64 file uses in place of the 32-bit header sizesEBU Tech 3306; ITU-R BS.2088Formal standard
What tools get wrong. For RF64 and BW64, a problem will occur if a parser reads the 32-bit header instead of the ds64 chunk. It will read it as nonsense. The RIFF size fields for both types are set to 0xFFFFFFFF. This indicates that the sizes are stored in the ds64 chunk. The ds64 table is the authority here. The second confusion involves the fmt chunk and its WAVE_FORMAT_EXTENSIBLE form. Don’t confuse the two…they’re not interchangeable. Multichannel and high-bit-depth files rely on the extensible form's channel mask and format GUID. The base header has no way to store that info.

Broadcast Wave and EBU metadata

Broadcast Wave (BWF) is an extension of the WAV format. It was developed by the European Broadcasting Union (EBU) for professional audio production and broadcasting. The bext chunk sits at the heart of the standard. The EBU also defines a numbered series of supplement chunks that add other types of metadata. The AES defines a separate chunk for radio automation.

ChunkCarriesDefining specificationStandardization
bextDescription, originator, origination date and time, time reference, SMPTE UMID (version 1 and later), EBU R 128 loudness (version 2), and coding historyEBU Tech 3285 (defining); ITU-R BS.1352 Annex 1Formal standard
ubxtUTF-8 multilingual broadcast extension, a companion to bextITU-R BS.1352 Annex 1Formal standard
mextMPEG-1 audio extension (Layer I/II in practice)EBU Tech 3285 Supplement 1Formal standard
qltyCapturing report: a quality report and a cue sheetEBU Tech 3285 Supplement 2Formal standard
levlPeak-envelope overview dataEBU Tech 3285 Supplement 3Formal standard
linkXML linkage across a set of related BWF filesEBU Tech 3285 Supplement 4Formal standard
axmlArbitrary XML 1.0, most often Audio Definition Model (ADM) dataEBU Tech 3285 Supplement 5; ITU-R BS.2088-2 section 5 for BW64Formal standard
dbmdDolby metadataEBU Tech 3285 Supplement 6Formal standard
cartRadio traffic and continuity data for broadcast automationAES46-2002Formal standard
r64mRF64 marker chunk, a cue replacement for large filesEBU Tech 3306Formal standard

What tools get wrong. The bext chunk has changed over time, so you must pay attention to its version number. Version 0 was introduced in 1997 and contains the original fields along with coding history. Version 1 has been in use since 2001. It adds the 64-byte SMPTE UMID. Lastly, Version 2 came out in 2011. It adds the EBU R 128 loudness fields. Those loudness fields obviously do not exist in versions 0 or 1. The same bytes are reserved there and are supposed to contain zeros. If a particular program reads them as loudness values anyway, then the loudness values are meaningless. Also, if a writer fails to zero the reserved area, it may leave actual garbage in those bytes.

mext is another chunk that can be misunderstood if you’re not careful. The MPEG coding information itself is stored in the MPEG extension to the fmt chunk and in the fact chunk. The broadcast-specific information is added by the mext chunk on top of that. This includes ancillary-data information, frame-size information, and homogeneity flags. Keep in mind that it does not replace the MPEG information stored in the other chunks. Also, EBU Supplement 1 refers to MPEG-1 audio rather than specifically to Layer 2. So a reader keying on "Layer 2" alone is working from a narrower label than the spec uses.

Lastly, cart contains more than the URL and tag text that many tools expose. AES46-2002 also defines eight post-timer entries along with a reserved area between the level-reference field and the URL. Those post-timers are part of the defined structure, regardless of whether a particular program chooses to show them or not.


Object audio and the Audio Definition Model

The Audio Definition Model (ADM) describes object-based and immersive audio. Its descriptive metadata is XML, which is usually carried in the axml chunk listed above. A separate chunk maps the file’s audio tracks to the ADM description.

ChunkCarriesDefining specificationStandardization
chnaA track map linking each audio track to ADM identifiers (track UID, track and channel format, pack format)ITU-R BS.2088-2 section 8; semantics from ITU-R BS.2076Formal standard
bxmlCompressed (gzip) XML, an alternative to axmlITU-R BS.2088-2 section 6Formal standard
sxmlSegment-associated XML; Serial ADM (S-ADM) is its principal payloadITU-R BS.2088-2 section 7 (chunk); ITU-R BS.2125-1 (S-ADM payload)Formal standard

What tools get wrong. Think of sxml as the general vehicle and S-ADM as one of the payloads that can ride inside of it. Treating the two as equivalent reverses the relationship. BS.2088-2 defines sxml as a container for segment-associated XML of any kind, with S-ADM listed as the principal named payload. BS.2125-1, by contrast, specifies the S-ADM format itself but never names or defines the chunk that carries it.

FourCC identifiers are case-sensitive. Legitimate ADM data can appear under the lowercase axml FourCC. However, an uppercase AXML variant shows up in some tooling. The two codes should be recognized separately and not merged or treated as interchangeable.


Cue points, regions, and sampler data

This family marks positions and ranges in the audio data, names them, and describes how a sampler should play the file. Most of it is part of the original 1991 baseline; the sampler and instrument chunks were added in Microsoft's 1994 Multimedia Standards Update.

ChunkCarriesDefining specificationStandardization
cueCue points, as bare sample positions with identifiersMicrosoft / IBM RIFF (MMPIDS 1.0, 1991)Formal standard
plstA playlist, an ordered play sequence of cue pointsMicrosoft / IBM RIFF (MMPIDS 1.0, 1991)Formal standard
LIST / adtlAssociated-data list: the container for the cue annotations belowMicrosoft / IBM RIFF (MMPIDS 1.0, 1991)Formal standard
lablA text label for a cue pointMicrosoft / IBM RIFF (MMPIDS 1.0, 1991)Formal standard
noteA text note for a cue pointMicrosoft / IBM RIFF (MMPIDS 1.0, 1991)Formal standard
ltxtLabeled text spanning a range of samplesMicrosoft / IBM RIFF (MMPIDS 1.0, 1991)Formal standard
fileInformation about a file tied to a cue point; the meaning is application-specificMicrosoft / IBM RIFF (MMPIDS 1.0, 1991)Formal standard
smplSampler data: unity note, fine tuning, loop points, SMPTE formatMicrosoft Multimedia Standards Update (1994)Formal standard
instInstrument data: unshifted note, gain, key and velocity rangesMicrosoft Multimedia Standards Update (1994)Formal standard
slntA run of silent samples, used within a wave listMicrosoft / IBM RIFF and WAVEFormal standard
wavlA wave list: an alternating sequence of data and slnt chunksMicrosoft / IBM RIFF and WAVEFormal standard
What tools get wrong. The cue point labels you see in an editor and their positions are stored in different locations. The positions and identifiers are located in the cue chunk. The labels, notes, and timed text are found in a LIST chunk of type adtl. Each entry uses its identifier to link back to the matching cue point. If an editor displays cue points as anonymous markers, it means the program read the cue chunk and ignored adtl or that the file had no adtl at all.

Tagged text and consumer metadata

The oldest and most common metadata in WAV files: simple tagged text, plus the Exif-audio list used by some cameras and recorders.

ChunkCarriesDefining specificationStandardization
LIST / INFOTagged text fields, each a four-character tag: title (INAM), artist (IART), comment (ICMT), genre (IGNR), creation date (ICRD), and many moreMicrosoft RIFF Multimedia File ReferenceFormal standard
LIST / exifExif-audio list (list type exif): version (ever), related image (erel), time (etim), maker (ecor), model (emdl), maker note (emnt), user comment (eucm)CIPA DC-008-2012 section 5.6.3Formal standard
What tools get wrong. The Exif-audio list is a CIPA-defined structure. Most WAV tools never show it, however. emnt (maker note) and eucm (user comment) are two of its fields to watch. eucm tells you how the text is encoded with its first eight bytes. Don’t fall into the trap by assuming it is ASCII. A reader must check those bytes for ASCII, JIS, Unicode, or undefined. For emnt, just report its existence and carry on, as there’s nothing for you to decode due to it being a manufacturer-specific block of data.

Production metadata: iXML

iXML is an open standard for embedding location recording metadata that is published by Gallery (UK). It’s also the home for several vendor-defined extensions, such as the Sony ASWG iXML Extension and Steinberg’s fields. These two both live inside the iXML payload as nested elements instead of as separate chunks.

ChunkCarriesDefining specificationStandardization
iXMLProduction-sound XML: scene, take, sound roll, project, sync and speed, and a track list, plus vendor extension blocksGallery iXML Specification, Revision 3.01 (Gallery UK, October 2021)Published spec
oXMLA nullified former iXML header, renamed so it is no longer read as active iXMLGallery iXML Specification (invalidation guideline)Published spec
Worth knowing. oXML shouldn’t be parsed. Sometimes a tool needs to lengthen the XML document structure because there is no space to fit data that needs to be added. When this happens, it’s quicker to just invalidate the original header and write the new iXML chunk at the end of the file. A faithful reader labels it as a former iXML chunk and leaves it alone.

Embedded standards: XMP and ID3

Because a WAV file is built on RIFF’s flexible container structure, two metadata standards from outside the WAV world are commonly carried as chunks: Adobe’s XMP and the ID3 tag familiar from MP3.

ChunkCarriesDefining specificationStandardization
_PMXAn Adobe XMP packet (the Extensible Metadata Platform)Wrapping: Adobe XMP Specification Part 3, Storage in Files (January 2020). Payload: Part 1, ISO 16684-1Published spec
ID3 / id3An embedded ID3v2 tag: attached picture (APIC), rating (POPM), and text framesPayload: ID3v2.3 and ID3v2.4. RIFF wrapping: convention, no standards-body definitionConvention
Worth knowing. Despite the name “XMP”, a WAV file carries the XMP metadata with the identifier _PMX, not XMP_, which would seem logical. The reason for the flip is a byte-order bug in the original implementation, as documented by Adobe. You can find XMP_ in a different container format. It’s used as the XMP atom in QuickTime.

Padding and miscellaneous chunks

This section is a bit of a catch-all. Three of the chunks (JUNK, PAD, FLLR) are used as filler or reserved space within the WAV file. The other four (CSET, DISP, MD5, PEAK) serve unrelated purposes that don’t fit neatly elsewhere.

ChunkCarriesDefining specificationStandardization
JUNKGeneric filler; also used as the reserved placeholder that an RF64 or BW64 file overwrites with ds64Microsoft / IBM RIFF (the placeholder use is derivative)Formal standard
PADPadding of arbitrary size, used to preserve alignmentMicrosoft Multimedia Standards Update (1994)Formal standard
CSETCharacter set: code page, language, dialectMicrosoft / IBM RIFFFormal standard
DISPA display hint: a title as text, or an icon as a device-independent bitmapMicrosoft Multimedia Standards Update (1994)Formal standard
FLLRFiller used to align the start of audio dataConvention; associated with Pro ToolsVendor-specific
MD5A 16-byte checksum of the audio dataConvention; documented by FADGI and BWF MetaEditConvention
PEAKAn editor's peak-overview cache for fast waveform drawingNo formal specification; used by Adobe and othersConvention

What tools get wrong. PEAK and BWF’s levl chunk both describe a waveform envelope. PEAK is an editor’s own peak-overview cache. Mixing them up will produce wrong overviews. Also, a known incompatibility exists between some tools' PEAK data and other libraries.


Vendor and proprietary chunks

A WAV file can also contain chunks written by some of the various editors, samplers, recorders, and other audio software. They’re not standardized chunks, but they belong to a particular company, developer, or product. What is known about them varies quite a bit. Some have been reverse-engineered and documented, while others we still know very little about. Regardless of how well or poorly they’re understood, the goal here is to document that they exist. If known, the software or hardware they are associated with and what they appear to contain will also be shown.

ChunkCarriesDefining specificationStandardization
acidACIDized-loop information: tempo, key, beat count, and meterNo published specification; layout is community-reverse-engineeredVendor-specific
strcSlice and transient markers used alongside acidNo published specification; community-observedVendor-specific
ovwfAn overview-waveform cache, reportedly associated with Apple Logic ProNo published specificationVendor-specific
SMEDSoundminer editing data; opaqueNo published specification (Soundminer)Vendor-specific
NMIXAn opaque binary chunk, reportedly associated with NetMix sound-library metadataNo published specificationVendor-specific
RLNDRoland sampler data (the SP-404 family); the chunk is padded so the audio data begins at a fixed offsetNo published specification (Roland); a community decoder existsVendor-specific
ResUApple Logic Pro project data, stored as compressed JSON (tempo, time signature)No published specification (Apple Logic Pro)Vendor-specific
minfMedia information, associated with Steinberg Cubase and NuendoNo published specificationVendor-specific
elm1Element data, associated with Steinberg Cubase and NuendoNo published specificationVendor-specific
tlstA trigger list: triggers that fire playback of cue points or playlist entries. Written by Sonic Foundry Sound ForgeNo formal specification; documented in the Sound Forge 4.5 manual (proposed by Sonic Foundry, never registered with Microsoft)Vendor-specific
regnRegion markers, associated with Sound ForgeNo published specificationVendor-specific
rpp1 / rppProject-state data, associated with Cockos ReaperNo published specificationVendor-specific
afspMetadata written by the AFsp audio toolkitNo published specification; toolkit-observedVendor-specific
olymRecorder metadata, associated with Olympus devicesNo published specification; tool-observedVendor-specific

What tools get wrong. The RLND chunk has an unusual structural quirk. It contains padding that places the audio data at a predetermined offset. If a parser ignores this padding, then it could read the wrong data as audio.


Related formats that are not WAV chunks

You should be aware that there are a few four-character codes that are NOT WAV chunks. They are different file formats altogether. For example, FORM is found at the beginning of an AIFF or AIFF-C file. fLaC is another one. It identifies a FLAC file. And OggS marks the beginning of an Ogg stream. Then there’s Sony’s Wave64. It’s related to WAV, but IS its own container. Instead of the RIFF/WAV four-character codes, it uses 128-bit GUIDs to identify its structures. This info is included here mainly to avoid confusion if you happen to come across them while inspecting audio files. Everything else on this page concerns chunks you can actually find inside a RIFF WAV file.


Specifications referenced


About this reference

This reference is compiled and maintained by the developer behind WAVScribe, a WAV metadata editor, and the free, read-only WAVScribe Viewer. If you spot an error or a chunk that should be added, corrections are welcome through the contact page.