Decoding and Interoperability

The advantage of Ambisonics encoding/decoding is that the same sound scene can be reproduced on an infinite variety of loudspeaker setups. Even first-order streams (for which 4 loudspeakers suffice) can be decoded on a setup with many more loudspeakers. Conversely, a 7th-order stream (for which a minimum of 64 loudspeakers are needed to reproduce the full precision of its 3D spatial resolution) can be decoded on 4 loudspeakers by taking only the first 4 channels of the stream (WYZX). Obviously, the precision will be reduced: the first-order stream decoded on 64 loudspeakers will still only have first-order precision, and the 7th-order stream decoded on 4 loudspeakers will then only have first-order precision.

The interoperability of Ambisonics allows any previously encoded stream to be decoded on any setup. Thus, a spatialized acoustic field encoded at order 1 requires a minimum of 3 loudspeakers for reproduction, while at order 2 it requires 5 loudspeakers (figures below).

In 2D: (N_{LS}ge 2 times N_{order} + 1 )

2D decoding of the order 1 Ambisonics stream of a spatialized mix across several loudspeaker systems. The spatial writing is preserved end to end.

An order 2 stream can however be decoded on fewer than 5 loudspeakers ((<5)) by taking only the first 4 channels (W, Y, Z, X). The encoding precision contributed by order 2 relative to order 1 will be lost (figure below).

2D decoding of the order 2 Ambisonics stream of a spatialized mix across several loudspeaker systems. The spatial writing is preserved end to end.

It is therefore important to remember that:

  • the higher the encoding order, the more precisely the spatialized field is sampled (spatially).
  • matching the number of loudspeakers to the encoding order preserves the precision of the latter.

Decoding is a critical operation for reproduction. It gives rise to a set of optimization parameters that will be developed further. These include the choice of normalisation (maxN, FuMa, N3D and SN3D), the decoding type which optimizes the apparent direction of sources (basic, maxRe, inPhase and mixed inPhase/maxRe), and the decoding method which optimizes the distribution of acoustic energy according to the loudspeaker array and its implementation irregularities (“Energy preserving” and “Allrad+“).

Normalisation sets the level ratio between the harmonics. SN3D is the adopted standard to guarantee digital interoperability between platforms (Ambix file format).

The decoding type essentially widens the listening zone — the sweet-spot size — by improving the apparent directivity of sources.

Improvement of the apparent direction of a source according to decoding type (basic, maxRe and inPhase).
Simulation: coherence zone (dark blue) for source direction reproduction without optimization (@2kHz): basic type.

Without optimization (figure above), basic type, the directional coherence zone in blue (the zone in which listeners perceive a source from the same direction) is narrow. Conversely, a maxRe optimization widens this zone (figure below).

Simulation: coherence zone (dark blue) for source direction reproduction with optimization (@2kHz): maxRe type.

Finally, the method recommended in the majority of situations is “energy preserving“.