Level/Time Relationship: Law of the First Wavefront
The precedence effect is a psychoacoustic phenomenon described by the law of the first wavefront. If two spatially distinct sources emit the same sound, our brain will only localize the source whose wavefront arrived at our ears first, even if the first signal is relatively weaker than the second. In summary, the source closest to the listener is localized.
If the time gap between the two wavefronts exceeds a certain duration, a sense of echo will appear (~20 to 30ms). The graph below shows the echo thresholds for speech (a) and white noise bursts (b) of different durations [1].

This perceptual property is widely used in sound reinforcement (following figure), as it allows, for example, reinforcing the sound from a loudspeaker S that is a bit too distant by adding a new loudspeaker (S_{del}) closer to a listener. By delaying the sound on this new loudspeaker (so that its wavefront reaches the listener after that of the first loudspeaker), it is possible to increase the sound level while preserving the sensation that the sound is coming from the first loudspeaker.

(S_{del}(t)) is the signal emitted by a delay fill loudspeaker (S_{del}). It is closer to the listener than S, which emits the original signal (S(t)).
If (S_{del}) emits the same signal as S, then (S_{del}(t)=S(t)). The listener will have the impression that the sound comes from (S_{del}) and not from S.
For the listener to still have the impression that the signal comes from S, the signal must first be delayed by a certain minimum duration (Delta t_{del}).
We therefore first have (S_{del}=S(t + Delta t_{del})).
The minimum delay (Delta t_{del}=t_1-t_2) corresponds to the path difference (d=d_1-d_2) of the sound wave emitted by (S) and (S_{del}) to reach the listener.
The time (T) taken by a sound wave to travel a certain distance (D) is given by: (T=D/c) where (c) is the speed of sound in air ≈ 340 m/s.
Thus, (Delta t_{del}=t_1-t_2=d_1/c-d_2/c=(d_1-d_2)/c). In practice, this delay must be greater than or equal to this minimum duration: (Delta t_{del} geq (d_1 – d_2)/c).
Obviously the level of (S_{del}) must also be adjusted. In practice, this level is lower, although under certain conditions described by the previous figure (b) — which depend on the nature of the signal (speech, music, broadband spectrum…) — it is possible to amplify the delayed original signal (S(t+Delta t_{del})) by up to approximately 10 dB.
Distance Perception
As we have seen above, the main localization cues (ITD, ILD, HRTF) allow us to determine the direction from which a sound originates. But these cues do not allow us to perceive the distance separating us from a sound source.
The perception of distance relies on numerous and complex phenomena, such as context, the auditory environment, the movement of a source or listener, familiarity with certain acoustic phenomena, or even simple visual perception. Nevertheless, a few key phenomena stand out[2]:
- Sound level differences, more precisely the variation in sound level during source movement.
- Spectral modification (loss of high frequencies with distance, due to air absorption).
- Cues related to acoustics and reverberation.
Reverberation and localization
When an acoustic source radiates in a room, its energy propagates in several directions before reaching our ears. We first perceive the direct sound, then the early reflections, and finally the diffuse field (following figure).
The direct sound comes from a single direction related to the position of the source (in reverberations, this is often called the “dry” signal).
The early reflections are caused by surfaces near the source. They come from several directions surrounding the source, providing complementary localization cues as well as information about the size and tone of the room.
Finally, the diffuse field is caused by multiple and repeated reflections off all the walls of the room. The high temporal density of reflections from all directions forms a continuous sound whose amplitude decreases mainly with wall absorption. It does not directly provide cues about the localization of the source, but rather information about the characteristics of the room: its size and tone.

The critical distance
Not to be confused with the notion of critical distance that will be developed in the context of WFS.
When a sound source is very far from the listener, the level of the reverberant sound can exceed that of the direct sound. Under these conditions localization becomes difficult or even impossible.
The critical distance defines the distance at which the level of the direct sound and the level of the reverberant field are equal.
Another consequence is that intelligibility — our ability to decode the information emitted (the words spoken by a speaker, for example) — can sometimes become impossible. In conclusion, the smaller this distance, the greater the room effect.
In a church, where this distance is small, a sound distribution system is necessary to listen to and understand a text.
Thus, there is no doubt that reverberation plays a role in the perception of the distance of a sound source.
Finally, under these conditions (i.e. beyond the critical distance), and if the reverberation time is sufficiently long, the level of the diffuse field (late reflections and reverberation) can exceed the level of the early reflections that carry directional information. Localizing a source then becomes extremely difficult or even impossible.
[1]J. Breebaart and C. Faller, Spatial Audio Processing: MPEG Surround and Other Applications, John Wiley & Sons, Ltd, 2007, p. 38.
[2] E. M. Wenzel, D. R. Begault and M. Godfroy-Cooper, Immersive Sound – The Art and Science of Binaural and Multi-Channel Audio: Perception of Spatial Sound, A. Roginska and P. Geluso, Eds., 2017.
