Splitting a speech signal into a plurality of point sources
By splitting the audio signal into sub-bands and distributing them to different locations, and combining this with the acoustic characteristics of the room, a multi-channel speaker signal is generated, which solves the problem of poor sound spatialization in virtual reality and augmented reality, and achieves a more realistic listening experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- APPLE INC
- Filing Date
- 2022-11-09
- Publication Date
- 2026-06-12
Smart Images

Figure CN116112861B_ABST
Abstract
Description
Technical Field
[0001] One aspect of this disclosure relates to spatializing sound. Background Technology
[0002] Spatial audio rendering (spatializing sound) can be described as the electronic processing of audio signals (such as microphone signals or other recorded or synthesized audio content) to generate multi-channel speaker driver signals that produce a more realistic sound perceived by the listener. For example, (a speaker's) speech signal can be electronically processed to generate virtual point sources of (human speech) perceived by the listener as emanating from a given location, such as to the listener's right or left rather than always forward or equally from all directions. Such sound is generated by spatial audio rendering algorithms that drive multi-channel speaker setups, such as stereo speakers, surround sound speakers, speaker arrays, or headphones. Summary of the Invention
[0003] Here, one aspect of this disclosure is a computer-implemented method for reproducing the sound of a data object that may produce a more realistic listening experience. An audio signal representing the sound of the data object is received by a sound engine. The object includes visual elements to be displayed, such as simulated real-world objects, like avatars. The sound engine splits the audio signal into two or more sub-band audio signals, including a first sub-band audio signal and a second sub-band audio signal. The first sub-band may be assigned to a first position within the visual element, and the second sub-band may be assigned to a second position within the visual element spaced apart from the first position. Multiple speaker driver signals are generated using the sub-band signals to produce the sound of the object.
[0004] In one respect, this is accomplished by processing the sub-band audio signals, for example, by spatializing each sub-band signal individually, such that the sound in the first sub-band originates from a different location than the sound in the second sub-band. Thus, taking a speech signal as an example, the speech signal from a single virtual point source (located on a virtual mouth) is split into two frequency domains or sub-band components assigned to two virtual point sources (one in the mouth and one in the chest). The mouth sub-band may be in a higher frequency range than the torso sub-band. The speaker driver signals can be dual-channel left and right headphone driver signals for driving headphones worn by a listener, or they can be speaker driver signals for a stereo or surround sound speaker system.
[0005] On the other hand, the loudspeaker driver signal can be a high-frequency signal and a low-frequency signal designed to drive the tweeter and woofer of a two-way loudspeaker system, respectively.
[0006] On the other hand, one or more cut-off frequencies defining the sub-band are set based on the room's acoustic characteristics, such as volume or size. The room's volume can be used to determine the frequency at which the sound diffuses around the room relative to its directionality. The cut-off frequencies defining the boundary between the low and high sub-bands can therefore vary depending on the room's size.
[0007] The room can be a virtual room, and the visual elements of the object are within that virtual room, both presented on a display. The listener can view the display and wear headphones (through which the sound of the object is reproduced). Alternatively, the room can be a real room, in which the listener of the reproduced sound is located, and the listener wears headphones while viewing an optical head-mounted display (such as in an augmented reality environment) in which the object is presented.
[0008] The above overview does not constitute an exhaustive list of all aspects of this disclosure. It is contemplated that this disclosure encompasses all systems and methods that can be practiced by all suitable combinations of the aspects outlined above and those disclosed in the detailed descriptions below and specifically pointed out in the claims section. Such combinations may have specific advantages not specifically set forth in the foregoing summary. Attached Figure Description
[0009] The aspects of this disclosure are illustrated by way of example and are not limited to the illustrations in the accompanying drawings, in which similar reference numerals indicate similar elements. It should be noted that references to “a” or “an” aspect in this disclosure do not necessarily refer to the same aspect, and each refers to at least one. Furthermore, for the sake of brevity and to reduce the total number of drawings, a given drawing may be used to illustrate more than one aspect of this disclosure, and for a given aspect, not all elements in that drawing may be necessary.
[0010] Figure 1 This is a block diagram of an audio system that splits the input audio signal associated with visual elements of a data object into at least two virtual sound sources and spatializes each source separately.
[0011] Figure 2 This is a block diagram of an audio system that splits the input speech signal and reproduces the speech through low-frequency speaker drivers and high-frequency speaker drivers.
[0012] Figure 3 It is a flowchart of a method for reproducing speech of a data object by splitting a speech signal into at least two sub-frequency bands for separate point sources. Detailed Implementation
[0013] Various aspects of this disclosure will now be explained with reference to the accompanying drawings. Where the shape, relative position, and other aspects of the described components are not explicitly defined, the scope of the invention is not limited to the components shown, which are for illustrative purposes only. Furthermore, while many details have been set forth, it should be understood that some aspects of this disclosure can be practiced without these details. In other instances, well-known circuits, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description.
[0014] One aspect of this disclosure is Figure 1 , Figure 1 This is a block diagram of an audio system that splits the input audio signal associated with visual elements of a data object into at least two virtual sound sources and spatializes each source. This paper describes the system by means of methodological operations (computer-implemented methods) performed by the system's data processor for spatializing the sound of the data object. This data processor can be configured by software (instructions stored in machine-readable memory) such as application development software or a simulated reality application created using application development software.
[0015] An input audio signal (e.g., a mono signal) is associated with or represents the sound of a data object (such as in a simulated reality application) represented by visual element 2. After being rendered by a video engine (not shown), visual element 2 of the data object appears on display 3. Visual element 2 can be a region of a graphical object (e.g., drawn on a 2D display), or it can be a volumetric graphical object of the data object (e.g., drawn on a 3D display). The data object can be, for example, a person, and visual element 2 is an avatar of that person, which... Figure 1 The image is depicted as having a head and torso. The audio signal represents the sound of the data object, which in the human example is human speech.
[0016] The audio system renders a single input audio signal into two or more virtual sound sources or point sources, as follows. Splitter 4 splits the audio signal into two or more sub-band audio signals (components of the input audio signal), including a first sub-band (sub-band A) and a second sub-band (sub-band B). The splitter can be implemented, for example, as a filter bank. Sub-band A can be in a higher frequency range than sub-band B within the human hearing range. As an example, the low-frequency band (sub-band B) can be in the range of 50Hz-200Hz. In another example, the low-frequency band is in the range of 100Hz-300Hz. The high-frequency band can be higher than those ranges.
[0017] Sub-band A is assigned to a first position within a visual element, which is within the region or volume of the visual element, while a second sub-band is assigned to a second position within the visual element spaced apart from the first position (but also within the region or volume of the visual element). As shown, sub-band A is spatialized as a virtual sound source A or point source located at the head or mouth of a person or avatar, while sub-band B is spatialized as a virtual sound source B located at the torso of a person or avatar. The system processes the audio signals of these two sub-bands and their associated metadata (including their respective virtual source positions) to cause the sound of sub-band A and the sound of sub-band B to be emitted from different positions, thereby generating a set of multi-channel speaker driver signals (two or more speaker driver signals) that drive the listening device to produce the sound of the data object. Note that the position of the virtual sound source can be equivalent to an azimuth or angle, and an elevation or angle, for example, as observed from a virtual listening position.
[0018] exist Figure 1 In the example, sub-bands A and B are spatialized, described by two spatializer blocks A and B, which receive the same virtual listening position but different virtual source positions and different audio signals as input. The outputs of spatializers A and B are combined with the outputs of one or more other spatializers C, etc., via combiner 7 (described by the summation symbol), so that the multichannel speaker signal contains a sound scene that may have other virtual sound sources C, etc. In the example shown, the multichannel speaker signal is a two-channel signal driving the left and right speakers of headphones, but in other versions, the listening device may be different, for example, a pair of speakers, a 5.1 surround sound speaker system.
[0019] Figure 1 This also illustrates another aspect of the disclosure, where the splitter 4 is controlled by the acoustic characteristics of a room. The room can be a virtual room, where the data object (its visual element 2) is presented on a display 3. Alternatively, the room can be a real room in which a spatialized sound listener is located. In this case, the listening device can be a headset worn by the listener, and the listener can view the real room through an optical head-mounted display (also worn by the listener) while the visual element 2 of the data object is being presented on display 3, thus overlaying the real room as in an augmented reality environment. In both cases, the processor can set one or more cutoff frequencies of the sub-band audio signal based on the room's acoustic characteristics. The room's acoustic characteristics can be a function of, for example, room size or volume (e.g., large vs. small), reverberation time, sound absorption characteristics, and room impulse response, among others.
[0020] Turn now Figure 2This is a block diagram of an example computer system where the audio signal is a speech signal. The speech signal is an audio signal whose content is primarily or significantly human speech, for example, a recording of a part of a conversation. Therefore, the speech signal does not contain music or special effects. The speech signal is associated with a visual element 2 as an avatar of a data object, such as an avatar in a simulated reality application, for example. In this system, as in... Figure 1 In this system, the data processor is configured to act as a splitter 4, which splits the speech signal into at least two components, such as a first sub-band signal in a first sub-band A and a second sub-band signal in a second sub-band B. It then generates multiple speaker driver signals, in this case, tweeter signals (for driving a tweeter represented by a smaller speaker symbol) and woofer signals (for driving a woofer represented by a larger speaker symbol). The tweeter and woofer form a dual-channel speaker system (e.g., integrated into the same housing of the listening device). Therefore, it is not performed as spatial acoustics that spatializes two sub-bands separately, but rather... Figure 2 The processor in the receiver causes sound from the first sub-band A to be emitted from the tweeter of the listening device, and causes sound from the second sub-band B to be emitted from the woofer of the listening device. The first sub-band A is a high-frequency band, and the second sub-band B is a low-frequency band, where the high-frequency band is higher than the low-frequency band. (As described above...) Figure 1 The description provides examples of these frequency bands. Furthermore, this can be combined with the above text. Figure 1 Modify the described features Figure 2 The splitter 4 is controlled by the acoustic characteristics of the room.
[0021] Figure 3 This is a flowchart of a method for reproducing speech of a data object by splitting a speech signal into at least two sub-frequency bands for separate point sources. The method can be executed by a data processor configured by instructions stored in the artifact and, in particular, in a machine-readable storage medium (memory). The method begins by receiving a speech signal of the data object (operation 9) and splitting the speech signal into a first sub-frequency band signal in a first sub-frequency band and a second sub-frequency band signal in a second sub-frequency band (operation 11). In one aspect, the processor also assigns the first sub-frequency band signal to a first position of a visual element of the data object (operation 13) and the second sub-frequency band signal to a second position of the visual element (operation 15). It generates multiple speaker driver signals to reproduce the sound of the data object in a single scene (operation 17). In one instance, the spatialization process generates speaker driver signals such that the sound of the first sub-frequency band signal is emitted from a first virtual position, and the sound of the second sub-frequency band signal is emitted from a second virtual position different from the first position.
[0022] In another example, instead of the sound of spatialized data objects, the sound of a first sub-band signal is generated by a high-frequency speaker driver (e.g., a tweeter), while the sound of a second sub-band signal is generated by a low-frequency speaker driver (e.g., a woofer) of a dual-channel or multi-way speaker system. These speaker drivers can be integrated into the same housing as the listening device, such as a laptop computer, tablet computer, or head-mounted device. In those cases, the listening device also has a (integrated or mounted) display.
[0023] Here, another aspect of this disclosure is to add audio processing effects to a signal processing chain performed on a sub-band A audio signal (e.g., rendered as a high-frequency band emitted from a source that is, in this case, an incarnation of the mouth) that has frequency-dependent directivity or frequency and gain-dependent directivity. Figure 1 In this context, such processing effects can be part of the spatial sound block A. This addition will affect speech equalization as the listener moves around the source; for example, when the listener is behind the avatar while it is speaking, the contrast is greater than when the listener is in front of the avatar. Adding frequency-dependent directionality effects to high-frequency band processing can result in a more realistic rendering of certain phonemes, particularly fricatives ("f", "th", "sh", "s"). Adding gain-dependent directionality to high-frequency band processing can result in a more realistic rendering of different levels of speech generation, for example, by making the speech more directional with a greater volume.
[0024] While certain aspects have been described and illustrated in the accompanying drawings, it should be understood that these aspects are merely illustrative and not limiting of the invention, and that the invention is not limited to the specific structures and arrangements shown and described, as various other modifications will be apparent to those skilled in the art. Therefore, the description is to be regarded as exemplary and not restrictive.
Claims
1. An audio system comprising a data processor configured to spatialize sound associated with visual elements being displayed on a display, the processor being configured to: The audio signal is split into multiple sub-frequency band audio signals, wherein the multiple sub-frequency band audio signals include a first sub-frequency band signal located in a first sub-frequency band and a second sub-frequency band signal located in a second sub-frequency band; and By processing the first sub-band signal and the second sub-band signal such that the first sub-band signal is spatialized to be emitted from a first position of the visual element and the second sub-band signal is spatialized to be emitted from a second position of the visual element different from the first position, a plurality of speaker driver signals are generated.
2. The system according to claim 1, wherein, In order to generate the speaker driver signal, the processor spatializes the first sub-band signal into a first virtual sound source at a first virtual location, and spatializes the second sub-band signal into a second virtual sound source at a second virtual location different from the first virtual location.
3. The system according to claim 2, wherein, The audio signal is a speech signal, and the visual element is an avatar.
4. The system according to claim 3, wherein, The first position in the avatar is in the head or mouth, and the second position in the avatar is in the torso.
5. The system according to claim 4, wherein, The first sub-frequency band is a high-frequency band and the second sub-frequency band is a low-frequency band, wherein the high-frequency band is higher than the low-frequency band.
6. The system according to claim 1, wherein, The audio signal is a speech signal, and the visual element is an embodiment associated with a data object in a simulated real-world application.
7. The system according to claim 6, wherein, The first position in the avatar is in the head or mouth, and the second position in the avatar is in the torso.
8. The system according to claim 7, wherein, The first sub-frequency band is a high-frequency band and the second sub-frequency band is a low-frequency band, wherein the high-frequency band is higher than the low-frequency band.
9. The system according to claim 8, wherein, The processor is configured to perform frequency-dependent directional processing on the first sub-band.
10. The system according to claim 8, wherein, The processor is configured to perform gain-related directional processing on the first sub-band.
11. The system according to claim 1, wherein, The processor is used for: The acoustic characteristics of the virtual room in which the visual elements are presented on the display or the real room in which the listener is located, where the spatialized sound is received; and Based on the acoustic characteristics, one or more cutoff frequencies are set for the plurality of sub-band audio signals.
12. The system according to claim 11, wherein, The acoustic characteristics include room size or room volume.
13. A method for reproducing the sound of a data object, the method comprising: The speech signal of the data object is split into a first sub-frequency band signal in the first sub-frequency band and a second sub-frequency band signal in the second sub-frequency band; Multiple speaker driver signals are generated by processing the first sub-band signal into a tweeter or high-frequency driver signal for a dual-channel speaker system and processing the second sub-band signal into a woofer or low-frequency driver signal for the dual-channel speaker system to generate the sound of the object through the dual-channel speaker system, wherein the data object is associated with a visual element in a simulated reality application, the visual element being an avatar; as well as The plurality of speaker driver signals are generated by processing the first sub-band signal and the second sub-band signal such that the first sub-band signal is spatialized to be emitted from a first position of the visual element and the second sub-band signal is spatialized to be emitted from a second position of the visual element different from the first position.
14. The method according to claim 13, wherein, The first sub-frequency band is a high-frequency band and the second sub-frequency band is a low-frequency band, wherein the high-frequency band is higher than the low-frequency band.
15. A machine-readable storage medium storing instructions that configure a processor to: The speech signal is split into a first sub-band signal in the first sub-band and a second sub-band signal in the second sub-band; Multiple speaker driver signals are generated to reproduce the sound of the speech signal, wherein the sound of the first sub-band signal is generated by a first speaker driver, and the sound of the second sub-band signal is generated by a second speaker driver. The voice signal is the voice signal of the avatar being displayed on the monitor; as well as The plurality of speaker driver signals are generated by processing the first sub-band signal and the second sub-band signal such that the first sub-band signal is spatialized to be emitted from a first position of the avatar and the second sub-band signal is spatialized to be emitted from a second position of the avatar different from the first position.
16. The machine-readable storage medium according to claim 15, wherein, The first speaker driver is located at the tweeter, and the second speaker driver is a woofer.
17. The machine-readable storage medium according to claim 15, wherein, The first sub-frequency band is a high-frequency band and the second sub-frequency band is a low-frequency band, wherein the high-frequency band is higher than the low-frequency band.
18. The machine-readable storage medium of claim 15, further comprising instructions for configuring the processor to perform the following operations: The acoustic characteristics of the virtual room where the avatar is presented on the display or the real room in which the listener is located, receiving the sound of the reproduced audio; and One or more cutoff frequencies are set for the first sub-band and the second sub-band based on the acoustic characteristics.
19. The machine-readable storage medium according to claim 18, wherein, The acoustic characteristics include room size or room volume.
Citation Information
Patent Citations
Mixed reality system with spatialized audio
CN109791441A
Device and method for improving communication through dichotic input of a speech signal
US20100262422A1
Surround sound simulation with virtual skeleton modeling
US20130208926A1