An immersive remote audio transmission system and method
By combining audio acquisition, processing, and transmission modules with deep neural networks and user location selection, the problem of viewers not being able to immerse themselves in the audio in remote audio transmission systems has been solved. This enables personalized adjustment and flexible masking of audio content, thereby improving the user experience.
Patent Information
- Application Number
- CN202310453357.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-25
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2043-04-25
AI Technical Summary
Existing remote audio transmission systems fail to provide viewers with an immersive experience and cannot adjust audio content according to their preferences, resulting in reduced engagement and choice.
It employs an audio acquisition module, a user unit module, a remote communication module, and an audio processing module. It separates the sound source through a deep neural network, combines user position selection and binaural effect, performs audio phase modulation and amplitude modulation processing, mixes the left and right channel audio, and provides recording and playback functions.
It enables viewers to select audio content according to their wishes, block out unwanted sound sources, provide an immersive audio experience, improve participation and flexibility, and support the application of combining recording and VR technology.
Smart Images

Figure CN116723229B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of audio transmission, and particularly relates to an immersive remote audio transmission system and method. Background Technology
[0002] With the development of internet technology, large-scale events are frequently held online. Currently, most remote audio transmission systems simply reliably transmit the event content, neglecting the question of whether the audience can experience a truly immersive participation. This significantly reduces audience engagement and diminishes the enthusiasm of enthusiasts. Furthermore, current online events transmit all sound indiscriminately, failing to filter out noise and interference, greatly reducing user choice and agency. Even if some sounds are unpleasant, users are forced to accept them. This is especially true for many live concerts; online audiences cannot enjoy the same immersive experience as in person, and may only be interested in certain instruments. Therefore, there is an urgent need for a remote audio transmission system that allows audiences to participate in various online events with an immersive experience, and to adjust their seating and block out unwanted sounds at any time. Summary of the Invention
[0003] The purpose of this invention is to address the shortcomings of existing technologies by proposing an immersive remote audio transmission system and method.
[0004] The objective of this invention is achieved through the following technical solutions: Firstly, this invention provides an immersive remote audio transmission system, which includes an audio storage module, an audio acquisition module, a user unit module, a remote communication module, and an audio processing module.
[0005] The audio acquisition module is used to acquire audio from all designated sound sources during the event.
[0006] The user unit module is used to provide users with on-site location selection services and to receive processed on-site audio transmitted to users through headphones;
[0007] The remote communication module is used to transmit the user's location selection and personalized configuration information and to transmit the processed live audio back to the user unit.
[0008] The audio processing module is used to separate the audio tracks of different sound sources, and obtain the distance from each sound source to the user's selected position based on the user's position selection information. Taking into account the binaural effect and the air attenuation of the audio, the module performs phase modulation and amplitude modulation on the audio of each sound source after separation. Finally, the processed audio is mixed to produce the left and right channel audio provided to the user.
[0009] The audio storage module is used to store the processed audio for recording and broadcasting purposes.
[0010] Furthermore, the audio acquisition module is at least one recording device, the audio storage module is at least one smart terminal, the user unit module is at least one pair of headphones and one smart terminal, and the remote communication module is at least two devices supporting wireless communication; the audio processing module includes at least one information processor for information communication and audio data processing during playback and recording; the audio processing module includes at least one database module containing distance information required for audio processing or a distance measuring module capable of measuring distance in real time.
[0011] Furthermore, the audio processing module is at least one integrated sound separation submodule, audio phase modulation submodule, and amplitude modulation and mixing submodule;
[0012] The sound separation submodule is used to separate sounds from different sound sources using a supervised learning deep neural network;
[0013] The audio phase modulation submodule is used to delay the obtained audio phase, and the audio phase modulation submodule is implemented using a phase modulation circuit;
[0014] The amplitude modulation and mixing submodule is used to mix the processed sounds from various sound sources, and to amplify or reduce the size of the synthesized sound according to the user's wishes through amplitude modulation operations. The amplitude modulation operation is implemented using an audio amplification circuit. Sound mixing is the nonlinear superposition of waveforms from multiple audio sources. During mixing, the input audio is first standardized in terms of sampling rate, bit width, and channel parameters, and then the PCM waves are mixed. The mixing methods include three methods: linear superposition followed by averaging, adaptive weighted averaging, and multi-channel mixing.
[0015] Secondly, the present invention also provides an immersive remote audio transmission method, the specific steps of which are as follows:
[0016] (1) Distance measurement: Based on the location of the audio acquisition module and the user's selectable location, measure the distance data from the location of the audio acquisition module to the location of each sound source; to make the user feel as if they are there, consider the binaural effect, measure the distance data from each sound source involved in the sound scene to the left and right ends of each position selected by the user; store the above distance data in the distance database, and for non-stationary sound sources, install a distance measuring module in advance to measure the distance from the sound source to the left and right ends of each provided position in real time, and transmit it to the audio processing module;
[0017] (2) User selection: The user selects or switches the location of the sound source in real time, and the user can also choose to block unwanted sound sources. The user unit module transmits the location information and blocking information selected by the user to the audio processing module in real time.
[0018] (3) On-site audio acquisition: The audio acquisition module acquires the audio of each sound source in real time and sends it to the audio processing module for further processing of the audio.
[0019] (4) Audio separation operation: Based on the audio acquisition results in step (3), the audio processing module separates the corresponding audio tracks according to the differences between the sound frequencies of each sound source, so that the audio of each sound source can be processed separately in the future.
[0020] (5) Relative distance calculation: Combining the distance measurement results in step (1) and the position information in step (2), the audio processing module subtracts the distance from each sound source to the left and right ends of the position from the measured distances of each sound source to the audio acquisition module, and calculates the distance that the acquired audio should still be transmitted, which is the relative distance.
[0021] (6) Amplitude modulation operation: Amplitude modulation operation is performed on the audio separation results in step (4) in sequence; considering the air attenuation of audio transmission, combined with the relative distance calculated in step (5), the information processor of the audio processing module calculates the amplitude attenuation of the obtained audio transmission to the left and right ends of the selected position respectively, and the audio amplitude modulation module of the audio processing module modulates the amplitude of the left and right channels according to the obtained attenuation, and obtains the left and right channel audio of each sound source after amplitude modulation respectively.
[0022] (7) Phase modulation operation: Based on the relative distance obtained in step (5), calculate the time required for the audio of each sound source to propagate to the left and right ends of the selected position, and perform phase delay operation on the left and right channel audio obtained in step (6) according to the time.
[0023] (8) Audio mixing operation: Combining the audio source information blocked by the user in step (2) and the left and right channel audio of each audio source obtained in step (7), the audio of all unblocked audio sources is mixed to form the live audio required by the user.
[0024] (9) Audio transmission and storage operation: The processed audio obtained in step (8) is transmitted to the user unit module through the remote communication module and finally transmitted to the user through the headphones. On the other hand, it is input into the audio storage module for recording and playback.
[0025] (10) User reselection: If the user is not satisfied with the obtained audio, he / she can reselect the position and adjust the sound source shielding in real time, and repeat steps (3) to (9).
[0026] (11) Recording mode: Combine the audio saved in step (9) and use the audio storage module to send the audio required by the user to the user unit module.
[0027] Furthermore, in step (4), the audio separation operation of the audio processing module collects audio from various sound sources separately through multiple microphones and uses the ICA independent component analysis method to separate the audio tracks of each sound source.
[0028] Furthermore, in step (4), the audio separation operation of the audio processing module is carried out by using supervised deep learning to train the sound separation network in advance, since each sound source involved in the sound scene is known in advance, and the audio separation work is carried out using the trained deep learning network.
[0029] Furthermore, in step (4), the audio separation operation of the audio processing module uses a dual-channel laser vibrometer to collect data from the sound-emitting part of each sound source, directly obtaining the audio of each sound source without the need for audio separation.
[0030] Furthermore, the relative distance in step (5) is:
[0031] L left =L l1 -L l2
[0032] L right =L r1 -L r2
[0033] Where L left L right These are the relative distances between the left and right ears, which are also the distances that need to be considered for subsequent amplitude modulation and phase modulation operations on the audio, L. l1 L r1 L represents the distance L from each instrument to the left and right sides of the user-specified position. l2 L r2 These represent the distances from the sound source acquisition device to the left and right sides of the designated location, respectively.
[0034] Furthermore, the relative distance calculation in step (5) is only for sound sources in fixed positions. If there are sound sources with changing positions, a laser rangefinder is used to measure the distance information in real time.
[0035] Thirdly, the present invention also provides an immersive remote audio transmission method based on VR technology, the specific steps of which are as follows:
[0036] (1) Distance measurement: Based on the location of the audio acquisition module and the user's selectable location, measure the distance data from the location of the audio acquisition module to the location of each sound source; to make the user feel as if they are there, consider the binaural effect, measure the distance data from each sound source involved in the sound scene to the left and right ends of each position selected by the user; store the above distance data in the distance database, and for non-stationary sound sources, install a distance measuring module in advance to measure the distance from the sound source to the left and right ends of each provided position in real time, and transmit it to the audio processing module;
[0037] (2) User selection: The user selects or switches positions in real time, and the user can also choose to block unwanted sound sources. The user unit module transmits the user's selected position information and blocking information to the audio processing module in real time.
[0038] (3) Virtual event scene presentation: Depending on the location selected by the user, the image mapping module in the VR glasses projects a virtual event scene into the user's eyes;
[0039] (4) On-site audio acquisition: The audio acquisition module acquires the audio from each sound source in real time and sends it to the audio processing module for further processing.
[0040] (5) Audio separation operation: Based on the audio acquisition results in step (4), the audio processing module separates the corresponding audio tracks according to the differences between the sound frequencies of each sound source, so that the audio of each sound source can be processed separately in the future.
[0041] (6) Real-time monitoring of user head movements: Sensors in the VR glasses capture the rotation of the user's head and transmit this information to the audio processing module;
[0042] (7) Recalculation of relative distance: Centered on the VR glasses worn by the user, where the z-axis represents the direction the user's face is facing, the x-axis represents the horizontal direction of the face, and the y-axis represents the vertical direction of the face, the relative distance needs to be recalculated when the user turns their head. The specific calculation formula is as follows:
[0043] If only horizontal head rotation is performed, the specific calculation is as follows:
[0044] Where a is the distance from the center of the face to the user's right and left ears; θ1 is the angle at which the head is turned to the right; L 左 L 右 L represents the distance from the sound source to the left and right ears before the head is turned. 转右 L 转左 θ represents the distance of the sound source from the user's right and left ears after turning their head; 转右 θ 转左 The connection between the sound source and the user's right and left ears is L. 转右1 L 转左1 The included angle between them;
[0045] In the above data, L 转右 L 转左 For data to be obtained; L 左 L 右 And 'a' represents pre-measured data; and before the user uses it, the L value is pre-measured when the user's head-turning angle θ1 changes between [-90°, 90°]. 转右1 L 转左1 and θ 转右 θ 转左 The corresponding values, where negative angles indicate that the user is turning their head to the left. The sensors in the VR system measure the user's head turning angle θ1 in real time, and then determine L at this point through a pre-determined correspondence. 转右1 L 转左1 and θ 转右 θ 转左 The corresponding value;
[0046] At this time L 转右 The calculation formula is as follows:
[0047]
[0048] Similarly, L 转左 The calculation formula is as follows:
[0049]
[0050] in
[0051]
[0052]
[0053] If only vertical head tilting is performed, the specific calculation is as follows:
[0054] Where a is the distance from the center of the face to the user's right and left ears; θ2 is the angle at which the user tilts their head back; L 左 L 右 L represents the distance from the sound source to the left and right ears before tilting the head back. 仰右 L 仰左 L represents the distance between the sound source and the user's right and left ears when the head is tilted back. 仰右1 L 仰左1 This represents the distance the central x-axis moves backward and upward during the head-tilting process;
[0055] In the above data, L 仰右 L 仰左 The quantities to be determined are a and L. 左 L 右 For data that has been measured beforehand; θ2 and L 仰右1 L 仰左1The VR sensor can be used to measure when the user is using VR;
[0056] At this time L 仰右 The calculation formula is as follows:
[0057]
[0058] At this time L 仰左 The calculation formula is as follows:
[0059]
[0060] If only the head tilts to the right, that is, the head tilts to the right shoulder, then the specific calculation is as follows:
[0061] Where a is the distance from the center of the face to the user's left and right ears; θ3 is the angle at which the user tilts their head to the right; L 左 L 右 L represents the distance from the sound source to the left and right ears before tilting the head; 歪右 L 歪左 θ represents the distance between the sound source and the user's right and left ears after tilting the head; 歪右 θ 歪左 The connection between the sound source and the user's right and left ears is L. 歪右1 L 歪左1 The included angle between them;
[0062] In the above data, L 歪右 L 歪左 For data to be obtained; L 左 L 右 And 'a' represents pre-measured data; and before the user uses it, the L value can be pre-measured when the user's head tilt angle θ3 changes between [-90°, 90°]. 歪右1 L 歪左1 and θ 歪右 θ 歪左 The corresponding values, where negative angles indicate the user tilts their head to the left. The VR sensor measures the user's tilt angle θ3 in real time, and L can be obtained from the pre-determined correspondence. 歪右1 L 歪左1 and θ 歪右 θ 歪左 The corresponding value.
[0063] At this time L 歪右 The calculation formula is as follows:
[0064]
[0065] Similarly, L 歪左 The calculation formula is as follows:
[0066]
[0067] in:
[0068]
[0069]
[0070] When a user turns their head left or right, tilts their head up or down, or tilts their head left or right, the VR system measures the corresponding head turning angle θ1, head tilting angle θ2, and head tilting angle θ3, and calculates the final relative distance value according to the above formula. This value is then transmitted to the audio processing module, which performs real-time audio processing based on the obtained relative distance to achieve an immersive experience for the user.
[0071] (8) Amplitude modulation operation: Amplitude modulation operation is performed on the audio separation results in step (5) in sequence; considering the air attenuation of audio transmission, combined with the relative distance calculated in step (7), the information processor of the audio processing module calculates the amplitude attenuation of the obtained audio transmitted to the user's left and right ears respectively, and the audio amplitude modulation module of the audio processing module modulates the amplitude of the left and right channels according to the obtained attenuation, and obtains the left and right channel audio of each sound source after amplitude modulation respectively.
[0072] (9) Phase modulation operation: Based on the relative distance obtained in step (7), calculate the time required for the audio from each sound source to travel to the user's left and right ears, and perform phase delay operation on the left and right channel audio obtained in step (8) after amplitude modulation according to the time.
[0073] (10) Audio mixing operation: Combine the audio source information blocked by the user in step (2) and the left and right channel audio of each audio source obtained in step (8) to mix the audio of all unblocked audio sources and combine them into the live audio required by the user.
[0074] (11) Audio transmission and storage operation: The processed audio obtained in step (10) is transmitted to the user unit module through the remote communication module and finally transmitted to the user through the headphones. On the other hand, it is input into the audio storage module for recording and playback.
[0075] (12) User reselection: If the user is not satisfied with the obtained audio, he / she can reselect the position and adjust the sound source shielding in real time, and repeat steps (3) to (11).
[0076] (13) Recording and broadcasting mode: Combine the audio saved in step (11) and use the audio storage module to send the audio required by the user to the user unit module.
[0077] The beneficial effects of this invention are:
[0078] 1. The immersive audio transmission system provides users with a realistic experience of the event, creating an immersive atmosphere.
[0079] 2. Users do not need to be physically present at the event; they can participate anytime, anywhere.
[0080] 3. Taking into account the human binaural effect and the air attenuation of sound, it allows users to perceive the distance of sound as if they were at the scene, thus improving the user experience.
[0081] 4. The system provides a sound source blocking function, allowing users to block certain sound sources involved in the activity at any time, greatly improving user operability and increasing the flexibility of the activity.
[0082] 5. The pre-recorded mode also allows users who cannot participate in the live event to still experience the event, enabling them to flexibly arrange their participation time without worrying about time conflicts.
[0083] 6. Users can switch seats in real time and experience the atmosphere of the event from different seats, making the user experience more diverse compared to participating in offline events.
[0084] 7. When combined with VR technology, any action of the user will change the final audio output, making the user feel as if they are really in the event. Attached Figure Description
[0085] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0086] Figure 1 This is a block diagram of the components of an immersive audio transmission system that only plays sound.
[0087] Figure 2 This is a block diagram of a method for combining an immersive audio transmission system with VR glasses.
[0088] Figure 3 This is a simplified diagram showing the change in distance when a user turns their head horizontally in VR mode.
[0089] Figure 4 This is a simplified diagram showing the change in distance when a user tilts their head vertically in VR mode.
[0090] Figure 5 This is a simplified diagram showing the change in distance when a user tilts their head (tilts their head towards their shoulder) while using VR mode.
[0091] Figure 6 This is a schematic diagram illustrating a specific scenario where an immersive audio transmission system simply plays sound.
[0092] Figure 7 This is a flowchart illustrating a specific example of a method for playing sound in an immersive audio transmission system.
[0093] Figure 8 This is a schematic diagram illustrating a specific scenario combining an immersive audio transmission system with VR glasses. Detailed Implementation
[0094] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0095] like Figure 1 As shown, the present invention provides an immersive remote audio transmission system, comprising an audio storage module, an audio acquisition module, a user unit module, a remote communication module, and an audio processing module.
[0096] The audio acquisition module is used to acquire the audio of all designated sound sources at the event; the audio acquisition module is at least one recording device used to acquire the sound of the event.
[0097] The audio storage module is used to store the processed audio for recording and broadcasting. The audio storage module is at least one smart terminal used to store the processed audio and to receive, send and process information when the user requests recording and broadcasting.
[0098] The user unit module is used to provide users with on-site location selection services and to receive processed on-site audio transmitted to users through headphones; the user unit module consists of at least one smart terminal and one pair of headphones, wherein the smart terminal is used to process and send customer seating and sound source selection information and to receive audio, and the headphones are used to send the received audio signal to the user in left and right channels.
[0099] The remote communication module is used to transmit user information and transmit the processed on-site audio back to the user unit; the remote communication module consists of at least two wireless communication modules responsible for signal transmission and reception.
[0100] The audio processing module is used to separate the audio tracks of different sound sources, and obtain the distance from each sound source to the user's selected location based on the user's location selection information. Taking into account the binaural effect and the air attenuation of the audio, the module performs phase modulation and amplitude modulation on the audio of each sound source after separation. Finally, the processed audio is mixed to produce the left and right channel audio provided to the user.
[0101] The audio processing module is at least one integrated system comprising a sound separation algorithm, an audio phase modulation, amplitude modulation and mixing module, a database module or real-time ranging module containing distance information required for audio processing, and an information processing module. The sound separation algorithm is used to separate audio tracks from different sound sources for separate processing. The audio phase modulation, amplitude modulation and mixing module is used to process the audio tracks from different sound sources separately and mix them into the final processed audio. The database module or ranging module is used to provide a basis for the phase modulation and amplitude modulation operations of the audio. The information processing module is used for sending and receiving information and processing data.
[0102] The sound separation submodule is mainly used to separate sounds from different sound sources. This process can be naturally expressed as a supervised learning problem. As the most powerful method of supervised learning, deep neural networks can be used to learn a mapping function from the original data signal to the separation target. Therefore, the sound separation submodule can be implemented by using deep learning-based speech separation techniques such as time-domain audio networks.
[0103] The audio phase modulation submodule is mainly used to delay the obtained audio phase, and the audio phase modulation submodule can be implemented using common phase modulation circuits.
[0104] The amplitude modulation (AM) and mixing submodule is mainly used to mix the processed sounds from various sound sources and amplify or reduce the size of the synthesized sound according to the user's wishes through amplitude modulation. The basic principle of sound mixing technology is to non-linearly superimpose the waveforms of multiple audio sources according to a certain algorithm. Typically, during mixing, it is necessary to first unify the sampling rate, bit width, and channels of the input audio before mixing the PCM waves. The main mixing methods include linear superposition followed by averaging, adaptive weighted averaging (allocating weights based on the characteristics of the input stream and then averaging), and multi-channel mixing. Amplitude modulation can be implemented using common audio amplification circuits.
[0105] This invention also provides a remote audio transmission method based on an immersive remote audio transmission system, the specific steps of which are as follows:
[0106] (1) Planning of remote activities: The event organizer plans the event content in advance, selects the event venue and the venue seats to be selected by users, and selects the location for the installation of audio acquisition devices; if necessary, the audio of the sound sources involved can be collected and trained in advance to lay the groundwork for subsequent separation operations.
[0107] (2) Distance measurement: Based on the location of the audio acquisition device selected in step (1) and the available seats for the user, measure the distance from the location of the audio acquisition device to each sound source; in order to make the user feel like they are there, considering the binaural effect, measure the distance data from each sound source involved in the activity to the left and right ends of each seat available to the user, and store the above location information in the distance database. For non-stationary sound sources, a distance measuring module should be pre-installed to measure the distance between the sound source and the left and right ends of each seat provided in real time, and transmit it to the audio processing module.
[0108] (3) User selection: After the event starts, users can select or switch seats in real time, and users can also choose to block certain sound sources. The seat information and blocking information selected by the user are transmitted to the audio processing module in real time.
[0109] (4) On-site audio acquisition: After the event starts, the audio acquisition device will acquire the audio of each sound source in real time and send it to the audio processor for further processing.
[0110] (5) Audio separation operation: Based on the audio acquisition results in step (4), the audio processor separates the corresponding audio tracks according to the different sound frequencies of each sound source, so as to process the audio of each sound source separately in the future; the audio separation operation of the audio processing module in this step can use the ICA independent component analysis method to separate the audio tracks of each sound source; however, if this method is used, multiple microphones are needed to collect and separate the audio of each sound source separately; since each sound source involved in the activity is known in advance, a supervised learning deep learning method can be adopted to pre-train commonly used sound separation network structures such as Conv-TasNet network or Dual-Path-RNN, and use the trained deep learning network to perform audio separation work at the beginning of the activity; or a dual-channel laser vibrometer can be used to collect the sound-emitting part of each sound source, so that the audio of each sound source can be obtained directly without audio separation operation, and the audio obtained by this method can completely ignore the interference of other sound sources, and the results are more accurate;
[0111] (6) Relative distance calculation: Combining the distance measurement results in step (2) and the seating information in step (3), the audio processor subtracts the distance from each sound source to the left and right ends of the seat from the measured distances of each sound source to the audio acquisition device, and calculates the remaining distance that the acquired audio should be transmitted, i.e., the relative distance; the specific calculation is as follows:
[0112] L left =L l1 -L l2
[0113] L right =Lr1 -L r2
[0114] Where L left L right These are the relative distances between the left and right ears, which are also the distances that need to be considered for subsequent amplitude modulation and phase modulation operations on the audio, L. l1 L r1 L represents the distance of each instrument to the left and right sides of its designated seat. l2 L r2 These are the distances from the sound source acquisition device to the left and right sides of the designated seat, respectively. The relative distance calculation is only for sound sources in fixed positions. If there are sound sources with changing positions, a laser rangefinder should be used to measure the distance information in real time.
[0115] (7) Amplitude modulation operation: Amplitude modulation operation is performed on the audio separation results in step (5) in sequence; considering the air attenuation of audio transmission, combined with the relative distance calculated in step (6), the information processor of the audio processor calculates the amplitude attenuation of the obtained audio transmission to the left and right ends of the selected seat respectively, and the audio amplitude modulation module of the audio processor modulates the left and right channels according to the obtained attenuation, and obtains the left and right channel audio of each sound source after amplitude modulation.
[0116] (8) Phase modulation operation: Based on the relative distance obtained in step (6), calculate the time required for the audio of various sound sources to propagate to the left and right sides of the selected seat. The audio phase modulation module performs phase delay operation on the left and right channel audio obtained in step (7) according to the time.
[0117] (9) Audio mixing operation: Combining the audio source information blocked by the user in step (3) and the left and right channel audio of each audio source obtained in step (8), the audio of all unblocked audio sources is mixed to form the live audio required by the user.
[0118] (10) Audio transmission and storage operation: The processed audio obtained in step (9) is transmitted to the user unit module through the remote communication module and finally transmitted to the user through the headphones. On the other hand, it is input into the audio storage module for recording and playback.
[0119] (11) User reselection: If the user is not satisfied with the obtained audio, he / she can reselect the seat and the sound source shielding situation, and transmit the audio to the audio acquisition module at the event site through the remote communication module via the user unit module for real-time adjustment, and repeat steps (4) to (10).
[0120] (12) Recording mode: Combining the audio saved in step (10), the audio storage module is used to send the audio required by the user to the user unit module for activities that have ended.
[0121] like Figure 2 As shown, the present invention also provides an immersive audio transmission method combined with VR technology; the user unit module incorporates VR glasses, which can sense the user's head turning and other dynamic information in real time, and transmit this information to the audio processing module. The audio processing module calculates the change in distance between the sound source and the user's ears based on the user's head turning, and reprocesses the audio accordingly to achieve an "immersive" effect where the audio changes as the user's posture changes.
[0122] The VR glasses are equipped with an image mapping function, which can present a three-dimensional virtual image of the event site to the user and detect the user's head movements in real time.
[0123] The audio processing module involves calculating the user's head movements to dynamically process the audio, making the user feel as if they are watching the event live.
[0124] The remaining modules are basically the same as those that only transmit audio.
[0125] The specific steps of an immersive audio transmission method combined with VR technology are as follows:
[0126] (1) Planning of remote activities: The event organizer plans the event content in advance, selects the event venue and the venue seats to be selected by users, and selects the location for the installation of audio acquisition devices; if necessary, the audio of the sound sources involved can be collected and trained in advance to lay the groundwork for subsequent separation operations.
[0127] (2) Distance measurement: Based on the location of the audio acquisition module and the available seats for the user, the distance data from the location of the audio acquisition device to the location of each sound source is measured; in order to make the user feel as if they are there, the binaural effect is considered, and the distance data from each sound source involved in the activity to the left and right ends of each seat that the user can select is measured respectively; the above distance data is stored in the distance database, and for non-stationary sound sources, a distance measuring module is pre-installed to measure the distance from the sound source to the left and right ends of each seat provided in real time, and transmits it to the audio processing module;
[0128] (3) User selection: After the activity starts, users can select or switch seats in real time, and users can also choose to block unwanted sound sources. The user unit module transmits the seat information and blocking information selected by the user to the audio processing module in real time.
[0129] (4) Virtual event scene presentation: Depending on the location selected by the user, the image mapping module in the VR glasses projects a virtual event scene into the user's eyes;
[0130] (5) On-site audio acquisition: After the event starts, the audio acquisition module acquires the audio of each sound source in real time and sends it to the audio processing module for further processing.
[0131] (6) Audio separation operation: Based on the audio acquisition results in step (5), the audio processing module separates the corresponding audio tracks according to the differences between the sound frequencies of each sound source, so that the audio of each sound source can be processed separately in the future.
[0132] (7) Real-time monitoring of user head movements: Sensors in the VR glasses capture the rotation of the user's head and transmit this information to the audio processing module;
[0133] (8) Recalculation of relative distance: Centered on the VR glasses worn by the user, where the z-axis represents the direction the user's face is facing, the x-axis represents the horizontal direction of the face, and the y-axis represents the vertical direction of the face, the relative distance needs to be recalculated when the user turns their head. The specific calculation formula is as follows:
[0134] like Figure 3 As shown, if only horizontal head rotation is performed, the specific calculation is as follows:
[0135] Where 'a' represents the distance from the center of the face to the user's right and left ears; z1, x1, and y1 represent the coordinate axes after horizontal head turning; θ1 represents the angle of head turning to the right; and the dashed line L... 左 L 右 The solid line L represents the distance from the sound source to the left and right ears before the head is turned. 转右 L 转左 L represents the distance between the sound source and the user's right and left ears after turning their head; 转右1 The lines connecting the center of the front of the head to the right ear and the center of the back of the head to the right ear form an isosceles triangle with a vertex angle of θ1; L 转左1 The lines connecting the center of the head before turning to the left ear and the center of the head after turning to the left ear form an isosceles triangle with a vertex angle of θ1; θ 转右 θ 转左 The connection between the sound source and the user's right and left ears is L. 转右1 L 转左1 The included angle between them;
[0136] In the above data, L 转右 L 转左 For data to be obtained; L 左 L 右 And 'a' represents data measured before the activity begins; and before the user uses the device, the L value is measured in advance when the user's head turning angle θ1 changes between [-90°, 90°]. 转右1 L 转左1 and θ 转右 θ 转左The corresponding value (where a negative angle indicates the user turns their head to the left) is then used. Later, when the user participates in the activity, the sensors in the VR system measure the user's head turning angle θ1 in real time. At this point, L is calculated using the pre-determined correspondence. 转右1 L 转左1 and θ 转右 θ 转左 The corresponding value.
[0137] At this time L 转右 The calculation formula is as follows:
[0138]
[0139] Similarly, L 转左 The calculation formula is as follows:
[0140]
[0141] in
[0142]
[0143]
[0144] like Figure 4 As shown, if only vertical tilting is performed, the specific calculation is as follows:
[0145] Where 'a' represents the distance from the center of the face to the user's right and left ears; z1, x1, and y1 represent the coordinate axes after tilting the head back; θ2 represents the angle at which the user tilts their head back; and the dashed line L... 左 L 右 The solid line L represents the distance between the sound source and the left and right ears before tilting the head back. 仰右 L 仰左 L represents the distance between the sound source and the user's right and left ears when the head is tilted back. 仰右1 L 仰左1 This represents the distance the central x-axis moves backward and upward during the head-tilting process;
[0146] In the above data, L 仰右 L 仰左 The quantities to be determined are a and L. 左 L 右 For data measured before the activity begins; θ2 and L 仰右1 L 仰左1 The results can be measured by the VR's sensors when the user is using VR.
[0147] At this time L 仰右 The calculation formula is as follows:
[0148]
[0149] At this time L 仰左The calculation formula is as follows:
[0150]
[0151] like Figure 5 As shown, if only the head tilts to the right (i.e., the head tilts to the right shoulder), the specific calculation is as follows:
[0152] Where 'a' represents the distance from the center of the face to the user's left and right ears; z1, x1, and y1 represent the coordinate axes after the head tilting action; θ3 represents the angle at which the user tilts their head to the right; and the dashed line L... 左 L 右 The solid line L represents the distance from the sound source to the left and right ears before tilting the head. 歪右 L 歪左 L represents the distance between the sound source and the user's right and left ears after tilting the head; 歪右1 The lines connecting the center of the head tilted forward to the right ear and the center of the head tilted backward to the right ear form an isosceles triangle with a vertex angle of θ3; L 歪左1 The lines connecting the center of the face in front of the head to the left ear and the center of the head behind the head to the left ear form an isosceles triangle with a vertex angle of θ3; θ 歪右 θ 歪左 The connection between the sound source and the user's right and left ears is L. 歪右1 L 歪左1 The included angle between them;
[0153] In the above data, L 歪右 L 歪左 For data to be obtained; L 左 L 右 And 'a' represents data measured before the activity begins; and before the user uses the device, the L value can be pre-measured when the user's head tilt angle θ3 changes between [-90°, 90°]. 歪右1 L 歪左1 and θ 歪右 θ 歪左 The corresponding value (where a negative angle indicates the user tilts their head to the left) is then used. Later, when the user participates in the activity, the VR sensor can measure the user's head tilt angle θ3 in real time. At this point, L can be calculated using the pre-determined correspondence. 歪右1 L 歪左1 and θ 歪右 θ 歪左 The corresponding value.
[0154] At this time L 歪右 The calculation formula is as follows:
[0155]
[0156] Similarly, L 歪左 The calculation formula is as follows:
[0157]
[0158] in:
[0159]
[0160]
[0161] In summary, when a user turns their head left and right, tilts their head up and down, or tilts their head left and right, the VR system measures the corresponding head turning angle θ1, head tilting angle θ2, and head tilting angle θ3, and calculates the final relative distance value according to the above formula. This value is then transmitted to the audio processing module, which performs real-time audio processing based on the obtained relative distance to achieve an immersive experience for the user.
[0162] (9) Amplitude modulation operation: Amplitude modulation operation is performed on the audio separation results in step (6) in sequence; considering the air attenuation of audio transmission, combined with the relative distance calculated in step (8), the information processor of the audio processing module calculates the amplitude attenuation of the obtained audio transmitted to the user's left and right ears respectively, and the audio amplitude modulation module of the audio processing module modulates the amplitude of the left and right channels according to the obtained attenuation, and obtains the left and right channel audio of each sound source after amplitude modulation respectively.
[0163] (10) Phase modulation operation: Based on the relative distance obtained in step (8), calculate the time required for the audio of each sound source to propagate to the user's left and right ears, and perform phase delay operation on the left and right channel audio obtained in step (9) according to the time.
[0164] (11) Audio mixing operation: Combine the audio source information blocked by the user in step (3) and the left and right channel audio of each audio source obtained in step (9) to mix the audio of all unblocked audio sources and combine them into the live audio required by the user.
[0165] (12) Audio transmission and storage operation: The processed audio obtained in step (11) is transmitted to the user unit module through the remote communication module and finally transmitted to the user through the headphones. On the other hand, it is input into the audio storage module for recording and playback.
[0166] (13) User reselection: If the user is not satisfied with the obtained audio, he / she can reselect the seat and adjust the sound source shielding in real time, and repeat steps (4) to (12).
[0167] (14) Recording mode: Combining the audio saved in step (12), the audio storage module is used to send the audio required by the user to the user unit module for activities that have ended.
[0168] Example 1: Using the remote audio transmission system of the present invention, an online remote listening session of a symphony concert is conducted in an audio-only transmission mode.
[0169] like Figure 6 As shown, the remote audio transmission system in this embodiment consists of a microphone, a smart terminal at the event site, a user's mobile terminal, headphones, a wireless communication module, an audio processor, and a mixer, respectively serving as an audio acquisition module, an audio storage module, a user unit module, a remote communication module, and an audio processing module. The smart terminal integrates a location database and an audio processor, and is externally connected to a mixer, ensuring that the smart terminal can complete real-time communication and audio storage with the microphone used for sound acquisition at the event site. Figure 7 As shown, the specific steps for listening to a concert online are as follows:
[0170] Upload event information: The event organizer plans the event in advance, determines the event venue, the seats to be provided, the audio sources involved in the event, and other information, and uploads the required information for users to choose from.
[0171] Pre-collect audio data of each instrument and train the audio separation network: The audio of each instrument can be pre-collected using a laser vibrometer before the activity begins. Specifically, the laser vibrometer is used to directly obtain the audio data of the instrument by pointing it at the sound-producing part of the instrument. This data is then used as training data to feed into the Conv-TasNet network or the Dual-Path-RNN network for pre-training. At the start of the activity, the pre-trained deep learning network is used to perform audio separation.
[0172] User Information Selection: Users can use their mobile devices to select their seats and desired audio sources according to the information uploaded by the event organizer. If a user is dissatisfied with the transmitted symphony audio, they can switch seats or change the muting information at any time. The mobile device will transmit this information to the smart terminal at the event site in real time for further audio processing.
[0173] Audio Acquisition: The microphone captures the sound from the symphony orchestra and transmits the results to a smart terminal for processing by an audio processor.
[0174] Audio Processing: The audio processor first separates the audio tracks of each instrument from the audio captured by the microphone. This separation can be achieved using the Conv-TasNet network or Dual-Path-RNN network trained in the previous step. Based on the distances of each instrument from the left and right ends of the seating area and from the microphone, the phase delay and amplitude attenuation required for transmitting the audio to the left and right channels are calculated, and amplitude modulation and phase modulation are performed accordingly. Subsequently, the processed audio from each instrument's left and right channels is stored in the smart terminal according to the seating information tags for future recording. Finally, based on the user's provided sound source masking settings, the processed instrument audio is selectively mixed and transmitted to the user.
[0175] Example 2: Using the remote audio transmission system of the present invention in combination with VR glasses to listen to a symphony concert online.
[0176] like Figure 8 As shown, the remote listening symphony system in this example consists of a VR headset containing an audio output device, sensors, and an image mapping device, an event site smart terminal, an audio processor, and a mixer. The specific steps are as follows:
[0177] The process of uploading event information, pre-training each instrument, audio acquisition, and audio processing is the same as described in Example 1.
[0178] User information selection: Based on the information uploaded by the event organizer, users can select their seats and desired audio sources according to their preferences. If users are dissatisfied with the transmitted symphony audio, they can switch seats and change the blocking information at any time.
[0179] VR Imaging: VR glasses project the "virtual symphony concert" that the user can see from their chosen position onto the user's eyes through an image mapping device, and the image will change as the user's seat changes.
[0180] User head turning information recognition: The sensors in the VR glasses recognize the user's head movements in real time and transmit the angle information of the user's head turning to the audio processing module of the smart terminal. The distance calculation module in the audio processing module will recalculate the distance information required for audio amplitude modulation.
[0181] The above embodiments are used to explain and illustrate the present invention, but not to limit the present invention. Any modifications and changes made to the present invention within the spirit and scope of the claims shall fall within the protection scope of the present invention.
Claims
1. An immersive remote audio transmission system, characterized in that, The system includes an audio storage module, an audio acquisition module, a user unit module, a remote communication module, and an audio processing module; The audio acquisition module is used to acquire audio from all designated sound sources during the event. The user unit module is used to provide users with on-site location selection services and to receive processed on-site audio transmitted to users through headphones; The remote communication module is used to transmit the user's location selection and personalized configuration information and to transmit the processed live audio back to the user unit. The audio processing module is used to separate audio tracks from different sound sources and obtain the distances from each sound source to the user's selected location based on the user's location selection information. The audio processor subtracts the distance from each sound source to the left and right ends of the seat from the measured distances from each sound source to the audio acquisition device, and calculates the remaining transmission distance, i.e., the relative distance, of the acquired audio. The specific calculation is as follows: Where L left L right These are the relative distances between the left and right ears, which are also the distances that need to be considered for subsequent amplitude modulation and phase modulation operations on the audio, L. l1 L r1 L represents the distance of each instrument to the left and right sides of its designated seat. l2 L r2 The distances from the sound source acquisition device to the left and right sides of the designated seat are respectively; the relative distance calculation is only for sound sources in fixed positions. If there are sound sources with changing positions, a laser rangefinder should be used to measure the distance information in real time; taking into account the binaural effect and the air attenuation of audio, the audio of each sound source after separation is phase-modulated and amplitude-modulated. Finally, the processed audio is mixed to produce the left and right channel audio provided to the user; the audio processing module is at least one set integrating the sound separation submodule, the audio phase-modulation submodule, and the amplitude modulation and mixing submodule; The sound separation submodule is used to separate sounds from different sound sources using a supervised learning deep neural network; The audio phase modulation submodule is used to delay the obtained audio phase, and the audio phase modulation submodule is implemented using a phase modulation circuit; The amplitude modulation and mixing submodule is used to mix the processed sounds from various sound sources, and to amplify or reduce the size of the synthesized sound according to the user's wishes through amplitude modulation operations, which are implemented using an audio amplification circuit. Sound mixing is the nonlinear superposition of waveforms from multiple audio sources. During mixing, the input audio is first standardized in terms of sampling rate, bit width, and channel parameters, and then the PCM waves are mixed. The mixing methods include three methods: linear superposition followed by averaging, adaptive weighted averaging, and multi-channel mixing. The audio storage module is used to store the processed audio for recording and broadcasting purposes.
2. The remote audio transmission system according to claim 1, characterized in that, The audio acquisition module is at least one recording device, the audio storage module is at least one smart terminal, the user unit module is at least one pair of headphones and one smart terminal, and the remote communication module is at least two devices that support wireless communication; the audio processing module includes at least one information processor for information communication and audio data processing during playback and recording; the audio processing module includes at least one database module containing distance information required for audio processing or a distance measuring module capable of measuring distance in real time.
3. A remote audio transmission method based on the immersive remote audio transmission system according to any one of claims 1-2, characterized in that, The specific steps of this method are as follows: (1) Distance measurement: Based on the location of the audio acquisition module and the user's selectable location, measure the distance data from the location of the audio acquisition module to the location of each sound source; in order to make the user feel like they are there, consider the binaural effect, measure the distance data from each sound source involved in the sound scene to the left and right ends of each position selected by the user; store the above distance data in the distance database, and for non-stationary sound sources, install a distance measuring module in advance to measure the distance from the sound source to the left and right ends of each provided position in real time, and transmit it to the audio processing module; (2) User selection: The user selects or switches the location of the sound source in real time, and the user can also choose to block unwanted sound sources. The user unit module transmits the location information and blocking information selected by the user to the audio processing module in real time. (3) On-site audio acquisition: The audio acquisition module acquires the audio of each sound source in real time and sends it to the audio processing module for further processing of the audio; (4) Audio separation operation: Based on the audio acquisition results in step (3), the audio processing module separates the corresponding audio tracks according to the differences between the sound frequencies of each sound source, so that the audio of each sound source can be processed separately in the future. (5) Relative distance calculation: Combining the distance measurement results in step (1) and the position information in step (2), the audio processing module subtracts the distance from each sound source to the left and right ends of the position from the measured distances of each sound source to the audio acquisition module, and calculates the distance that the acquired audio should still be transmitted, which is the relative distance. (6) Amplitude modulation operation: Amplitude modulation operation is performed on the audio separation results in step (4) in sequence; considering the air attenuation of audio transmission, combined with the relative distance calculated in step (5), the information processor of the audio processing module calculates the amplitude attenuation of the obtained audio transmission to the left and right ends of the selected position respectively, and the audio amplitude modulation module of the audio processing module modulates the amplitude of the left and right channels according to the obtained attenuation, and obtains the left and right channel audio of each sound source after amplitude modulation respectively. (7) Phase modulation operation: Based on the relative distance obtained in step (5), calculate the time required for the audio of each sound source to propagate to the left and right ends of the selected position, and perform phase delay operation on the left and right channel audio obtained in step (6) according to the time. (8) Audio mixing operation: Combining the audio source information blocked by the user in step (2) and the left and right channel audio of each audio source obtained in step (7), the audio of all unblocked audio sources is mixed to form the live audio required by the user. (9) Audio transmission and storage operation: The processed audio obtained in step (8) is transmitted to the user unit module through the remote communication module and finally transmitted to the user through the headphones. On the other hand, it is input into the audio storage module for recording and playback. (10) User reselection: If the user is not satisfied with the obtained audio, he / she can reselect the position and adjust the sound source shielding in real time, and repeat steps (3) to (9). (11) Recording and broadcasting mode: Combine the audio saved in step (9) and use the audio storage module to send the audio required by the user to the user unit module.
4. The remote audio transmission method according to claim 3, characterized in that, In step (4), the audio processing module performs an audio separation operation by using multiple microphones to collect audio from various sound sources separately and using the ICA independent component analysis method to separate the audio tracks of each sound source.
5. The remote audio transmission method according to claim 3, characterized in that, In step (4), the audio separation operation of the audio processing module is carried out by supervising deep learning method to train the sound separation network in advance, since each sound source involved in the sound scene is known in advance. The trained deep learning network is then used to perform the audio separation operation.
6. The remote audio transmission method according to claim 3, characterized in that, In step (4), the audio separation operation of the audio processing module uses a dual-channel laser vibrometer to collect data from the sound source of each sound source, directly obtaining the audio of each sound source without the need for audio separation.
7. The remote audio transmission method according to claim 3, characterized in that, The relative distance in step (5) is: Where L left L right These are the relative distances between the left and right ears, which are also the distances that need to be considered for subsequent amplitude modulation and phase modulation operations on the audio, L. l1 L r1 L represents the distance L from each instrument to the left and right sides of the user-specified position. l2 L r2 These represent the distances from the sound source acquisition device to the left and right sides of the designated location, respectively.
8. The remote audio transmission method according to claim 3, characterized in that, The relative distance calculation in step (5) is only for sound sources in fixed positions. If there are sound sources with changing positions, a laser rangefinder is used to measure the distance information in real time.
9. A VR-based immersive remote audio transmission method based on the remote audio transmission method of claim 3, characterized in that, The specific steps of this method are as follows: (1) Distance measurement: Based on the location of the audio acquisition module and the user's selectable location, measure the distance data from the location of the audio acquisition module to the location of each sound source; in order to make the user feel like they are there, consider the binaural effect, measure the distance data from each sound source involved in the sound scene to the left and right ends of each position selected by the user; store the above distance data in the distance database, and for non-stationary sound sources, install a distance measuring module in advance to measure the distance from the sound source to the left and right ends of each provided position in real time, and transmit it to the audio processing module; (2) User selection: The user selects or switches positions in real time, and the user can also select to block unwanted sound sources. The user unit module transmits the user's selected position information and blocking information to the audio processing module in real time. (3) Virtual event scene presentation: Depending on the location selected by the user, the image mapping module in the VR glasses projects the virtual event scene into the user's eyes; (4) On-site audio acquisition: The audio acquisition module acquires the audio of each sound source in real time and sends it to the audio processing module for further processing of the audio; (5) Audio separation operation: Based on the audio acquisition results in step (4), the audio processing module separates the corresponding audio tracks according to the differences between the sound frequencies of each sound source, so that the audio of each sound source can be processed separately in the future. (6) Real-time monitoring of user head movements: Sensors in the VR glasses capture the rotation of the user's head and transmit this information to the audio processing module; (7) Recalculation of relative distance: Taking the VR glasses worn by the user as the center, where the z-axis represents the direction the user's face is facing, the x-axis represents the horizontal direction of the face, and the y-axis represents the vertical direction of the face, the relative distance needs to be recalculated when the user turns their head. The specific calculation formula is as follows: If only horizontal head rotation is performed, the specific calculation is as follows: Where a is the distance from the center of the face to the user's right and left ears; θ1 is the angle at which the head is turned to the right; L 左 L 右 L represents the distance from the sound source to the left and right ears before the head is turned. 转右 L 转左 θ represents the distance of the sound source from the user's right and left ears after turning their head; 转右 θ 转左 The connection between the sound source and the user's right and left ears is L. 转右1 L 转左1 The included angle between them; In the above data, L 转右 L 转左 For data to be obtained; L 左 L 右 And 'a' represents pre-measured data; and before the user uses it, the L value is pre-measured when the user's head-turning angle θ1 changes between [-90°, 90°]. 转右1 L 转左1 and θ 转右 θ 转左 The corresponding values, where negative angles indicate that the user is turning their head to the left. The sensors in the VR system measure the user's head turning angle θ1 in real time, and then determine L at this point through a pre-determined correspondence. 转右1 L 转左1 and θ 转右 θ 转左 The corresponding value; At this time L 转右 The calculation formula is as follows: Similarly, L 转左 The calculation formula is as follows: in If only vertical head tilting is performed, the specific calculation is as follows: Where a is the distance from the center of the face to the user's right and left ears; θ2 is the angle at which the user tilts their head back; L 左 L 右 L represents the distance from the sound source to the left and right ears before tilting the head back. 仰右 L 仰左 L represents the distance between the sound source and the user's right and left ears when the head is tilted back. 仰右1 L 仰左1 This represents the distance the central x-axis moves backward and upward during the head-tilting process; In the above data, L 仰右 L 仰左 The quantities to be determined are a and L. 左 L 右 Data that has been measured and completed in advance; θ2 and L 仰右1 L 仰左1 The VR sensor can be used to measure when the user is using VR; At this time L 仰右 The calculation formula is as follows: At this time L 仰左 The calculation formula is as follows: If only the head tilts to the right, that is, the head tilts to the right shoulder, then the specific calculation is as follows: Where a is the distance from the center of the face to the user's left and right ears; θ3 is the angle at which the user tilts their head to the right; L 左 L 右 L represents the distance from the sound source to the left and right ears before tilting the head; 歪右 L 歪左 θ represents the distance between the sound source and the user's right and left ears after tilting the head; 歪右 θ 歪左 The connection between the sound source and the user's right and left ears is L. 歪右1 L 歪左1 The included angle between them; In the above data, L 歪右 L 歪左 For data to be obtained; L 左 L 右 And 'a' represents pre-measured data; and before the user uses it, the L value can be pre-measured when the user's head tilt angle θ3 changes between [-90°, 90°]. 歪右1 L 歪左1 and θ 歪右 θ 歪左 The corresponding values, where negative angles indicate the user tilts their head to the left. The VR sensor measures the user's tilt angle θ3 in real time, and L can be obtained from the pre-determined correspondence. 歪右1 L 歪左1 and θ 歪右 θ 歪左 The corresponding value; At this time L 歪右 The calculation formula is as follows: Similarly, L 歪左 The calculation formula is as follows: in: When a user turns their head left or right, tilts their head up or down, or tilts their head left or right, the VR system measures the corresponding head turning angle θ1, head tilting angle θ2, and head tilting angle θ3, and calculates the final relative distance value according to the above formula. This value is then transmitted to the audio processing module, which performs real-time audio processing based on the obtained relative distance to achieve an "immersive" experience for the user. (8) Amplitude modulation operation: Amplitude modulation operation is performed on the audio separation results in step (5) in sequence; considering the air attenuation of audio transmission, combined with the relative distance calculated in step (7), the information processor of the audio processing module calculates the amplitude attenuation of the audio transmitted to the user's left and right ears respectively, and the audio amplitude modulation module of the audio processing module modulates the amplitude of the left and right channels according to the obtained attenuation, and obtains the left and right channel audio of each sound source after amplitude modulation respectively. (9) Phase modulation operation: Based on the relative distance obtained in step (7), calculate the time required for the audio of each sound source to propagate to the user's left and right ears, and perform phase delay operation on the left and right channel audio obtained in step (8) according to the time. (10) Audio mixing operation: Combining the audio source information blocked by the user in step (2) and the left and right channel audio of each audio source obtained in step (8), the audio of all unblocked audio sources is mixed to form the live audio required by the user. (11) Audio transmission and storage operation: The processed audio obtained in step (10) is transmitted to the user unit module through the remote communication module and finally transmitted to the user through the headphones. On the other hand, it is input into the audio storage module for recording and playback. (12) User reselection: If the user is not satisfied with the obtained audio, he / she can reselect the position and adjust the sound source shielding in real time, and repeat steps (3) to (11). (13) Recording and broadcasting mode: Combine the audio saved in step (11) and use the audio storage module to send the audio required by the user to the user unit module.
Citation Information
Patent Citations
An audio scene apparatus
CN105378826A
Simulation System, Sound Processing Method And Information Storage Medium
CN107277736A