Spatial audio capture method and system, wearable device and storage medium
By building a virtual audio space and determining the target sound source based on multimodal environmental information, the problem of low accuracy in spatial audio capture technology is solved, and higher spatial audio capture accuracy and immersion are achieved.
Patent Information
- Application Number
- CN202510306989.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-15
- Publication Date
- 2025-05-30
AI Technical Summary
Traditional audio capture technology has low accuracy when capturing spatial audio, making it difficult to meet the needs of modern entertainment and communication systems for immersion and realism.
By building a virtual audio space based on multimodal environmental information, multiple virtual sound sources in the virtual audio space assist in determining the target sound source, and building a real audio space based on the target sound source, achieving more accurate spatial audio positioning and immersion improvement.
Improves the accuracy of spatial audio capture, enhances immersion and user experience quality in virtual reality and augmented reality applications, while reducing the computing burden and power consumption of wearable devices.
Smart Images

Figure CN120075726A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the technical field of audio processing, and more specifically, relates to a method and system for spatial audio capture, a wearable device, and a storage medium. Background Art
[0002] With the rapid development of technology, virtual reality (VR) and augmented reality (AR) technologies have gradually entered people's lives and are widely used in many fields such as entertainment, education, and communication. In these application scenarios, users' requirements for immersive experiences are getting higher and higher. Whether exploring fantasy scenes in the virtual world or interacting by superimposing virtual information on the real scene, users expect to obtain a more real and immersive audio experience. As a key element in enhancing immersion, the importance of spatial audio is becoming increasingly prominent.
[0003] However, traditional audio capture technologies still have the problem of low accuracy when capturing spatial audio, making it difficult to meet the requirements of modern entertainment and communication systems for immersion and realism. Summary of the Invention
[0004] The purpose of the present disclosure is to provide a method and system for spatial audio capture, a wearable device, and a storage medium to improve the accuracy of spatial audio capture.
[0005] In the first aspect of the embodiments of the present disclosure, a method for spatial audio capture is provided, which is applied to a wearable device and includes: Constructing a virtual audio space based on multimodal environmental information; the virtual audio space is a space composed of multiple virtual sound sources; In response to receiving a first audio signal and a second audio signal, determining a target sound source from the virtual audio space based on the first audio signal and the second audio signal; the first audio signal and the second audio signal are audio signals collected by different audio collection devices; the audio collection devices are arranged to simulate the positional relationship of human ears; Constructing a real audio space based on the target sound source; the number of sound sources in the virtual audio space is greater than or equal to the number of sound sources in the real audio space.
[0006] In the second aspect of the embodiments of the present disclosure, a system for spatial audio capture is provided, which is applied to a wearable device and includes: A virtual audio space construction module for constructing a virtual audio space based on multimodal environmental information; the virtual audio space is a space composed of multiple virtual sound sources; A spatial audio capture module, configured to determine a target sound source from the virtual audio space based on the first audio signal and the second audio signal in response to receiving the first audio signal and the second audio signal; the first audio signal and the second audio signal are audio signals collected by different audio capture devices; the audio capture devices are arranged by simulating the positional relationship of human ears. A real - world audio space construction module, configured to construct a real - world audio space based on the target sound source; the number of sound sources in the virtual audio space is greater than or equal to the number of sound sources in the real - world audio space.
[0007] In a third aspect of the embodiments of the present disclosure, a wearable device is provided, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the above - mentioned spatial audio capture method are implemented.
[0008] In a fourth aspect of the embodiments of the present disclosure, a computer - readable storage medium is provided. The computer - readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above - mentioned spatial audio capture method are implemented.
[0009] The beneficial effects of the spatial audio capture method, system, wearable device, and storage medium provided by the embodiments of the present disclosure are as follows: On the one hand, the embodiments of the present disclosure construct a virtual audio space based on multi - modal environmental information. When constructing the virtual audio space, all environmental entity substances reflected in the multi - modal environmental information are regarded as potential sound sources, which can provide a data basis for generating rich simulated sound effects. Combining the audio signals collected by audio capture devices arranged by simulating the positional relationship of human ears, and determining the target sound source corresponding to the audio signal with the assistance of the virtual audio space, more accurate spatial audio positioning can be achieved. The finally constructed real - world audio space can highly restore the sound direction and environmental sound effects in the real world. In virtual reality or augmented reality application scenarios, the immersion can be enhanced and the user experience quality can be improved.
[0010] On the other hand, the number of sound sources in the virtual audio space is greater than or equal to the number of sound sources in the real - world audio space, realizing comprehensive simulation in the virtual audio space and on - demand processing in the real - world audio space, reducing unnecessary audio calculations. Since wearable devices have limited resources, such a processing method reduces the computational burden and power consumption, extends the battery life, ensures the stable operation of the system, and enables users to enjoy better audio services. Description of the Drawings
[0011] To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the accompanying drawings required for use in the embodiments or the description of the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0012] Figure 1 Schematic flowchart of a spatial audio capture method provided by an embodiment of the present disclosure; Figure 2 Block diagram of the structure of a spatial audio capture system provided by an embodiment of the present disclosure; Figure 3 Schematic block diagram of a wearable device provided by an embodiment of the present disclosure. Detailed implementation manners
[0013] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures and technologies are presented to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.
[0014] To make the purpose, technical solutions, and advantages of the present disclosure clearer, the following will be described through specific embodiments with reference to the accompanying drawings.
[0015] Please refer to Figure 1 , Figure 1 Schematic flowchart of a spatial audio capture method provided by an embodiment of the present disclosure. This method is applied to a wearable device and may include S101 to S103.
[0016] S101: Construct a virtual audio space based on multi-modal environmental information. The virtual audio space is a space composed of multiple virtual sound sources.
[0017] In this embodiment, the multi-modal environmental information is environmental information data from different sources and in different forms, such as image-based environmental information, text-based environmental information, and depth data-based environmental information.
[0018] Exemplarily, the image environmental information and depth data environmental information can be collected through the image sensor and depth sensor on the wearable device respectively. The text-based environmental information, image environmental information, and depth data environmental information can also be obtained through data transmission from a smart device wirelessly connected to the wearable device.
[0019] In this embodiment, the virtual audio space includes multiple object targets, and each object target is regarded as a sound source. Each object target corresponds to a virtual sound source information and position information. The virtual sound source information is associated with the attribute / type of the object target.
[0020] Exemplarily, the object targets can include any sound-producing objects such as people, animals, plants, and devices. The virtual sound source information can include the attribute information of the object target, the audio parameter range information, and the environmental interaction parameter information.
[0021] In this embodiment, through multi-modal environmental information collection, various information about the object targets is obtained. For example, the attribute and type of the object target are identified through image environmental information, and its position is determined using depth data environmental information.
[0022] For each object target, the corresponding virtual sound source information is determined according to its attribute and type. For example, when it is recognized that the object target is a moving car, according to the attribute of the object target of the car, the virtual sound source information of the car driving sound is associated, including the frequency characteristics of the sound, the law of intensity change, etc.
[0023] When constructing the virtual audio space, the virtual sound source information and position information corresponding to each object target are integrated. Through audio rendering technology, the propagation characteristics of sound in the space are simulated, such as the attenuation and reflection of sound. According to the position information of the object target, the time difference and intensity difference of the sound propagating to different positions are calculated, so that the virtual audio space can truly reflect the spatial distribution of the sounds of each object target.
[0024] Exemplarily, image recognition technology is used to analyze the image data in the multi-modal environmental information to identify various object targets, and the attributes and types of each object target are marked. At the same time, based on the depth data environmental information, the position information of each object target is determined to form spatial position coordinates (x, y, z).
[0025] A virtual sound source information library is established, and for object targets of different attributes and types, the corresponding virtual sound source information is pre-stored. When an object target is recognized, the matching virtual sound source information is retrieved from the virtual sound source information library.
[0026] In the virtual audio space, a virtual sound source information is assigned to each object target, and its virtual sound source information and position information are associated. Using an audio rendering algorithm, according to the position of the object target, the changes when the sound propagates to different positions are simulated. As the object targets in the environment change, the virtual audio space is updated in real time.
[0027] S102: In response to receiving the first audio signal and the second audio signal, determine a target sound source from the virtual audio space based on the first audio signal and the second audio signal. The first audio signal and the second audio signal are audio signals collected by different audio acquisition devices. The audio acquisition devices are arranged by simulating the positional relationship of human ears.
[0028] In this embodiment, the first audio signal and the second audio signal are electrical sound signals collected by different audio acquisition devices that simulate the positional relationship of human ears. The target sound source is the actual sound-emitting source selected from the virtual audio space based on the first audio signal and the second audio signal.
[0029] Exemplarily, in a classroom, a teacher is giving a lesson. Based on the classroom environment information, a virtual audio space including multiple target objects such as desks, chairs, students, teachers, fans, blackboards, etc. can be constructed. And through the first audio signal and the second audio signal collected in this environment, it can be analyzed that the target sound sources include the teacher and the fan.
[0030] S103: Construct a real audio space based on the target sound source. The number of sound sources in the virtual audio space is greater than or equal to the number of sound sources in the real audio space.
[0031] In this embodiment, the real audio space is an audio space constructed based on the target sound source. The construction of the real audio space can provide users with an immersive audio experience closely fitting the real scene.
[0032] In this embodiment, after determining the target sound source from the virtual audio space, other virtual sound sources in the virtual audio space except the target sound source are removed, thereby constructing an audio space that restores reality and providing users with a real audio perception.
[0033] It can be concluded from the above that, on the one hand, the embodiments of the present disclosure construct a virtual audio space based on multi-modal environment information. When constructing the virtual audio space, all environmental entity substances reflected in the multi-modal environment information are regarded as potential sound sources, which can provide a data basis for generating rich simulated sound effects. Combining the audio signals collected by the audio acquisition devices arranged by simulating the positional relationship of human ears, and using the virtual audio space to assist in determining the target sound source corresponding to the audio signal, more accurate spatial audio positioning can be achieved. The finally constructed real audio space can highly restore the sound direction and environmental sound effects in the real world. It can enhance the immersion and improve the user experience quality in virtual reality or augmented reality application scenarios.
[0034] On the other hand, the number of sound sources in the virtual audio space is greater than or equal to that in the real audio space, enabling comprehensive simulation in the virtual audio space and on-demand processing in the real audio space, thereby reducing unnecessary audio calculations. Wearable devices have limited resources, and such a processing method reduces the computational burden and power consumption, extends battery life, ensures stable system operation, and allows users to enjoy better-quality audio services.
[0035] In an embodiment of the present disclosure, the multi-modal environmental information includes image environmental information and depth data environmental information.
[0036] Constructing a virtual audio space based on the multi-modal environmental information includes: Obtaining text description environmental information based on the image environmental information.
[0037] Determining multiple virtual sound source information based on the image environmental information and the text description environmental information.
[0038] Determining the spatial position information corresponding to the multiple virtual sound source information based on the image environmental information and the depth data environmental information.
[0039] Constructing a virtual audio space based on the multiple virtual sound source information and the spatial position information corresponding to each virtual sound source information.
[0040] In this embodiment, the multiple virtual sound source information includes virtual sound source information of dynamic targets and virtual sound source information of static targets. The spatial position information corresponding to the multiple virtual sound source information includes the movement range information of dynamic targets and the position information of static targets.
[0041] Determining the spatial position information corresponding to the multiple virtual sound source information based on the image environmental information and the depth data environmental information includes: Determining the position information of static targets based on the depth data environmental information.
[0042] Determining the movement range information of dynamic targets based on the depth data environmental information and the image environmental information.
[0043] In this embodiment, the image environment information is the environmental image data collected by an image sensor. The collection of the image environment information is used to identify the objects, scenes, and spatial structures in the environment. The depth data environment information is the distance data between the object and the sensor obtained through a depth sensor, which is used to determine the spatial position and motion state of the object. The text description environment information is the text description generated by analyzing the image environment information using image recognition and natural language processing technologies. The text description environment information may include information such as object categories, scene types, and relative position relationships between objects. The spatial position information is the position data of the virtual sound source in the virtual audio space. For static targets, it is fixed position coordinates, and for dynamic targets, it is the movement range information, which may include the initial position, movement direction, and movable area, etc.
[0044] In this embodiment, the environmental data is collected using an image sensor and a depth sensor. The image environment information can be processed to generate a text description to assist in understanding the image content. By analyzing the image and depth data, the object targets (dynamic targets or static targets) and scenes in the environment can be identified. Multiple dynamic targets and static targets are used as potential sound sources, and a virtual sound source information is constructed for each object target to determine the position information of each object target. The virtual sound source information and the position information are combined, and at the same time, an audio rendering technology is used to construct a virtual audio space.
[0045] Exemplarily, the image sensor is used to collect the image environment information at a specific frame rate and resolution, and the depth sensor synchronously obtains the depth data environment information. Image recognition technology is used to extract the object and scene features in the image, and the text description environment information is generated based on natural language processing technology.
[0046] Based on the image environment information and the text description environment information, the potential sound sources in the environment are judged, and the information such as the sound source type, sound intensity, and frequency range corresponding to the potential sound sources is determined to construct the virtual sound source information. The position coordinates of the static target are determined according to the depth data environment information; by analyzing multiple frames of depth data and image environment information, the initial position, movement direction, and maximum movement distance of the dynamic target are determined to obtain its movement range information. A virtual audio space is constructed based on the virtual sound source information and the corresponding spatial position information.
[0047] This embodiment can accurately simulate the spatial audio distribution in the real environment, enhance the user's immersion, and bring a more realistic audio experience to users in the fields of virtual reality, augmented reality, etc. By distinguishing and processing the dynamic targets and static targets in the environment, this embodiment can effectively handle the complex audio problems brought by movable sound sources, reduce the calculation difficulty, meet the requirements of various application scenarios, and can also improve the accuracy and efficiency of spatial audio capture.
[0048] In one embodiment of the present disclosure, the multiple virtual sound source information includes virtual sound source information of a dynamic target and virtual sound source information of a static target. The spatial position information corresponding to the multiple virtual sound source information includes the movement range information of the dynamic target and the position information of the static target.
[0049] Constructing a virtual audio space based on the multiple virtual sound source information and the spatial position information corresponding to each virtual sound source information includes: Constructing a static virtual audio space based on the position information of the static target and the virtual sound source information of the static target.
[0050] Constructing a virtual audio space based on the movement range information of the dynamic target, the virtual sound source information of the dynamic target, and the static virtual audio space.
[0051] In this embodiment, constructing a virtual audio space based on the movement range information of the dynamic target, the virtual sound source information of the dynamic target, and the static virtual audio space includes: Constructing a dynamic virtual audio space based on the movement range information of the dynamic target and the virtual sound source information of the dynamic target.
[0052] Constructing a virtual audio space based on the dynamic virtual audio space and the static virtual audio space.
[0053] In this embodiment, constructing a dynamic virtual audio space based on the movement range information of the dynamic target and the virtual sound source information of the dynamic target includes: Constructing a three-dimensional coordinate matrix based on the movement range information of the dynamic target and the target unit distance.
[0054] Determining the position sequence of the dynamic target based on the position information corresponding to all coordinates in the three-dimensional coordinate matrix.
[0055] Constructing a dynamic virtual audio space based on the position sequence of the dynamic target and the virtual sound source information.
[0056] In this embodiment, the virtual sound source information of the dynamic target includes virtual sound source related information corresponding to a movable object in the real scene. The virtual sound source information of the dynamic target may include sound source type, sound intensity, frequency range, and sound change characteristics, etc. The virtual sound source information of the static target includes virtual sound source information corresponding to a stationary object. The virtual sound source information of the static target may include a fixed sound source type, a stable sound intensity, a specific frequency range, and its sound characteristics do not change significantly over time.
[0057] The movement range information of the dynamic target is used to describe the activity area of the dynamic target in space. The position information of the static target is to determine the fixed position of the static target in the virtual audio space with three-dimensional coordinates (x, y, z).
[0058] A static virtual audio space is an audio space constructed based on static targets, providing a stable audio background. A dynamic virtual audio space is constructed based on dynamic targets and may include a position sequence of the dynamic targets, virtual sound source information, and parameters of sound propagation characteristics varying with position.
[0059] The target unit distance is the distance unit used to construct a three-dimensional coordinate matrix. The target unit distance is a preset basic distance scale for measuring the position change of a dynamic target. For example, if it is set to 0.1 meters, it represents the minimum interval distance of the position change of the dynamic target. The three-dimensional coordinate matrix is a set of three-dimensional coordinates formed by dividing the moving range of the dynamic target according to the target unit distance, and each coordinate represents a possible position where the dynamic target may appear. The position sequence of the dynamic target includes an ordered arrangement of the possible positions of the dynamic target in the three-dimensional coordinate matrix, reflecting its movement trajectory within the moving range.
[0060] In this embodiment, a static virtual audio space is constructed according to the position of the static target and the virtual sound source information, and a stable sound background in the virtual environment can also be simulated. For the dynamic target, a three-dimensional coordinate matrix is constructed through the moving range information and the target unit distance, its possible positions are determined to form a position sequence, and a dynamic virtual audio space is constructed in combination with the virtual sound source information, and a dynamic sound effect can also be simulated. The dynamic virtual audio space is superimposed on the static virtual audio space, and the sounds of the two are fused to form a complete virtual audio space, simulating the sound scene of the combination of movement and stillness in the real environment.
[0061] Exemplarily, an image sensor and a depth sensor are used to collect environmental data, and the moving range information of the dynamic target, the virtual sound source information, the position information of the static target, and the virtual sound source information are analyzed and obtained.
[0062] According to the position information and the virtual sound source information of the static target, the virtual sound source position and sound attributes are set. Considering the spatial acoustic characteristics, an audio rendering technology is used to construct a static virtual audio space.
[0063] Based on the moving range information of the dynamic target, grids are divided in the three-dimensional space according to the target unit distance to generate a three-dimensional coordinate matrix, and all position coordinates where the dynamic target may appear are determined. According to the movement law of the dynamic target (such as according to the movement direction vector and the maximum movement distance), the position sequence of the dynamic target is extracted from the three-dimensional coordinate matrix to clarify the movement trajectory. For each position in the position sequence, in combination with the virtual sound source information of the dynamic target, considering the sound propagation characteristics such as intensity attenuation caused by distance and Doppler effect generated by movement, an audio rendering technology is used to construct a dynamic virtual audio space.
[0064] Fuse the dynamic virtual audio space and the static virtual audio space, adjust the volume balance between the two, unify the reverberation time, reflection coefficient, etc., and integrate the sound propagation effects to form the final virtual audio space.
[0065] Exemplarily, in a VR game of post-apocalyptic survival, the player is in a dilapidated city street. The static virtual audio space constructs stable background sounds such as the continuous wind sound in the abandoned buildings on the street and the current sound of the equipment that has not lost power in the distance. Dynamic targets such as zombies and running survivors. The movement range of the zombies is set according to their AI and they move within a certain range around the area where the player is located. A three-dimensional coordinate matrix is constructed through their movement range and the set target unit distance to determine the position sequence, and the dynamic virtual audio space is constructed by combining virtual sound source information such as the growling sound of the zombies. The player can hear the change in the sound when the zombies approach from different directions, which is combined with the static background sound, greatly enhancing the immersion and tension of the game.
[0066] In this embodiment, by distinguishing dynamic and static targets, constructing corresponding virtual audio spaces and fusing them respectively, the sound direction and characteristics can be accurately presented, and the real environment sound can be simulated. Classify and process the dynamic and static targets, and directly construct the dynamic virtual audio space by combining the virtual sound source information after determining the position sequence of the dynamic targets, avoiding repeated calculations, improving the speed of constructing the virtual audio space, realizing efficient spatial audio capture and construction, and enhancing the user experience.
[0067] In an embodiment of the present disclosure, determining a target sound source from a virtual audio space based on a first audio signal and a second audio signal includes: Determine the number of target sound sources based on the first audio signal and the second audio signal.
[0068] Perform sound source separation on the first audio signal based on the number of target sound sources to obtain multiple first sub-audio signals containing independent sound sources.
[0069] Perform sound source separation on the second audio signal based on the number of target sound sources to obtain multiple second sub-audio signals containing independent sound sources.
[0070] For each first sub-audio signal and the second sub-audio signal corresponding to the first sub-audio signal: Determine the target sound source from the virtual audio space based on the first sub-audio signal and the second sub-audio signal.
[0071] In this embodiment, determining the target sound source from the virtual audio space based on the first sub-audio signal and the second sub-audio signal includes: Calculate the time difference information between the first sub-audio signal and the second sub-audio signal, and calculate the intensity difference information between the first sub-audio signal and the second sub-audio signal.
[0072] Determine the first sound source azimuth area based on the time difference information and the intensity difference information.
[0073] Determine the second sound source azimuth area from the first sound source azimuth area based on the first spectrum information of the first sub-audio signal and the second spectrum information of the second sub-audio signal.
[0074] Determine the target sound source from the virtual audio space based on the second sound source azimuth area.
[0075] In this embodiment, determining the first sound source azimuth area based on the time difference information and the intensity difference information includes: Determine the first sound source azimuth range based on the time difference information, and determine the second sound source azimuth range based on the intensity difference information.
[0076] Take the intersection of the first sound source azimuth range and the second sound source azimuth range as the first sound source azimuth area.
[0077] In this embodiment, determining the target sound source from the virtual audio space based on the second sound source azimuth area includes: determining a set of virtual sound source information from the virtual audio space based on the second sound source azimuth area.
[0078] If the set of virtual sound source information contains one piece of virtual sound source information, determine the target sound source based on this virtual sound source information and the corresponding position information.
[0079] If the set of virtual sound source information contains two or more pieces of virtual sound source information, determine the audio feature based on the first sub-audio signal or the second sub-audio signal.
[0080] Calculate the matching degree between the audio feature and each piece of virtual sound source information, and screen out the target virtual sound source information with the largest matching degree.
[0081] Determine the target sound source based on the target virtual sound source information and the position information corresponding to the target virtual sound source information.
[0082] In this embodiment, the number of target sound sources refers to the number of independent sound sources identified from the mixed audio signal (the first audio signal and the second audio signal). The time difference information is the difference in arrival times of the first sub-audio signal and the second sub-audio signal, and is used to judge the sound source direction. The intensity difference information is the difference in intensities of the two sub-audio signals, and is also used for sound source direction judgment. The first sound source azimuth area is the spatial azimuth area where the sound source may exist determined according to the time difference and intensity difference information.
[0083] The first spectral information and the second spectral information respectively represent the frequency components of the first sub-audio signal and the second sub-audio signal and the amplitude distribution of each frequency component. The second sound source azimuth region is a more precise possible region of the sound source further refined based on the first sound source azimuth region in combination with the spectral information, which is more precise than the first sound source azimuth region. The virtual sound source information set includes the set of virtual sound source information located within the second sound source azimuth region in the virtual audio space. The audio feature is a feature extracted from the first sub-audio signal or the second sub-audio signal for matching the virtual sound source, and the audio feature may include the energy distribution of frequency bands, harmonic structure, etc.
[0084] In this embodiment, using the principle of binaural effect, there are differences in the time and intensity when sound arrives at two audio acquisition devices (simulating the distribution of human ears) at different positions. By analyzing these time and intensity differences, the azimuth range of the sound source can be initially determined. Combining with the spectral information of the audio signal, the sound source can be located more precisely. In the virtual audio space, the possible virtual sound source information is screened according to the determined sound source azimuth region, and by calculating the matching degree between the audio feature and the virtual sound source information, the most matching virtual sound source is found, thereby determining the target sound source.
[0085] Exemplarily, audio acquisition devices simulating the positional relationship of human ears are used to obtain the first audio signal and the second audio signal, and preprocessing such as denoising and amplification is performed on the first audio signal and the second audio signal to improve the signal quality.
[0086] Based on methods such as independent component analysis or clustering analysis, the features of the first audio signal and the second audio signal are analyzed to determine the number of independent sound sources contained therein. According to the number of target sound sources, blind source separation techniques such as the FastICA algorithm are used to separate the first audio signal into multiple first sub-audio signals, and the second audio signal into multiple second sub-audio signals. Each sub-audio signal corresponds to an independent sound source, and each independent sound source corresponds to a first sub-audio signal and a second sub-audio signal.
[0087] For each pair of corresponding first sub-audio signal and second sub-audio signal, the time difference information is calculated through the cross-correlation algorithm, and the intensity difference information is calculated by comparing the signal amplitudes. According to the time difference information, the first sound source azimuth range is determined using the triangulation principle; according to the intensity difference information, combined with the attenuation characteristics of sound propagation, the second sound source azimuth range is determined, and the intersection of the two is taken to obtain the first sound source azimuth region. The spectral information of the first sub-audio signal and the second sub-audio signal is analyzed, and the distribution differences of different frequency components are compared to further narrow down the range within the first sound source azimuth region to determine the second sound source azimuth region.
[0088] In the virtual audio space, virtual sound source information located in the second sound source azimuth region is filtered out to form a virtual sound source information set. If there is only one piece of virtual sound source information in the set, it is directly determined as the target sound source; if there are multiple pieces of virtual sound source information, the audio features of the first sub-audio signal or the second sub-audio signal are extracted, the matching degrees with each piece of virtual sound source information are calculated, and the virtual sound source with the highest matching degree is selected as the target sound source, and its position information is used as the position of the target sound source. Repeat the above operations to obtain all the sound source information in the first sub-audio signal and the second sub-audio signal, and construct the real audio space.
[0089] Exemplarily, in a multiplayer online shooting VR game, the wearable devices worn by players collect the first audio signal and the second audio signal, and the signals are mixed with enemy footsteps, gunshots, and teammate shouts. Based on the first audio signal and the second audio signal, the number of target sound sources is determined, and the first audio signal and the second audio signal are separated respectively. For the first sub-audio signal and the second sub-audio signal corresponding to the footsteps, the time difference and intensity difference information are calculated to determine the approximate azimuth region of the footsteps. The azimuth range is further narrowed down through the spectral information of the first sub-audio signal and the second sub-audio signal to obtain the second sound source azimuth region where the enemy is located.
[0090] If there is only one piece of virtual sound source information in the corresponding region of the second sound source azimuth region in the virtual audio space pre-constructed according to the game scene information, the position of the target sound source is determined based on this virtual sound source information. Utilizing the binaural effect, the first sub-audio signal (left ear) and the second sub-audio signal (right ear) are parameter-adjusted and replayed to the player according to the position of the sound source in combination with the human ear characteristics of the player, so as to simulate the real binaural effect, enabling the player to perceive the position of the sound source relative to himself through the audio heard by both ears.
[0091] For example, if the sound source is on the left side of the player, then the sound reaching the left ear will be earlier and stronger than that reaching the right ear. By adjusting the time and intensity parameters of the first sub-audio signal and the second sub-audio signal, the player can perceive that the sound comes from the left side.
[0092] Sound will attenuate as the distance increases during propagation, and the attenuation degrees of different frequency components are different. After determining the position of the target sound source according to the virtual sound source information, in combination with the sound propagation characteristics, the first sub-audio signal and the second sub-audio signal are processed. For the sound source that is farther away from the player, the signal intensity is reduced, and the frequency components are appropriately adjusted, which can simulate the attenuation and change of the sound during propagation, enabling the player to perceive the distance of the sound source.
[0093] The ear shapes and head structures of each person are different, and there are also differences in the perception of sound. The audio signal is personalized according to the human ear characteristics of the player (such as the head-related transfer function HRTF). A personalized HRTF model is constructed using the ear shape and head characteristics of the player. This HRTF model can simulate the propagation effect of sound around the player's head and ears, perform filtering and other processing on the first sub-audio signal and the second sub-audio signal, making the audio more in line with the player's personal auditory perception habits, and enhancing the accuracy and immersion of sound localization.
[0094] In this embodiment, the azimuth of the sound source is gradually and accurately determined by combining the time difference, intensity difference, and spectral information. The target sound source is selected from the virtual audio space according to the azimuth of the sound source, thereby achieving precise positioning. Based on the binaural effect, sound propagation characteristics, and personal HRTF model, the audio signal is processed, enabling the user to truly feel the azimuth and distance of the sound, enhancing the immersive experience. This embodiment can process audio signals with a mixture of multiple sounds, separate and locate different sound sources, is applicable to complex audio environments, meets the high requirements for spatial audio localization in multiple scenarios such as VR games and immersive exhibitions, and improves the quality and realism of the user experience.
[0095] In an embodiment of the present disclosure, determining a target sound source from a virtual audio space based on a second sound source azimuth region includes: Determining a target sound source azimuth region corresponding to the second sound source azimuth region in the virtual audio space based on the position corresponding to the audio acquisition device in the virtual audio space.
[0096] Determining target virtual sound source information and target position information corresponding to the target virtual sound source information based on the target sound source azimuth region.
[0097] Determining the target sound source based on the target virtual sound source information and the target position information.
[0098] In this embodiment, the position corresponding to the audio acquisition device in the virtual audio space refers to the positioning point of the device that simulates the human ear to collect audio signals in the virtual audio space, which is an important reference for subsequent conversion of the azimuth region. The target sound source azimuth region is the azimuth region corresponding to the second sound source azimuth region in the virtual audio space obtained by conversion according to the coordinate system and mapping relationship of the virtual audio space. The target virtual sound source information represents the detailed information of the virtual sound source located in the target sound source azimuth region in the virtual audio space. The target position information refers to the specific coordinate position of the target virtual sound source in the virtual audio space.
[0099] In this embodiment, the audio acquisition device collects sounds to generate audio signals, and determines the second sound source azimuth area through processing. Since the virtual audio space has its own independent coordinates and space settings, it is necessary to convert the second sound source azimuth area determined based on the actual space into the virtual audio space according to specific mapping rules to obtain the target sound source azimuth area. In the target sound source azimuth area of the virtual audio space, search for the matching target virtual sound source information and the corresponding target position information, and finally determine the target sound source to realize the association between the virtual audio space and the actual sound localization.
[0100] Exemplarily, obtain the three-dimensional coordinates of the audio acquisition device in the virtual audio space as the reference point for subsequent azimuth area conversion. According to the construction rules of the virtual audio space and the mapping relationship with the actual space, convert the spatial coordinate range of the second sound source azimuth area into the coordinate range of the target sound source azimuth area in the virtual audio space. In the virtual audio space, according to the coordinate range of the target sound source azimuth area, screen out the target virtual sound source information and the corresponding target position information in this area. Integrate the screened target virtual sound source information and target position information to determine the final target sound source.
[0101] In this embodiment, by converting the second sound source azimuth area into the virtual audio space, the target sound source can be accurately located. In scenarios such as VR and AR, the immersion and realism of the audio are enhanced, and the interaction experience and perception effect of the user in the virtual scenario are improved.
[0102] Corresponding to the spatial audio capture method in the above embodiment, Figure 2 is a structural block diagram of a spatial audio capture system provided by an embodiment of the present disclosure. For the convenience of description, only the parts related to the embodiments of the present disclosure are shown. Refer to Figure 2 The spatial audio capture system 20 is applied to a wearable device and includes: a virtual audio space construction module 21, a spatial audio capture module 22, and a real audio space construction module 23.
[0103] Among them, the virtual audio space construction module 21 is used to construct a virtual audio space based on multi-modal environment information. The virtual audio space is a space composed of multiple virtual sound sources.
[0104] The spatial audio capture module 22 is used to determine a target sound source from the virtual audio space in response to receiving a first audio signal and a second audio signal. The first audio signal and the second audio signal are audio signals collected by different audio acquisition devices. The audio acquisition devices are arranged by simulating the positional relationship of human ears.
[0105] The real audio space construction module 23 is used to construct a real audio space based on the target sound source. The number of sound sources in the virtual audio space is greater than or equal to the number of sound sources in the real audio space.
[0106] In one embodiment of the present disclosure, the multi-modal environmental information includes image environmental information and depth data environmental information. The virtual audio space construction module 21 is specifically configured to obtain text description environmental information based on the image environmental information.
[0107] Determine a plurality of virtual sound source information based on the image environmental information and the text description environmental information.
[0108] Determine spatial position information corresponding to the plurality of virtual sound source information based on the image environmental information and the depth data environmental information.
[0109] Construct a virtual audio space based on the plurality of virtual sound source information and the spatial position information corresponding to each virtual sound source information.
[0110] In one embodiment of the present disclosure, the plurality of virtual sound source information includes virtual sound source information of a dynamic target and virtual sound source information of a static target. The spatial position information corresponding to the plurality of virtual sound source information includes movement range information of the dynamic target and position information of the static target. The virtual audio space construction module 21 is specifically further configured to determine the position information of the static target based on the depth data environmental information.
[0111] Determine the movement range information of the dynamic target based on the depth data environmental information and the image environmental information.
[0112] In one embodiment of the present disclosure, the plurality of virtual sound source information includes virtual sound source information of a dynamic target and virtual sound source information of a static target. The spatial position information corresponding to the plurality of virtual sound source information includes movement range information of the dynamic target and position information of the static target. The virtual audio space construction module 21 is specifically further configured to construct a static virtual audio space based on the position information of the static target and the virtual sound source information of the static target.
[0113] Construct a virtual audio space based on the movement range information of the dynamic target, the virtual sound source information of the dynamic target, and the static virtual audio space.
[0114] In one embodiment of the present disclosure, the spatial audio capture module 22 is specifically configured to determine the number of target sound sources based on the first audio signal and the second audio signal.
[0115] Perform sound source separation on the first audio signal based on the number of target sound sources to obtain a plurality of first sub-audio signals each containing an independent sound source.
[0116] Perform sound source separation on the second audio signal based on the number of target sound sources to obtain a plurality of second sub-audio signals each containing an independent sound source.
[0117] For each first sub-audio signal and the second sub-audio signal corresponding to the first sub-audio signal: Determine a target sound source from a virtual audio space based on a first sub-audio signal and a second sub-audio signal.
[0118] In an embodiment of the present disclosure, the spatial audio capture module 22 is further specifically configured to calculate the time difference information between the first sub-audio signal and the second sub-audio signal, and calculate the intensity difference information between the first sub-audio signal and the second sub-audio signal.
[0119] Determine a first sound source azimuth area based on the time difference information and the intensity difference information.
[0120] Determine a second sound source azimuth area from the first sound source azimuth area based on the first spectrum information of the first sub-audio signal and the second spectrum information of the second sub-audio signal.
[0121] Determine the target sound source from the virtual audio space based on the second sound source azimuth area.
[0122] In an embodiment of the present disclosure, the spatial audio capture module 22 is further specifically configured to determine a target sound source azimuth area corresponding to the second sound source azimuth area in the virtual audio space based on the position corresponding to the audio capture device in the virtual audio space.
[0123] Determine target virtual sound source information and target position information corresponding to the target virtual sound source information based on the target sound source azimuth area.
[0124] Determine the target sound source based on the target virtual sound source information and the target position information.
[0125] See Figure 3 , Figure 3 is a schematic block diagram of a wearable device provided by an embodiment of the present disclosure. As Figure 3 shown, the wearable device 300 in this embodiment may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 communicate with each other through a communication bus 305. The memory 304 is used to store a computer program, and the computer program includes program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of each module in the above-mentioned system embodiments, such as Figure 2 the functions of the modules 21 to 23 shown.
[0126] It should be understood that in the embodiments of the present disclosure, the so-called processor 301 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0127] The input device 302 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the direction information of the fingerprint of the user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.
[0128] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.
[0129] In a specific implementation, the processor 301, the input device 302, and the output device 303 described in the embodiments of the present disclosure may implement the implementation manners described in the first embodiment and the second embodiment of the spatial audio capture method provided by the embodiments of the present disclosure, and may also implement the implementation manner of the wearable device 300 described in the embodiments of the present disclosure, which will not be elaborated herein.
[0130] In another embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the methods of the above embodiments are implemented. It can also be completed by instructing relevant hardware through the computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0131] The computer-readable storage medium can be the internal storage unit of the wearable device in any of the foregoing embodiments, such as the hard disk or memory of the wearable device. The computer-readable storage medium can also be an external storage device of the wearable device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on the wearable device. Further, the computer-readable storage medium can also include both the internal storage unit and the external storage device of the wearable device. The computer-readable storage medium is used to store the computer program and other programs and data required by the wearable device. The computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.
[0132] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.
[0133] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the wearable device and the unit described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0134] In several embodiments provided by the present application, it should be understood that the disclosed wearable device and method can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection between each other can be an indirect coupling or communication connection through some interfaces or units, or can also be an electrical, mechanical or other form of connection.
[0135] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can also be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present disclosure.
[0136] In addition, in each embodiment of the present disclosure, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0137] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A spatial audio capture method, characterized in that: Applied to wearable devices, including: Constructing a virtual audio space based on multimodal environmental information; the virtual audio space is a space composed of multiple virtual sound sources; In response to receiving a first audio signal and a second audio signal, determining a target sound source from the virtual audio space based on the first audio signal and the second audio signal; the first audio signal and the second audio signal are audio signals collected by different audio collection devices; the audio collection devices are set to simulate the position relationship of human ears; A real audio space is constructed based on the target sound source; the number of sound sources in the virtual audio space is greater than or equal to the number of sound sources in the real audio space.
2. The spatial audio capturing method according to claim 1, wherein: The multimodal environment information includes image environment information and depth data environment information; The constructing of a virtual audio space based on multimodal environment information includes: Obtaining text description environment information based on the image environment information; Determine a plurality of virtual sound source information based on the image environment information and the text description environment information; Determine spatial position information corresponding to the plurality of virtual sound source information based on the image environment information and the depth data environment information; A virtual audio space is constructed based on the multiple virtual sound source information and the spatial position information corresponding to each virtual sound source information.
3. The spatial audio capturing method according to claim 2, wherein: The multiple virtual sound source information includes virtual sound source information of dynamic targets and virtual sound source information of static targets; the spatial position information corresponding to the multiple virtual sound source information includes moving range information of dynamic targets and position information of static targets; The determining, based on the image environment information and the depth data environment information, the spatial position information corresponding to the plurality of virtual sound source information comprises: Determine the position information of the static target based on the depth data environment information; The moving range information of the dynamic object is determined based on the depth data environment information and the image environment information.
4. The spatial audio capturing method according to claim 2, wherein: The multiple virtual sound source information includes virtual sound source information of dynamic targets and virtual sound source information of static targets; the spatial position information corresponding to the multiple virtual sound source information includes moving range information of dynamic targets and position information of static targets; The step of constructing a virtual audio space based on the plurality of virtual sound source information and the spatial position information corresponding to each virtual sound source information includes: Constructing a static virtual audio space based on the position information of the static target and the virtual sound source information of the static target; A virtual audio space is constructed based on the moving range information of the dynamic object, the virtual sound source information of the dynamic object, and the static virtual audio space.
5. The spatial audio capturing method according to claim 1, wherein: The determining a target sound source from the virtual audio space based on the first audio signal and the second audio signal includes: determining the number of target sound sources based on the first audio signal and the second audio signal; Performing sound source separation on the first audio signal based on the target sound source quantity to obtain a plurality of first sub-audio signals containing independent sound sources; Performing sound source separation on the second audio signal based on the target sound source quantity to obtain a plurality of second sub-audio signals containing independent sound sources; For each first sub audio signal and the second sub audio signal corresponding to the first sub audio signal: A target sound source is determined from the virtual audio space based on the first sub audio signal and the second sub audio signal.
6. The spatial audio capturing method according to claim 5, characterized in that: The determining a target sound source from the virtual audio space based on the first sub audio signal and the second sub audio signal comprises: Calculating time difference information between the first sub audio signal and the second sub audio signal, and calculating intensity difference information between the first sub audio signal and the second sub audio signal; Determine a first sound source azimuth area based on the time difference information and the intensity difference information; Determine a second sound source orientation area from the first sound source orientation area based on first spectrum information of the first sub audio signal and second spectrum information of the second sub audio signal; A target sound source is determined from the virtual audio space based on the second sound source orientation area.
7. The spatial audio capturing method according to claim 6, wherein: The determining a target sound source from the virtual audio space based on the second sound source orientation area includes: Determine, based on the corresponding position of the audio acquisition device in the virtual audio space, a target sound source orientation area corresponding to the second sound source orientation area in the virtual audio space; Determine target virtual sound source information and target position information corresponding to the target virtual sound source information based on the target sound source orientation area; A target sound source is determined based on the target virtual sound source information and the target position information.
8. A spatial audio capture system, characterized in that: Applied to wearable devices, including: A virtual audio space construction module, used to construct a virtual audio space based on multimodal environment information; the virtual audio space is a space composed of multiple virtual sound sources; A spatial audio capture module, configured to determine a target sound source from the virtual audio space based on the first audio signal and the second audio signal in response to receiving a first audio signal and a second audio signal; the first audio signal and the second audio signal are audio signals collected by different audio collection devices; the audio collection devices are set to simulate the position relationship of human ears; A real audio space construction module is used to construct a real audio space based on the target sound source; the number of sound sources in the virtual audio space is greater than or equal to the number of sound sources in the real audio space.
9. A wearable device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.