Audio processing method and electronic device

CN119497030BActive Publication Date: 2026-09-18HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311035340.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-15
Publication Date
2026-09-18
Estimated Expiration
2043-08-15

AI Technical Summary

Technical Problem

然而,现有计算技术生成的声学响应往往与真实场景中的声学响应存在一定差异,这使得用户体验到空间音频信号品质较低、临场感较弱

Benefits of technology

[0077] The processor is used to execute the audio processing method in the first aspect or any possible implementation thereof;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119497030B_ABST
    Figure CN119497030B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an audio processing method and an electronic device. The method comprises: first, obtaining a source audio signal, a corresponding sound source position of the source audio signal in a virtual scene, and a corresponding first position of a user in the virtual scene, wherein the virtual scene is obtained by modeling a real scene; obtaining a first acoustic feature according to the sound source position and the first position, wherein the first acoustic feature is obtained by adjusting a second acoustic feature according to conversion information, the second acoustic feature is obtained by acoustic feature extraction according to the first position, the virtual scene, and the sound source position, and the conversion information is used to describe a conversion relationship between acoustic features of the same position in the virtual scene and the real scene; and then, generating a spatial audio signal according to the first acoustic feature and the source audio signal. In this way, the user can experience a spatial audio signal with high quality and strong sense of presence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing, and more particularly to an audio processing method and an electronic device. Background Technology

[0002] Currently, many electronic devices (such as mobile phones, augmented reality (AR) devices, virtual reality (VR) devices, etc.) have spatial audio capabilities, which can render source audio signals into spatial audio signals, enabling users to perceive the location, distance, and spatial sense of sound images in the audio, thus bringing users an immersive listening experience.

[0003] Typically, rendering a source audio signal into a spatial audio signal involves rendering the source audio signal using the acoustic response of the target space. However, the acoustic response generated by existing computational techniques often differs from the acoustic response in a real-world scene, resulting in users experiencing lower quality spatial audio signals and a weaker sense of presence. Summary of the Invention

[0004] To address the aforementioned technical problems, this application provides an audio processing method and an electronic device that enables users to experience high-quality spatial audio signals with a strong sense of presence.

[0005] It should be noted that the application scenarios of this application may include: scenarios where users use AR / VR devices to experience AR / VR projects (such as AR / VR science lectures, AR / VR cinemas, AR / VR concerts, etc.), and scenarios where users use terminal devices such as mobile phones, tablets, laptops, personal computers, and smartwatches with headphones to experience audio projects (such as crosstalk, movies, concerts, music concerts, etc.).

[0006] In a first aspect, embodiments of this application provide an audio processing method, the method comprising: firstly, acquiring a source audio signal, a sound source position corresponding to the source audio signal in a virtual scene, and a first position corresponding to the user in the virtual scene, wherein the virtual scene is obtained by modeling a real scene; acquiring a first acoustic feature based on the sound source position and the first position, the first acoustic feature being obtained by adjusting a second acoustic feature based on conversion information, the second acoustic feature being obtained by acoustic feature extraction based on the first position, the virtual scene, and the sound source position, the conversion information being used to describe the conversion relationship between acoustic features at the same position in the virtual scene and the real scene; and then generating a spatial audio signal based on the first acoustic feature and the source audio signal.

[0007] In other words, this application adjusts the acoustic features (i.e., the second acoustic feature) of the user's current first position in the virtual scene based on the conversion relationship between acoustic features at the same location in the virtual scene and the real scene. This makes the acoustic features (i.e., the first acoustic feature) of the user's first position in the virtual scene closer to the acoustic features of the user's position in the real scene corresponding to the virtual scene. As a result, the spatial audio signal subsequently generated based on the adjusted acoustic features (i.e., the first acoustic feature) is closer to the spatial audio signal heard by the user at the position in the real scene corresponding to the virtual scene, enabling the user to experience a high-quality spatial audio signal with a strong sense of presence.

[0008] It should be noted that the conversion information can be used to adjust the second acoustic feature at any first location in the virtual environment to obtain the corresponding first acoustic feature.

[0009] For example, the virtual scene in which the user is located can refer to the virtual scene corresponding to the audio project or AR / VR project that the user selects (or experiences). For instance, if the audio project is a concert, the corresponding virtual scene is a virtual stadium or virtual studio; if the audio project is a movie, the corresponding virtual scene is a virtual cinema; if the AR / VR project is a concert, the corresponding virtual scene is a virtual concert hall; if the AR / VR project is a lecture, the corresponding virtual scene is a virtual lecture hall, and so on. Alternatively, the virtual scene in which the user is located can refer to the virtual scene corresponding to the audio signal in the user's experience space.

[0010] For example, virtual scenes are obtained by modeling real-world scenes. The methods for generating virtual scenes can be varied, and this application does not limit this. Real-world scenes refer to objectively existing scenes, such as cinemas, concert halls, and lecture halls; virtual scenes are scenes constructed using virtual technology and are not objectively existing scenes. The virtual scene corresponding to the listening project (or AR / VR project) selected (or experienced) by the user is obtained by modeling a real-world scene. For example, a virtual concert hall is obtained by modeling a real concert hall; another example is a virtual cinema, which is obtained by modeling a real cinema; yet another example is a virtual lecture hall, which is obtained by modeling a real lecture hall; and so on.

[0011] It should be noted that the real-world scene in which the user is currently located may or may not be the same as the virtual scene corresponding to the audio project (or AR / VR project) selected (or experienced) by the user; this application does not impose any restrictions on this. Alternatively, the real-world scene in which the user is currently located may or may not be the same as the real-world scene used to model and obtain the virtual scene corresponding to the audio project (or AR / VR project) selected (or experienced) by the user; this application does not impose any restrictions on this.

[0012] For example, when the real scene where the user is currently in is the same as the virtual scene corresponding to the audio project (or AR / VR project) selected (or experienced) by the user, then the user's position in the real scene is the same as the corresponding position in the virtual scene, that is, the user's position in the real scene is the first position.

[0013] For example, when the user's current real-world scene is not the same as the virtual scene corresponding to the audio project (or AR / VR project) selected (or experienced) by the user: In one possible approach, a default position can be determined as the user's first position in the virtual scene; the default position can be set according to needs, for example, it could be the position with the best audio experience in the virtual scene, and this application does not limit this. In another possible approach, the terminal device can provide position options in the virtual scene, allowing the user to select the position option corresponding to the desired experience location in the virtual scene; subsequently, the terminal device can use the position corresponding to the user's selected position option as the user's first position in the virtual scene.

[0014] For example, the modeling methods can include a variety of methods, such as manual modeling, modeling based on visual information, modeling based on auditory information, etc., and this application does not limit them.

[0015] For example, spatial audio processing (such as rendering) can be performed based on the first acoustic features and the source audio signal to generate a spatial audio signal (such as a binaural rendering signal).

[0016] For example, the source audio signal may include audio files (such as music files, crosstalk files, etc.) or audio files contained in multimedia files (such as audio files contained in movie files, etc.).

[0017] For example, the first acoustic feature can be pre-generated and stored in a database, or it can be generated in real time; this application does not limit this.

[0018] For example, the second acoustic feature can be pre-generated and stored in a database, or it can be generated in real time; this application does not limit this.

[0019] For example, the conversion information can be pre-generated and stored in a database, or it can be generated in real time; this application does not limit this.

[0020] It should be noted that the first acoustic feature, the second acoustic feature, and the conversion information can be stored in the same database or in different databases.

[0021] For example, acoustic feature extraction methods may include, but are not limited to, acoustic simulation, machine learning, numerical calculation, etc., and this application does not limit them.

[0022] According to the first aspect, the conversion information is determined by analyzing a set of acoustic features, which includes a third acoustic feature and a fourth acoustic feature. The fourth acoustic feature is the acoustic feature at the second position in the real scene, and the third acoustic feature is the acoustic feature at the second position in the virtual scene. Thus, by analyzing the acoustic features at the same position in both the real and virtual scenes, the accurate conversion relationship between the acoustic features at the same position in the real and virtual scenes can be determined.

[0023] According to the first aspect, or any implementation of the first aspect above, there are one or more acoustic feature groups, one or more second positions, and the second positions corresponding to the third acoustic feature and the fourth acoustic feature belonging to the same acoustic feature group are the same.

[0024] For example, when the real-world scenario is a uniform room (e.g., a square room with uniform wall materials, where the room can be understood as an indoor scene), and the differences in the actual acoustic features at different locations are relatively small, there can be only one first location, thus reducing workload. Conversely, when the real-world scenario is a non-uniform room (e.g., an asymmetrical room with diverse wall materials), and the differences in the actual acoustic features at different locations are relatively large, there can be multiple second locations, thus increasing the accuracy of the generated conversion information. It should be understood that the number of second locations can be determined according to requirements, and this application does not impose any restrictions on this.

[0025] For example, when sound source devices are deployed at different locations (hereinafter referred to as preset sound source locations) in a real-world scenario, the fourth acoustic feature at the same location is different; correspondingly, when virtual sound sources are deployed at different preset sound source locations in a virtual scenario, the third acoustic feature at the same location is different. That is, in a real-world scenario, each second location corresponds to one or more fourth acoustic features; in a virtual scenario, each second location corresponds to one or more third acoustic features. Therefore, a second location can correspond to one or more acoustic feature groups, and the second locations corresponding to the third and fourth acoustic features belonging to the same acoustic feature group are the same, and the corresponding preset sound source locations are also the same.

[0026] It should be noted that this application only needs to analyze the conversion relationship between the acoustic features of one or more locations (i.e., one or more second locations) in the virtual scene and the real scene to obtain conversion information; and the conversion information has universal applicability in any location in the same room (i.e. the same virtual scene), rather than being used only to adjust the acoustic features of the first location that is close to the second location.

[0027] According to the first aspect, or any implementation of the first aspect above, the conversion information includes a conversion function, which is obtained by analyzing the signal processing results of the third acoustic feature and the fourth acoustic feature.

[0028] For example, signal processing can be performed on the third and fourth acoustic features in an acoustic feature group to obtain signal processing results. Then, the signal processing results of the third and fourth acoustic features are analyzed to obtain a transfer function. This transfer function can then be used as transfer information. For instance, for the third acoustic feature including a third acoustic response such as the Room Impulse Response (RIR), and the fourth acoustic feature including a fourth acoustic response such as RIR, where the third acoustic feature is represented by RIR 1 and the fourth acoustic feature by RIR 2, frequency domain transformation can be performed on RIR 1 and RIR 2 respectively to obtain the frequency domain response of RIR 1 (called frequency response 1) and the frequency domain response of RIR 2 (called frequency response 2). Next, in one possible approach, the frequency response transfer function can be calculated using frequency response 1 and frequency response 2, and this frequency response transfer function can be used as the transfer function. In another possible approach, signal analysis can be performed on frequency response 1 and frequency response 2 separately, and then the transfer function can be determined based on the signal analysis results.

[0029] According to the first aspect, or any implementation of the first aspect above, the conversion information includes the feature change rate, which is the rate of change of the third acoustic feature relative to the fourth acoustic feature.

[0030] For example, numerical analysis can be performed on the third and fourth acoustic features in an acoustic feature group to obtain the feature change rate; in this case, the feature change rate can be used as conversion information. Specifically, the change rate of the third acoustic feature in the acoustic feature group relative to the fourth acoustic feature in the acoustic feature group can be calculated as the feature change rate. For example, for the third acoustic parameter such as direct-mixing ratio included in the third acoustic feature, and the fourth acoustic parameter such as direct-mixing ratio included in the fourth acoustic feature, where the direct-mixing ratio in the third acoustic parameter is represented by DRR 1 and the direct-mixing ratio in the fourth acoustic parameter is represented by DRR 2; the corresponding feature change rate can be (DRR 1-DRR 2) / DRR2, or DRR1 / DRR2. It should be understood that the conversion information can include multiple feature change rates (multiple feature change rates correspond one-to-one with multiple acoustic parameters), and this application does not limit this.

[0031] According to the first aspect, or any implementation of the first aspect above, the transformation information includes the model output information obtained by inputting the third acoustic feature and the fourth acoustic feature into the model.

[0032] For example, machine learning can be used to process the third and fourth acoustic features in an acoustic feature set to obtain conversion information. Specifically, the third and fourth acoustic features in an acoustic feature set can be input into an AI model (or machine learning model, hereinafter referred to as the second model). The second model processes the third and fourth acoustic features in the acoustic feature set and outputs conversion information; that is, the output information of the second model is the conversion information.

[0033] For example, the model output information may include transformation functions, feature change rates, and other forms of information, which are not limited in this application.

[0034] It should be understood that other methods can also be used to generate conversion information, and this application does not impose any restrictions on this.

[0035] According to the first aspect, or any implementation of the first aspect above, the first acoustic feature includes at least one of the following: acoustic response, energy information, acoustic parameters, or acoustic features.

[0036] Acoustic responses may include, but are not limited to: room impulse response, binaural room impulse response (BRIR), higher order ambisonics (HOA), etc.

[0037] Energy information can include energy distribution in multiple frequency bands and directions corresponding to the acoustic response, energy attenuation, etc.

[0038] Acoustic parameters can include environmental acoustic parameters and binaural acoustic parameters. Environmental acoustic parameters include, but are not limited to: Direct-to-Reverberation energy ratio (DRR), reverberation time (e.g., T20 (time to decay 20dB of energy), T30, T60, etc.), and clarity (e.g., C50 (sound energy ratio before and after 50ms), C80, etc.). Binaural acoustic parameters can include, but are not limited to: Interaural Time Difference (ITD), Interaural Level Difference (ILD), and Interaural Cross Correlation (IACC), etc.

[0039] The acoustic features may include numerical analysis features obtained by performing numerical analysis on acoustic parameters (e.g., principal component analysis (PCA) analysis); and may also include the output information of the first model obtained by inputting at least one of acoustic response, energy information or acoustic parameters into the AI ​​model (hereinafter referred to as the first model).

[0040] According to the first aspect, or any implementation of the first aspect above, the fourth acoustic feature is obtained by processing the test audio signal received at the second position in a real scene.

[0041] For example, a sound source device can be deployed in a real-world scenario. After the user moves to the second location with the receiving device, the sound source device can be controlled to play a test audio signal. Correspondingly, the receiving device at the second location can receive the test audio signal. Then, the receiving device can process the received test audio signal to obtain a fourth acoustic feature. It should be understood that the fourth acoustic feature is a real acoustic feature.

[0042] For example, the sound source device plays test audio signals such as pulse signals, noise signals, and frequency sweep signals, and the receiving device in the second position receives and processes the test audio signals to obtain acoustic responses in formats such as RIR, BRIR, and HOA.

[0043] For example, a sound source device plays sound signals such as pulse signals, noise signals, and frequency sweep signals, and a receiving device in a second position receives and processes the test audio signals to obtain energy information, acoustic parameters, and acoustic characteristics.

[0044] For example, a sound source device plays natural audio signals (such as voice, songs, instrumental music, videos, etc.), and a receiving device in a second position receives and processes the test audio signal to obtain acoustic responses in formats such as RIR, BRIR, and HOA.

[0045] For example, a sound source device plays natural audio signals (such as voice, songs, instrumental music, videos, etc.), and a receiving device in a second position receives and processes the test audio signals to obtain energy information, acoustic parameters, and acoustic characteristics.

[0046] Among them, sound source devices include, but are not limited to: speaker devices (such as home audio systems), terminal devices with external playback functions (such as tablets, large screens, etc.), professional acoustic measurement equipment, etc.

[0047] The receiving devices include, but are not limited to: terminal devices including microphones (such as tablets, mobile phones, AR / VR devices, etc.), professional acoustic measurement equipment, etc.

[0048] According to the first aspect, or any of the above implementations of the first aspect, the fourth acoustic feature is a user-defined input.

[0049] Since the acoustic features obtained from actual measurements in some real-world scenarios may not be the most accurate match for the acoustic features that users actually hear, we directly input acoustic features that match the users' actual hearing in some real-world scenarios. This makes the adjusted acoustic features closer to the acoustic features that match the users' actual hearing in some real-world scenarios. In this way, the spatial audio signal generated subsequently based on the adjusted acoustic features is closer to the spatial audio signal that the user hears in the real-world scenario corresponding to some virtual scenarios.

[0050] For example, "certain real-world scenarios" may include: real-world scenarios with unavoidable interference noise, real-world scenarios where the visual representation of the surface material does not match its actual acoustic effects, and real-world scenarios that serve as multi-functional spaces (e.g., real-world scenarios that serve as lecture halls, concert halls, and film screening rooms).

[0051] According to the first aspect, or any implementation of the first aspect above, the first acoustic feature is obtained based on the sound source location and the first location, including: determining one or more acoustic feature groups, one acoustic feature group including a third acoustic feature and a fourth acoustic feature, the fourth acoustic feature being the acoustic feature of the second location in the real scene, and the third acoustic feature being the acoustic feature of the second location in the virtual scene; analyzing the third acoustic feature and the fourth acoustic feature in one or more acoustic feature groups to obtain conversion information; extracting acoustic features based on the first location, the virtual scene, and the sound source location to obtain a second acoustic feature; and adjusting the second acoustic feature based on the conversion information to obtain the first acoustic feature.

[0052] In this case, the real scene where the user is currently in is the same as the virtual scene corresponding to the audio project (or AR / VR project) selected (or experienced) by the user, and thus conversion information, second acoustic features and first acoustic features can be generated in real time.

[0053] According to the first aspect, or any implementation of the first aspect above, obtaining the first acoustic feature based on the sound source location and the first location includes: selecting a second acoustic feature set corresponding to a virtual scene from multiple second acoustic feature sets stored in a database; wherein, the multiple second acoustic feature sets correspond one-to-one with multiple preset virtual scenes, a second acoustic feature set includes multiple second preset acoustic features, and a second acoustic feature set is obtained by acoustic feature extraction based on a preset virtual scene, multiple fourth locations, and multiple preset sound source locations, and a fourth location and a preset sound source location are used to determine a second preset acoustic feature in a second acoustic feature set; selecting the second acoustic feature from the second acoustic feature set corresponding to the virtual scene based on the sound source location and the first location; selecting conversion information from multiple preset conversion information stored in the database based on the virtual scene, the multiple preset conversion information corresponding one-to-one with multiple preset virtual scenes; and adjusting the second acoustic feature based on the conversion information to obtain the first acoustic feature.

[0054] In this scenario, the user's current real-world environment and the virtual environment corresponding to the audio project (or AR / VR project) selected (or experienced) by the user may or may not be the same. This eliminates the need for the terminal device to generate conversion information and second acoustic features in real time, reducing the time required to generate spatial audio signals and improving the user experience. Furthermore, it allows for a more efficient way to invoke virtual environments that may be used multiple times at different times, by different users, and in different applications.

[0055] According to the first aspect, or any implementation of the first aspect above, obtaining the first acoustic feature based on the sound source location and the first location includes: selecting a first acoustic feature set corresponding to a virtual scene from multiple first acoustic feature sets stored in a database; wherein, the multiple first acoustic feature sets correspond one-to-one with multiple preset virtual scenes, the multiple first acoustic feature sets correspond one-to-one with multiple second acoustic feature sets, a second acoustic feature set is obtained by acoustic feature extraction based on a preset virtual scene, multiple third locations, and multiple preset sound source locations, a third location and a preset sound source location are used to determine a second preset acoustic feature in a second acoustic feature set, and a first preset acoustic feature in a first acoustic feature set is obtained by adjusting a second preset acoustic feature in the corresponding second acoustic feature set according to conversion information; selecting the first acoustic feature from the first acoustic feature set corresponding to the virtual scene based on the sound source location and the first location.

[0056] In this scenario, the user's current real-world environment and the virtual environment corresponding to the audio project (or AR / VR project) selected (or experienced) by the user may or may not be the same. This eliminates the need for the terminal device to generate conversion information, second acoustic features, and first acoustic features in real time, further reducing the time required to generate spatial audio signals and improving the user experience. Additionally, it allows for a more efficient way to invoke virtual environments that may be used multiple times at different times, by different users, and in different applications.

[0057] Secondly, embodiments of this application provide an audio processing apparatus, the apparatus comprising:

[0058] The first acquisition module is used to acquire the source audio signal, the sound source position of the source audio signal in the virtual scene, and the first position of the user in the virtual scene. The virtual scene is obtained by modeling the real scene.

[0059] The second acquisition module is used to acquire a first acoustic feature based on the sound source location and the first location. The first acoustic feature is obtained by adjusting the second acoustic feature based on the conversion information. The second acoustic feature is obtained by extracting acoustic features based on the first location, the virtual scene, and the sound source location. The conversion information is used to describe the conversion relationship between acoustic features at the same location in the virtual scene and the real scene.

[0060] The audio signal generation module is used to generate a spatial audio signal based on the first acoustic features and the source audio signal.

[0061] It should be understood that the audio processing device of the second aspect can execute any of the implementations of the first aspect, which will not be elaborated here.

[0062] The second aspect and any implementation thereof correspond to the first aspect and any implementation thereof, respectively. The technical effects of the second aspect and any implementation thereof are similar to those of the first aspect and any implementation thereof, and will not be repeated here.

[0063] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor, the memory being coupled to the processor; the memory storing program instructions, when executed by the processor, causing the electronic device to perform the audio processing method in the first aspect or any possible implementation of the first aspect.

[0064] The third aspect and any implementation thereof correspond to the first aspect and any implementation thereof, respectively. The technical effects of the third aspect and any implementation thereof can be found in the technical effects of the first aspect and any implementation thereof, as described above, and will not be repeated here.

[0065] Fourthly, embodiments of this application provide a chip including one or more interface circuits and one or more processors; the one or more processors receive or send data through the one or more interface circuits, and when the one or more processors execute computer instructions, the steps of the audio processing method in the first aspect or any possible implementation of the first aspect are executed.

[0066] The fourth aspect and any implementation thereof correspond to the first aspect and any implementation thereof, respectively. The technical effects of the fourth aspect and any implementation thereof are similar to those of the first aspect and any implementation thereof, and will not be repeated here.

[0067] Fifthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when run on a computer or processor, causes the computer or processor to perform the audio processing method in the first aspect or any possible implementation thereof.

[0068] The fifth aspect and any implementation thereof correspond to the first aspect and any implementation thereof, respectively. The technical effects of the fifth aspect and any implementation thereof are similar to those of the first aspect and any implementation thereof, and will not be repeated here.

[0069] In a sixth aspect, embodiments of this application provide a computer program product, which includes computer instructions that, when executed by a computer or processor, cause the computer or processor to perform the audio processing method in the first aspect or any possible implementation thereof.

[0070] The sixth aspect and any implementation thereof correspond to the first aspect and any implementation thereof, respectively. The technical effects of the sixth aspect and any implementation thereof are similar to those of the first aspect and any implementation thereof, and will not be repeated here.

[0071] Seventhly, embodiments of this application provide an augmented reality (AR) device, which includes: a display module, an image acquisition module, earphones, and a processor, wherein:

[0072] The processor is used to execute the audio processing method in the first aspect or any possible implementation thereof;

[0073] The headphones are used to play spatial audio signals generated by the audio processing method in the first aspect or any possible implementation of the first aspect.

[0074] For example, the display module can be used to display images; the image acquisition module is used to acquire images.

[0075] The seventh aspect and any implementation thereof correspond to the first aspect and any implementation thereof, respectively. The technical effects of the seventh aspect and any implementation thereof are similar to those of the first aspect and any implementation thereof, and will not be repeated here.

[0076] Eighthly, embodiments of this application provide an augmented reality (VR) device, which includes: a display module, an image acquisition module, headphones, and a processor, wherein:

[0077] The processor is used to execute the audio processing method in the first aspect or any possible implementation thereof;

[0078] The headphones are used to play spatial audio signals generated by the audio processing method in the first aspect or any possible implementation of the first aspect.

[0079] For example, the display module can be used to display images; the image acquisition module is used to acquire images.

[0080] The eighth aspect and any implementation thereof correspond to the first aspect and any implementation thereof, respectively. The technical effects corresponding to the eighth aspect and any implementation thereof are similar to those corresponding to the first aspect and any implementation thereof, and will not be repeated here. Attached Figure Description

[0081] Figure 1A This is a schematic diagram illustrating an application scenario;

[0082] Figure 1B This is a schematic diagram illustrating an application scenario;

[0083] Figure 1C This is a schematic diagram illustrating an application scenario;

[0084] Figure 1D This is a schematic diagram illustrating an application scenario;

[0085] Figure 2This is a schematic diagram illustrating an audio processing procedure as an example.

[0086] Figure 3 This is a schematic diagram illustrating an audio processing procedure as an example.

[0087] Figure 4 This is a schematic diagram illustrating an audio processing procedure as an example.

[0088] Figure 5 This is a schematic diagram illustrating an audio processing procedure as an example.

[0089] Figure 6A A schematic diagram of the frequency response curve of an acoustic feature as an example;

[0090] Figure 6B A schematic diagram of the frequency response curve of an acoustic feature as an example;

[0091] Figure 7 This is a schematic diagram of an audio processing device as an example.

[0092] Figure 8 This is a schematic diagram of the structure of an exemplary device. Detailed Implementation

[0093] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0094] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.

[0095] The terms "first" and "second," etc., used in the specification and claims of this application are used to distinguish different objects, not to describe a specific order of objects. For example, "first target object" and "second target object," etc., are used to distinguish different target objects, not to describe a specific order of target objects.

[0096] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0097] In the description of the embodiments in this application, unless otherwise stated, "multiple" means two or more. For example, multiple processing units means two or more processing units; multiple systems means two or more systems.

[0098] Figure 1A This is a schematic diagram illustrating an exemplary application scenario. Figure 1A The image illustrates an application scenario where a user uses an AR device to experience an AR concert. It should be understood that users can also use AR devices to experience other AR projects such as AR science lectures, AR cinemas, etc., and this application does not impose any limitations on this.

[0099] in, Figure 1A The real-world scene where the user is located and the virtual scene corresponding to the AR project of the user experience are the same scene. That is, the real-world scene where the user is located is a concert hall, and the virtual scene corresponding to the AR project of the user experience is a virtual concert hall.

[0100] Reference Figure 1A For example, when a user wishes to experience an AR concert in an empty concert hall, they can wear an AR device. The user can then select the AR concert hall as the desired AR experience within the device, and choose the music file they wish to play. The AR device can then process the selected music file, such as rendering it, to obtain a spatial audio signal, which is then played through headphones. Furthermore, the AR device can display the performance on a screen, allowing the user to see the performers playing from their corresponding positions on the concert hall stage. In this way, by combining visual and auditory elements, the user can experience an immersive concert experience.

[0101] It should be understood that, Figure 1A In the context of a concert hall, there may be one or more users experiencing an AR concert using AR devices, and this application does not impose any restrictions on this.

[0102] Figure 1B This is a schematic diagram illustrating an exemplary application scenario. Figure 1B The image illustrates an application scenario where a user uses VR devices to experience a VR cinema. It should be understood that users can also use VR devices to experience other VR projects such as VR concert halls, VR science lectures, etc., and this application does not impose any restrictions on this.

[0103] in, Figure 1B The real-world scenario where the user is located is not the same as the virtual scenario corresponding to the VR project they are experiencing. That is, the real-world scenario where the user is located is the living room, while the virtual scenario corresponding to the VR project they are experiencing is a virtual movie theater.

[0104] Reference Figure 1BFor example, when a user wants to experience a VR cinema in their living room, they can wear a VR device. Then, the user can select the VR cinema as the desired VR experience within the VR device, choose the desired movie file, and play it. Subsequently, the VR device can process the audio files contained in the movie file, such as rendering, to obtain spatial audio signals, which are then played through headphones. The VR device can also display the movie images on a screen. In this way, through the combination of visual and auditory experiences, the user can have an immersive movie experience.

[0105] Figure 1C This is a schematic diagram illustrating an exemplary application scenario. Figure 1C The image shows an application scenario where users experience a concert using (mobile phone + headphones).

[0106] It should be understood that users can also use (mobile phone + headphones) to experience other audio-visual content such as crosstalk, movies, concerts, etc., and this application does not limit this. Furthermore, the mobile phone is merely one example of this application; this application can also use tablets, laptops, personal computers, smartwatches, and other terminal devices + headphones for audio-visual experiences.

[0107] in, Figure 1C The real-world scenario where the user is located is the same as the virtual scenario corresponding to the listening project of the user experience. That is, the real-world scenario where the user is located is a concert hall, and the virtual scenario corresponding to the listening project of the user experience is a virtual concert hall.

[0108] Reference Figure 1C For example, when a user wishes to experience a concert in an empty concert hall, they can wear headphones and connect them to their mobile phone. The user can then select the concert as the desired audio experience on their phone, choose the desired music file, and play it. The phone can then process the selected music file, such as rendering it, to obtain a spatial audio signal. This spatial audio signal can then be sent to the headphones for playback. The phone can also play video footage of the concert performance. In this way, the user can immerse themselves in a concert using their mobile phone in a concert hall.

[0109] It should be understood that, Figure 1C There can be one or more users using their mobile phones to experience a concert in a concert hall, and this application does not impose any restrictions on this.

[0110] Figure 1D This is a schematic diagram illustrating an exemplary application scenario. Figure 1D The image shows an application scenario where users experience crosstalk using (mobile phone + headphones).

[0111] It should be understood that users can also use (mobile phone + headphones) to experience other audio-visual experiences such as concerts, movies, and live performances, and this application does not limit this. Furthermore, the mobile phone is merely one example of this application; this application can also utilize tablets, laptops, personal computers, smartwatches, and other terminal devices combined with headphones for audio listening.

[0112] in, Figure 1D The real-world scenario where the user is located is not the same as the virtual scenario corresponding to the listening project of the user experience. That is, the real-world scenario where the user is located is a bedroom, while the virtual scenario corresponding to the listening project of the user experience is a virtual studio.

[0113] Reference Figure 1D For example, when a user wants to experience crosstalk in their bedroom, they can wear headphones and connect them to their mobile phone. Then, the user can select crosstalk as the desired audio experience on their phone, choose the desired crosstalk file, and play it. The phone can then process the selected crosstalk file, such as rendering it, to obtain a spatial audio signal. This spatial audio signal can then be sent to the headphones for playback. The phone can also play the crosstalk video. In this way, the user can immerse themselves in crosstalk using their mobile phone in their bedroom.

[0114] It should be understood that this application can also be applied to other listening scenarios or other virtual scenarios, and this application does not limit it.

[0115] The process of generating spatial audio signals is explained below.

[0116] Figure 2 This is a schematic diagram illustrating an audio processing procedure as an example. Figure 2 The steps in the embodiments can be executed by terminal devices such as AR devices, VR devices, mobile phones, tablets, laptops, personal computers, and smartwatches.

[0117] S201, obtain the source audio signal, the sound source position of the source audio signal in the virtual scene, and the first position of the user in the virtual scene, wherein the virtual scene is obtained by modeling the real scene.

[0118] For example, when a user needs to listen to audio, they can wear headphones and connect their phone to the headphones; then, they can select the desired audio item and the audio file they wish to experience from the terminal device such as a mobile phone, tablet, laptop, personal computer, or smartwatch; and then, they can perform the playback operation. When a user needs to experience AR / VR projects, they can wear AR / VR devices, then select the desired AR / VR project and the multimedia file they wish to experience; and then, they can perform the playback operation.

[0119] For example, after the user performs a playback operation, the terminal device can respond to the user's operation by obtaining the audio file selected by the user (or the audio file contained in the multimedia file selected by the user), that is, obtaining the source audio signal; and the terminal device can also obtain the first position of the user in the virtual scene and the sound source position of the source audio signal in the virtual scene, so as to generate a spatial audio signal that matches the first position and the sound source position in the subsequent generation.

[0120] For example, the virtual scene in S201 where the user is located can refer to the virtual scene corresponding to the audio project or AR / VR project selected (or experienced) by the user. For instance, if the audio project is a concert, the corresponding virtual scene is a virtual stadium or virtual studio; if the audio project is a movie, the corresponding virtual scene is a virtual cinema; if the AR / VR project is a concert, the corresponding virtual scene is a virtual concert hall; if the AR / VR project is a lecture, the corresponding virtual scene is a virtual lecture hall, and so on. Alternatively, the virtual scene in S201 where the user is located can refer to the virtual scene corresponding to the audio signal in the user experience space.

[0121] For example, virtual scenes are obtained by modeling real-world scenes; the methods for generating virtual scenes can include various types, which are not limited in this application and will be described in detail later. Real-world scenes refer to objectively existing scenes, such as cinemas, concert halls, and lecture halls; virtual scenes are scenes constructed using virtual technology and are not objectively existing scenes. The virtual scene corresponding to the listening project (or AR / VR project) selected (or experienced) by the user is obtained by modeling real-world scenes. For example, a virtual concert hall is obtained by modeling a real concert hall; another example is a virtual cinema, which is obtained by modeling a real cinema; yet another example is a virtual lecture hall, which is obtained by modeling a real lecture hall; and so on.

[0122] It should be noted that the real-world scene in which the user is currently located and the virtual scene corresponding to the audio project (or AR / VR project) selected (or experienced) by the user can be the same scene, as mentioned above. Figure 1A and Figure 1C As shown in the examples, the scenarios described above may also differ. Figure 1B and Figure 1D As shown in the embodiments, this application does not impose any limitations on this. In other words, the real-world scene where the user is currently located, and the real-world scene corresponding to the virtual scene corresponding to the audio project (or AR / VR project) selected (or experienced) by the user, used for modeling, can be the same scene as described above. Figure 1A and Figure 1C As shown in the examples, the scenarios described above may also differ. Figure 1B and Figure 1D As shown in the examples, this application does not impose any limitations on this.

[0123] For example, when the real scene where the user is currently in is the same as the virtual scene corresponding to the audio project (or AR / VR project) selected (or experienced) by the user, then the user's position in the real scene is the same as the corresponding position in the virtual scene, that is, the user's position in the real scene is the first position.

[0124] For example, when the user's current real-world scene is not the same as the virtual scene corresponding to the audio project (or AR / VR project) selected (or experienced) by the user: In one possible approach, a default position can be determined as the user's first position in the virtual scene; the default position can be set according to needs, for example, it could be the position with the best audio experience in the virtual scene, and this application does not limit this. In another possible approach, the terminal device can provide position options in the virtual scene, so that the user can select the position option corresponding to the desired experience position in the virtual scene before performing playback operations; subsequently, the terminal device can use the position option selected by the user as the user's first position in the virtual scene.

[0125] S202, based on the sound source location and the first location, obtain the first acoustic feature. The first acoustic feature is obtained by adjusting the second acoustic feature according to the conversion information. The second acoustic feature is obtained by extracting acoustic features based on the first location, the virtual scene, and the sound source location. The conversion information is used to describe the conversion relationship between acoustic features at the same location in the virtual scene and the real scene.

[0126] Next, a first acoustic feature can be obtained, which can be used for spatial audio processing to generate a spatial audio signal.

[0127] For example, the first acoustic feature includes, but is not limited to, acoustic response, energy information, acoustic parameters or acoustic characteristics, etc., and this application does not limit it.

[0128] Acoustic responses may include, but are not limited to: Room Impulse Response (RIR), Binaural Room Impulse Response (BRIR), Higher Order Ambisonics (HOA), etc.

[0129] Energy information can include energy distribution in multiple frequency bands and directions corresponding to the acoustic response, energy attenuation, etc.

[0130] Acoustic parameters can include environmental acoustic parameters and binaural acoustic parameters. Environmental acoustic parameters include, but are not limited to: Direct-to-Reverberation energy ratio (DRR), reverberation time (e.g., T20 (time to decay 20dB of energy), T30, T60, etc.), and clarity (e.g., C50 (sound energy ratio before and after 50ms), C80, etc.). Binaural acoustic parameters can include, but are not limited to: Interaural Time Difference (ITD), Interaural Level Difference (ILD), and Interaural Cross Correlation (IACC), etc.

[0131] The acoustic features may include numerical analysis features obtained by performing numerical analysis on acoustic parameters (e.g., principal component analysis (PCA)); and may also include the output information of the first model obtained by inputting at least one of acoustic response, energy information or acoustic parameters into the AI ​​model (hereinafter referred to as the first model).

[0132] For ease of explanation, the acoustic response contained in the first acoustic feature is referred to as the first acoustic response; the energy information contained in the first acoustic feature is referred to as the first energy information; the acoustic parameters contained in the first acoustic feature are referred to as the first acoustic parameters; and the acoustic features contained in the first acoustic feature are referred to as the first acoustic feature.

[0133] In S202, one possible approach is to first acquire conversion information, which can be used to describe the conversion relationship between acoustic features at the same location in the virtual scene and the real scene. It should be noted that the conversion information can be pre-generated and stored in a database, or it can be generated in real time; this application does not impose any restrictions on this. The generation method of the conversion information will be explained later. Next, a second acoustic feature is acquired; the second acoustic feature is the acoustic feature of the first location in the virtual scene, which can be obtained by acoustic feature extraction based on the first location, the virtual scene, and the sound source location. The acoustic feature extraction method can include, but is not limited to, acoustic simulation, machine learning, numerical calculation, etc.; this application does not impose any restrictions on this. It should be noted that the second acoustic feature can be pre-generated and stored in a database, or it can be generated in real time; this application does not impose any restrictions on this. Afterwards, the second acoustic feature can be adjusted according to the conversion information to obtain the first acoustic feature. In this way, the acoustic feature of the first location in the virtual scene can be adjusted to be closer to the acoustic feature of the first location in the real scene, thereby improving the quality and presence of the spatial audio signal and thus enhancing the user experience. In this context, the acoustic features at any location in the virtual scene can be called virtual acoustic features, while the acoustic features at any location in the real scene can be called real acoustic features; that is, the first acoustic feature and the second acoustic feature are virtual acoustic features.

[0134] In S202, one possible approach is to adjust the second preset acoustic features (one or more second preset acoustic features per location, and multiple second preset acoustic features per location correspond to multiple virtual sound source locations) of each location in various preset virtual scenes in advance based on each preset conversion information (one preset conversion information corresponds to one preset virtual scene), to obtain the first preset acoustic features of each location in various virtual scenes and store them in the database; in this way, the first acoustic features can be directly obtained from the database based on the virtual scene, the sound source location, and the first location.

[0135] It should be understood that the second acoustic feature may include, but is not limited to: second acoustic response, second energy information, second acoustic parameter, or second acoustic feature.

[0136] It should be noted that the conversion information, the second acoustic feature, and the first acoustic feature can be stored in the same database or in different databases; this application does not impose any restrictions on this.

[0137] S203 generates a spatial audio signal based on the first acoustic characteristics and the source audio signal.

[0138] Then, spatial audio processing (such as rendering) can be performed based on the first acoustic features and the source audio signal to obtain spatial audio signals (such as binaural rendering signals).

[0139] Then, AR / VR devices can play the spatial audio signal through their built-in headphones; mobile phones, tablets, laptops, smartwatches and other terminal devices can play the spatial audio signal through headphones connected to them; in this way, users can hear the spatial audio signal through headphones.

[0140] The following explains the process of generating the conversion information and the process of adjusting the second acoustic feature to obtain the first acoustic feature.

[0141] Figure 3 This is a schematic diagram illustrating an audio processing procedure as an example. Figure 3 In this embodiment, the real-world scene in which the user is currently located is the same as the virtual scene corresponding to the audio project (or AR / VR project) selected (or experienced) by the user (i.e., the application scene can be as described above). Figure 1A and Figure 1C In this case, the conversion information, the second acoustic feature, and the first acoustic feature can all be generated in real time.

[0142] S301, based on real-world scenarios, constructs virtual scenarios.

[0143] One possible approach is manual modeling. For example, users can input information such as the size, position, and surface material of objects in a real-world scene into modeling software, which then creates a 3D model based on the user's input, resulting in a virtual scene.

[0144] One possible approach is to model based on visual information. For example, a user wearing an AR / VR device can capture a sequence of images that cover a real-world scene by moving or rotating the device; the AR / VR device can then perform 3D modeling based on the image sequence to obtain a virtual scene.

[0145] It should be understood that other methods can also be used to construct virtual scenes, and this application does not impose any restrictions on this.

[0146] It should be noted that S301 can be a step executed in advance or a step executed in real time, and this application does not limit it.

[0147] S302, acquire fourth acoustic features at one or more second locations in a real scene.

[0148] In one possible approach, the fourth acoustic feature can be obtained by processing the audio signal received at a second location in a real-world scene.

[0149] For example, a sound source device can be deployed in a real-world scenario. After the user moves to the second location with the receiving device, the sound source device can be controlled to play a test audio signal. Correspondingly, the receiving device at the second location can receive the test audio signal. Then, the receiving device can process the received test audio signal to obtain a fourth acoustic feature. It should be understood that the fourth acoustic feature is a real acoustic feature.

[0150] For example, the sound source device plays test audio signals such as pulse signals, noise signals, and frequency sweep signals, and the receiving device in the second position receives and processes the test audio signals to obtain acoustic responses in formats such as RIR, BRIR, and HOA.

[0151] For example, a sound source device plays sound signals such as pulse signals, noise signals, and frequency sweep signals, and a receiving device in a second position receives and processes the test audio signals to obtain energy information, acoustic parameters, and acoustic characteristics.

[0152] For example, a sound source device plays natural audio signals (such as voice, songs, instrumental music, videos, etc.), and a receiving device in a second position receives and processes the test audio signal to obtain acoustic responses in formats such as RIR, BRIR, and HOA.

[0153] For example, a sound source device plays natural audio signals (such as voice, songs, instrumental music, videos, etc.), and a receiving device in a second position receives and processes the test audio signals to obtain energy information, acoustic parameters, and acoustic characteristics.

[0154] Among them, sound source devices include, but are not limited to: speaker devices (such as home audio systems), terminal devices with external playback functions (such as tablets, large screens, etc.), professional acoustic measurement equipment, etc.

[0155] The receiving devices include, but are not limited to: terminal devices including microphones (such as tablets, mobile phones, AR / VR devices, etc.), professional acoustic measurement equipment, etc.

[0156] The second position can be one or more. For example, when the real-world scene is a uniform room (e.g., a square room with uniform wall materials, where the room can be understood as an indoor scene) and the differences in the actual acoustic features between different positions are relatively small, the first position can be one, thus reducing workload. Conversely, when the real-world scene is a non-uniform room (e.g., an asymmetrical room with diverse wall materials) and the differences in the actual acoustic features between different positions are relatively large, the second position can be multiple, thus increasing the accuracy of the generated conversion information. It should be understood that the number of second positions can be determined according to requirements, and this application does not impose any restrictions on this.

[0157] In one possible approach, when the terminal device is an AR / VR device, the AR / VR device can display measurement instructions on its screen, which may include a second location. The user can then wear the AR / VR device and move according to the second location displayed on the device's screen. During the user's movement, the AR / VR device can perform real-time positioning to obtain location information and provide movement prompts based on this information. When the AR / VR device determines that the user has moved to the second location based on the obtained location information, it can prompt the user to stop moving. After the AR / VR device receives a test audio signal at a second location and processes it to obtain the fourth acoustic feature, it can play a test success message. If there are multiple second locations, the AR / VR device can prompt the user to move to the next second location until the AR / VR device obtains the fourth acoustic feature for all second locations in the real scene. If the AR / VR device does not receive a test audio signal or obtain the fourth acoustic feature at a second location, it can play a test failure message and prompt the user to test again, i.e., prompting the user to remain at the current second location to receive the test audio signal. In this scenario, AR / VR devices can record location information for one or more second locations to facilitate the subsequent determination of third acoustic features for one or more second locations within the virtual scene.

[0158] In one possible approach, when the terminal device is an AR / VR device, the AR / VR device can display measurement instructions on its screen, which may include the number of second locations. The user can then select a corresponding number of second locations within the AR / VR device based on the displayed number; subsequently, the user can move to the selected second locations while wearing the AR / VR device. During the user's movement, the AR / VR device can perform real-time positioning to obtain location information and provide movement prompts based on this information. When the AR / VR device determines that the user has moved to a second location based on the obtained location information, it can prompt the user to stop moving. After the AR / VR device receives a test audio signal at a second location and processes it to obtain the fourth acoustic feature, it can play a test success message; and when there are multiple second locations, the AR / VR device can prompt the user to move to the next second location, until the AR / VR device obtains the fourth acoustic feature for all second locations in the real scene. If the AR / VR device does not receive a test audio signal or obtain the fourth acoustic feature at a second location, it can play a test failure message and prompt the user to test again, i.e., prompting the user to remain at the current second location to receive the test audio signal. In this scenario, AR / VR devices can record location information for one or more second locations to facilitate the subsequent determination of third acoustic features for one or more second locations within the virtual scene.

[0159] In one possible approach, when the terminal device is a tablet, laptop, or mobile phone—devices with low positioning accuracy—a virtual scene can be displayed on the terminal device, and a second location can be marked within the virtual scene. The user then moves the terminal device to the second location marked in the virtual scene and stops. Recording is then performed on the terminal device, which can then receive a test audio signal at the second location. After receiving the test audio signal at the second location and processing it to obtain the fourth acoustic feature, a test success message can be played. If there are multiple second locations, the terminal device can prompt the user to move to the next second location, until the terminal device obtains the fourth acoustic features for all second locations in the real scene. If the terminal device does not receive a test audio signal or obtain the fourth acoustic feature at a second location, a test failure message can be played, prompting the user to test again, i.e., to remain at the current second location to receive the test audio signal. In this case, the user can input the location information of one or more second locations on the terminal device (e.g., distance from a wall, the direction of an object relative to the user, the user's height, etc.) to facilitate the subsequent determination of the third acoustic features of one or more second locations in the virtual scene.

[0160] It should be understood that the fourth acoustic feature may include, but is not limited to: fourth acoustic response, fourth energy information, fourth acoustic parameter, or fourth acoustic feature.

[0161] In addition, the location information of the sound source device in the real scene can be recorded to facilitate the subsequent determination of the third acoustic features of one or more second locations in the virtual scene.

[0162] In one possible approach, the fourth acoustic feature is a user-defined input.

[0163] Since the acoustic features obtained from actual measurements in certain real-world scenarios may not be the most accurate match for the acoustic features that users actually hear, we directly input acoustic features that match the user's actual hearing in certain real-world scenarios. This makes the adjusted acoustic features closer to the acoustic features that match the user's actual hearing in certain real-world scenarios. In this way, the spatial audio signal generated subsequently based on the adjusted acoustic features is closer to the spatial audio signal that the user hears in the real-world scenario corresponding to certain virtual scenarios.

[0164] For example, "certain real-world scenarios" may include:

[0165] 1. Real-world scenarios with unavoidable interference noise. For example, unavoidable interference noise exists in or near the real-world scenario (such as construction in the next room). The fourth acoustic feature to be user-defined can be determined using the design parameters of this real-world scenario and relevant historical data.

[0166] 2. Real-world scenarios where the visual representation of the surface material does not match its actual acoustic effects. In such real-world scenarios, if the fourth acoustic feature is obtained by processing the test audio signal received at the second position in the real-world scenario, then when the subsequent spatial audio signal is generated and played, it may lead to a discrepancy between the user's auditory perception and visual perception.

[0167] 3. A real-world scenario serving as a multi-functional space; for example, a real-world scenario that functions as a lecture hall, concert hall, and film screening room. Because users have different auditory requirements for spatial audio signals in different virtual scenarios, the virtual acoustic characteristics of different virtual scenarios differ. If the actual acoustic characteristics are measured in the same real-world scenario for different virtual scenarios, the resulting generated spatial audio signals may lead to a discrepancy between the user's auditory and visual perception.

[0168] It should be noted that, in one possible approach, the sound source device can be deployed at one location (hereinafter referred to as the preset sound source location); in this way, one second location can correspond to one fourth acoustic feature. In another possible approach, the sound source device can be deployed at multiple preset sound source locations; in this way, one second location can correspond to multiple fourth acoustic features.

[0169] S303, acquire the third acoustic features of one or more second locations in the virtual scene.

[0170] Accordingly, acoustic features can be extracted based on the location information of the sound source device recorded in S302, the location information of the second location recorded in S302, the test audio signal played by the sound source device in S302, and the virtual scene to determine one or more third acoustic features for the second location. Similarly, one second location can correspond to one or more third acoustic features. It should be understood that the third acoustic features are virtual acoustic features.

[0171] It should be understood that the third acoustic feature may include, but is not limited to: third acoustic response, third energy information, third acoustic parameter, or third acoustic feature.

[0172] After executing S302 and S303, one or more acoustic feature groups can be obtained for a second position. An acoustic feature group includes a third acoustic feature and a fourth acoustic feature. The third acoustic feature and the fourth acoustic feature belonging to the same acoustic feature group have the same second position and the same preset sound source position.

[0173] S304, analyze the third and fourth acoustic features in one or more acoustic feature groups to obtain conversion information.

[0174] For example, this application can use signal processing, machine learning, numerical analysis, and other methods to analyze and compare the third and fourth acoustic features in one or more acoustic feature groups to obtain conversion information.

[0175] The following uses an acoustic feature set as an example to illustrate the method for generating conversion information.

[0176] In one possible approach, the third and fourth acoustic features in an acoustic feature set can be processed separately to obtain signal processing results. Then, the signal processing results of the third and fourth acoustic features are analyzed to obtain a transfer function. This transfer function can then be used as the transfer information. For example, for a third acoustic response such as RIR included in the third acoustic feature, and a fourth acoustic response such as RIR included in the fourth acoustic feature, where the RIR in the third acoustic response is represented by RIR 1 and the RIR in the fourth acoustic response by RIR 2, frequency domain transformations can be performed on RIR 1 and RIR 2 respectively to obtain the frequency domain response of RIR 1 (called frequency response 1) and the frequency domain response of RIR 2 (called frequency response 2). Next, in one possible approach, the frequency response transfer function can be calculated using frequency response 1 and frequency response 2, and this frequency response transfer function can be used as the transfer function. In another possible approach, signal analysis can be performed on frequency response 1 and frequency response 2 separately, and then the transfer function can be determined based on the signal analysis results.

[0177] In one possible approach, numerical analysis can be performed on the third and fourth acoustic features in an acoustic feature set to obtain the feature change rate; in this case, the feature change rate can be used as conversion information. Specifically, the change rate of the third acoustic feature relative to the fourth acoustic feature in the acoustic feature set can be calculated as the feature change rate. For example, for a third acoustic parameter such as direct-to-mix ratio included in the third acoustic feature, and a fourth acoustic parameter such as direct-to-mix ratio included in the fourth acoustic feature, where the direct-to-mix ratio in the third acoustic parameter is represented by DRR 1 and the direct-to-mix ratio in the fourth acoustic parameter is represented by DRR 2; the corresponding feature change rate can be (DRR 1-DRR 2) / DRR 2, or DRR 1 / DRR 2. It should be understood that the conversion information can include multiple feature change rates (multiple feature change rates correspond one-to-one with multiple acoustic parameters), and this application does not limit this.

[0178] One possible approach is to use machine learning to process the third and fourth acoustic features in an acoustic feature set to obtain conversion information. Specifically, the third and fourth acoustic features in an acoustic feature set can be input into an AI model (or machine learning model, hereinafter referred to as the second model). The second model processes the third and fourth acoustic features in the acoustic feature set and outputs conversion information; that is, the output information of the second model is the conversion information.

[0179] It should be understood that other methods can also be used to generate conversion information, and this application does not impose any restrictions on this.

[0180] S305, acquire the source audio signal, the sound source position of the source audio signal in the virtual scene, and the first position of the user in the virtual scene.

[0181] For example, S305 can be described with reference to the above description of S201, and will not be repeated here.

[0182] S306. Based on the first position, the sound source position of the source audio signal, and the virtual scene, acoustic features are extracted to obtain the second acoustic feature.

[0183] For example, the second acoustic feature can be understood as the virtual acoustic feature at the first position when the virtual sound source at the sound source location plays the source audio signal in the virtual scene.

[0184] S307, adjust the second acoustic feature according to the conversion information to obtain the first acoustic feature.

[0185] For example, the second acoustic feature can be adjusted according to the conversion information to obtain the first acoustic feature; in this way, when the virtual sound source at the sound source location in the virtual scene plays the source audio signal, the virtual acoustic feature at the first position can be adjusted to be close to or consistent with the real acoustic feature at the first position when the real sound source at the sound source location in the real scene plays the source audio signal.

[0186] For example, different adjustment methods can be used for different types of second acoustic features. For instance, when the second acoustic feature is a second acoustic response, a frequency response curve adjustment algorithm based on signal processing can be used. Specifically, the second acoustic response can be frequency-domain transformed to obtain the frequency domain response of the second acoustic response (hereinafter referred to as frequency response 3); then, the frequency response 3 can be adjusted according to the transformation function determined above to obtain the adjusted frequency response 3; after that, the adjusted frequency response 3 is transformed into the first acoustic response, that is, the first acoustic feature.

[0187] For example, when the second acoustic feature includes a second acoustic response and a second acoustic parameter, a reverberation adjustment algorithm based on signal processing can be used. Specifically, for the second DRR in the second acoustic parameter, the second DRR can be adjusted according to the rate of change of the direct-to-mixing ratio calculated above to obtain the first DRR (that is, the first acoustic parameter included in the first acoustic feature); then, the second acoustic response can be adjusted according to the first DRR to obtain the first acoustic response. It should be understood that other second acoustic parameters can also be adjusted according to the corresponding feature change rate, and this application does not limit this.

[0188] For example, the adjustment algorithm can be an AI model-based adjustment algorithm. Specifically, the second acoustic feature and the conversion information can be input into the AI ​​model (hereinafter referred to as the third model), and the third model can adjust the second acoustic feature according to the conversion information and output the first acoustic feature.

[0189] It should be understood that other adjustment algorithms can also be used to adjust the second acoustic feature to obtain the first acoustic feature, and this application does not limit this.

[0190] It should be noted that the conversion information obtained by S304 can be used to adjust the second acoustic feature of any first position in the virtual environment to obtain the first acoustic feature. In other words, S304 only needs to analyze the third and fourth acoustic features of one or more second positions, and the resulting conversion information has universal applicability to any position in the same room (i.e., virtual scene), rather than being used only to adjust the acoustic features of the first position that is close to the second position.

[0191] S308 generates a spatial audio signal based on the first acoustic characteristics and the source audio signal.

[0192] For example, when the first acoustic feature is the first acoustic response, the first acoustic response can be convolved with the source audio signal to obtain the spatial audio signal.

[0193] For example, when the first acoustic feature includes a first acoustic response and first energy information, the energy of the spatial audio signal can be adjusted according to the first energy information; then, the adjusted first acoustic response is convolved with the source audio to obtain the spatial audio signal.

[0194] For example, when the first acoustic feature includes the first acoustic parameter and the first acoustic feature, a first acoustic response can be generated based on the first acoustic parameter and the first acoustic feature; then, the first acoustic response can be convolved with the source audio to obtain a spatial audio signal.

[0195] For example, when the first acoustic feature includes a first acoustic response, a first acoustic parameter, and a first acoustic feature, the first acoustic response can be optimized first based on the first acoustic parameter and the first acoustic feature. Then, the optimized first acoustic response is convolved with the source audio to obtain the spatial audio signal.

[0196] It should be understood that there are many ways to generate spatial audio signals based on the first acoustic features, and this application does not limit this.

[0197] One possible approach is to pre-generate and store pre-defined conversion information corresponding to multiple preset virtual scenes in a database, and pre-generate and store second preset acoustic features for multiple locations within these preset virtual scenes in the database. During application, the conversion information corresponding to the current virtual scene can be retrieved from the database, and the first position within the current virtual scene can be determined in real-time. This allows for the retrieval of second acoustic features matching the first position and the sound source position from the database, which are then adjusted in real-time to obtain a first acoustic feature matching the first position and the sound source position within the current virtual scene. Utilizing a database can shorten the time required to generate spatial audio signals, improving the user experience. Furthermore, it allows for a more efficient way to invoke virtual scenes that may be used multiple times at different times, by different users, and in different applications.

[0198] Figure 4 This is a schematic diagram illustrating an audio processing procedure as an example. Figure 4 In this embodiment, it is not limited to whether the user's current real-world scene is the same as the virtual scene corresponding to the audio project (or AR / VR project) selected (or experienced) by the user. In other words, Figure 4 The application scenarios of the embodiments can be those described above. Figures 1A to 1D .

[0199] S401, acquire the source audio signal, the sound source position of the source audio signal in the virtual scene, and the first position of the user in the virtual scene.

[0200] S402, Select the second acoustic feature set corresponding to the virtual scene from multiple second acoustic feature sets stored in the database.

[0201] S403, based on the sound source location and the first position of the source audio signal, select the second acoustic feature from the second acoustic feature set corresponding to the virtual scene.

[0202] For example, multiple sets of second acoustic features can be generated in advance for various preset virtual scenes; wherein, the multiple sets of second acoustic features correspond one-to-one with the multiple preset virtual scenes.

[0203] Specifically, for a preset virtual scene, multiple locations (hereinafter referred to as fourth locations) can be pre-identified within the preset virtual scene. For example, if the preset virtual scene is divided into a 10*10 grid, the intersection of the grid lines or the center point of the grid can be used as the fourth location. Next, for each fourth location, acoustic features can be extracted based on the fourth location, the preset sound source locations in the preset virtual scene, and the preset virtual scene itself, to obtain the second preset acoustic features for that fourth location. There can be multiple preset sound source locations in the preset virtual scene, thus, multiple second preset acoustic features can be obtained for each fourth location. Following this method, multiple second preset acoustic features for fourth locations can be obtained; these multiple second preset acoustic features for fourth locations can form a second acoustic feature set corresponding to the preset virtual scene. Then, the second acoustic feature set can be stored in a database, and a correspondence between the second acoustic feature set and the preset virtual scene can be established.

[0204] In accordance with the above method, multiple sets of second acoustic features can be generated; and multiple sets of second acoustic features can be stored in a database, and a correspondence between multiple sets of second acoustic features and corresponding preset virtual scenes can be established.

[0205] For example, a second acoustic feature set corresponding to the current virtual scene can be selected from the database first; then, based on the sound source position and the fourth position of the source audio signal, a second preset acoustic feature with the same or closest preset sound source position and the same or closest fourth position as the first position can be selected from the selected second acoustic feature set as the second acoustic feature.

[0206] S404, Based on the virtual scene, select conversion information from multiple preset conversion information stored in the database; wherein, the multiple preset conversion information corresponds one-to-one with multiple preset virtual scenes.

[0207] For example, for one of a variety of preset virtual scenes, preset conversion information can be generated in advance according to the descriptions in S302 to S304 above, and the preset conversion information can be stored in a database, and a correspondence between the preset conversion information and the preset virtual scene can be established. During the execution of S404, preset conversion information that is the same as the current virtual scene in the corresponding preset virtual scene can be selected from the database as the conversion information corresponding to the current virtual scene.

[0208] S405, adjust the second acoustic feature according to the conversion information to obtain the first acoustic feature.

[0209] S406 generates a spatial audio signal based on the first acoustic feature and the source audio signal.

[0210] S405 to S406 can be referred to the descriptions of S307 to S308 above, and will not be repeated here.

[0211] One possible approach is to pre-generate first acoustic features for multiple locations within various preset virtual scenes and store them in a database. During application, the database can be used to retrieve the first acoustic features matching the first location and sound source location in the current virtual scene, further reducing the time required to generate spatial audio signals and improving user experience. Alternatively, this approach allows for a more efficient way to invoke virtual scenes that may be used multiple times at different times, by different users, and in different applications.

[0212] Figure 5 This is a schematic diagram illustrating an audio processing procedure as an example. Figure 5 In this embodiment, it is not limited to whether the user's current real-world scene is the same as the virtual scene corresponding to the audio project (or AR / VR project) selected (or experienced) by the user. In other words, Figure 5 The application scenarios of the embodiments can be those described above. Figures 1A to 1D .

[0213] S501, acquire the source audio signal, the sound source position of the source audio signal in the virtual scene, and the first position of the user in the virtual scene.

[0214] S502, select the first acoustic feature set corresponding to the virtual scene from multiple first acoustic feature sets stored in the database.

[0215] S503, based on the sound source location and the first position of the source audio signal, select the first acoustic feature from the first acoustic feature set corresponding to the virtual scene.

[0216] For example, for one of a variety of preset virtual scenarios, according to the above... Figure 3 The example generates corresponding conversion information, and according to Figure 4 The embodiment generates a corresponding second acoustic feature set. Then, according to the transformation information corresponding to the preset virtual scene, the second preset acoustic features at multiple positions (hereinafter referred to as third positions) in the second acoustic feature set corresponding to the preset virtual scene are adjusted to obtain multiple first preset acoustic features at the third positions. These multiple first preset acoustic features at the third positions can form a first acoustic feature set, which is a first acoustic feature set corresponding to the preset virtual scene. Then, the first acoustic feature set corresponding to the preset virtual scene can be stored in a database, and a correspondence between the first acoustic feature set and the preset virtual scene can be established.

[0217] It should be understood that there can be multiple preset sound source locations in this preset virtual scene. Thus, for a third location, multiple second preset acoustic features can be obtained. Correspondingly, for a third location, multiple first preset acoustic features can be obtained, and each first preset acoustic feature corresponds to a preset acoustic location.

[0218] For example, a first acoustic feature set corresponding to the current virtual scene can be filtered from the database first; then, from the filtered first acoustic feature set, a first preset acoustic feature that is the same as or closest to the preset sound source position and the third position is the same as or closest to the first position is selected as the first acoustic feature.

[0219] S504 generates a spatial audio signal based on the first acoustic characteristics and the source audio signal.

[0220] S504 can be referred to the description of S308 above, and will not be repeated here.

[0221] Figure 6A and Figure 6B This is a schematic diagram of the frequency response curve of an acoustic feature as an example. Figure 6A and Figure 6B The curves show the frequency response curves of the real acoustic features (curve 1), the adjusted virtual acoustic features (i.e., the first acoustic feature) (curve 2), and the unadjusted virtual acoustic features (i.e., the second acoustic feature) (curve 3) when the user is in different initial positions within a virtual scene. Comparing curves 1, 2, and 3, it is evident that the frequency response curve of the adjusted virtual acoustic feature is closer to that of the real acoustic feature; that is, by using the method provided in this application, the virtual acoustic features can be made closer to the real acoustic features, thereby improving the quality of the spatial audio signal.

[0222] Figure 7 This is a schematic diagram of an audio processing device as an example.

[0223] The first acquisition module 701 is used to acquire the source audio signal, the sound source position of the source audio signal in the virtual scene, and the first position of the user in the virtual scene, wherein the virtual scene is obtained by modeling the real scene;

[0224] The second acquisition module 702 is used to acquire a first acoustic feature based on the sound source location and the first location. The first acoustic feature is obtained by adjusting the second acoustic feature based on the conversion information. The second acoustic feature is obtained by extracting acoustic features based on the first location, the virtual scene, and the sound source location. The conversion information is used to describe the conversion relationship between acoustic features at the same location in the virtual scene and the real scene.

[0225] The audio signal generation module 703 is used to generate a spatial audio signal based on the first acoustic features and the source audio signal.

[0226] For example, the conversion information is determined by analyzing a set of acoustic features, which includes a third acoustic feature and a fourth acoustic feature. The fourth acoustic feature is the acoustic feature of the second location in the real scene, and the third acoustic feature is the acoustic feature of the second location in the virtual scene.

[0227] For example, there may be one or more acoustic feature groups, one or more second positions, and the second positions corresponding to the third and fourth acoustic features belonging to the same acoustic feature group are the same.

[0228] For example, the conversion information includes a conversion function, which is obtained by analyzing the signal processing results of the third and fourth acoustic features.

[0229] For example, the conversion information includes the feature change rate, which is the rate of change of the third acoustic feature relative to the fourth acoustic feature.

[0230] For example, the transformation information includes model output information obtained by inputting the third and fourth acoustic features into the model.

[0231] For example, the first acoustic feature includes at least one of the following: acoustic response, energy information, acoustic parameters, or acoustic characteristics.

[0232] For example, the fourth acoustic feature is obtained by processing a test audio signal received at a second location in a real scene.

[0233] For example, the fourth acoustic feature is a user-defined input.

[0234] For example, the second acquisition module 702 is specifically used to determine one or more acoustic feature groups, one acoustic feature group including a third acoustic feature and a fourth acoustic feature, the fourth acoustic feature being the acoustic feature of the second position in the real scene, and the third acoustic feature being the acoustic feature of the second position in the virtual scene; analyze the third acoustic feature and the fourth acoustic feature in one or more acoustic feature groups to obtain conversion information; extract acoustic features based on the first position, the virtual scene and the sound source position to obtain a second acoustic feature; adjust the second acoustic feature according to the conversion information to obtain a first acoustic feature.

[0235] For example, the second acquisition module 702 is specifically used to select a second acoustic feature set corresponding to a virtual scene from multiple second acoustic feature sets stored in the database; wherein, the multiple second acoustic feature sets correspond one-to-one with multiple preset virtual scenes, and a second acoustic feature set includes multiple second preset acoustic features, and a second acoustic feature set is obtained by acoustic feature extraction based on a preset virtual scene, multiple fourth positions, and multiple preset sound source positions, and a fourth position and a preset sound source position are used to determine a second preset acoustic feature in a second acoustic feature set; select a second acoustic feature from the second acoustic feature set corresponding to the virtual scene according to the sound source position and the first position; select conversion information from multiple preset conversion information stored in the database according to the virtual scene, and the multiple preset conversion information correspond one-to-one with multiple preset virtual scenes; adjust the second acoustic feature according to the conversion information to obtain a first acoustic feature.

[0236] For example, the second acquisition module 702 is specifically used to acquire the first acoustic feature, including:

[0237] From multiple sets of first acoustic features stored in the database, a first acoustic feature set corresponding to the virtual scene is selected. These multiple first acoustic feature sets correspond one-to-one with multiple preset virtual scenes, and each set also corresponds one-to-one with multiple second acoustic feature sets. A second acoustic feature set is obtained by extracting acoustic features based on a preset virtual scene, multiple third positions, and multiple preset sound source positions. A fourth position and a preset sound source position are used to determine a second preset acoustic feature within a second acoustic feature set. A first preset acoustic feature within a first acoustic feature set is obtained by adjusting a second preset acoustic feature within the corresponding second acoustic feature set based on conversion information. Based on the sound source position and the first position, a first acoustic feature is selected from the first acoustic feature set corresponding to the virtual scene.

[0238] This application also provides an AR device, which includes a display module, an image acquisition module, headphones, and a processor, wherein: the processor is used to execute the audio processing method as described in the above embodiments; and the headphones are used to play the spatial audio signal generated by the audio processing method described in the above embodiments.

[0239] This application also provides a VR device, which includes a display module, an image acquisition module, headphones, and a processor, wherein: the processor is used to execute the audio processing method as described in the above embodiments; and the headphones are used to play the spatial audio signal generated by the audio processing method described in the above embodiments.

[0240] In one example, Figure 8A schematic block diagram illustrating an embodiment of the present application shows an apparatus 800. The apparatus 800 may include a processor 801 and a transceiver / transceiver pin 802, and optionally, a memory 803.

[0241] The various components of device 800 are coupled together via bus 804, which includes a data bus, a power bus, a control bus, and a status signal bus. However, for clarity, all buses are referred to as bus 804 in the figure.

[0242] Optionally, the memory 803 can be used to store instructions from the foregoing method embodiments. The processor 801 can be used to execute the instructions in the memory 803, control the receive pin to receive signals, and control the transmit pin to transmit signals.

[0243] The device 800 may be an electronic device or a chip of an electronic device in the above method embodiments.

[0244] All relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.

[0245] This application also provides a chip, including one or more interface circuits and one or more processors; the one or more processors receive or send data through the one or more interface circuits, and when the one or more processors execute computer instructions, the steps of the above-described related method steps that implement the method in the above embodiments are executed. The interface circuit is a transceiver / transceiver pin 802.

[0246] This embodiment also provides a computer-readable storage medium storing computer instructions. When the computer instructions are executed on an electronic device, the electronic device performs the aforementioned method steps to implement the methods described in the above embodiments.

[0247] This embodiment also provides a computer program product containing computer instructions that, when executed by a computer or processor, cause the computer to perform the aforementioned related steps to implement the methods described in the above embodiments.

[0248] In addition, embodiments of this application also provide an apparatus, which may specifically be a chip, component, or module. The apparatus may include a connected processor and a memory; wherein the memory is used to store computer execution instructions, and when the apparatus is running, the processor may execute the computer execution instructions stored in the memory to cause the chip to execute the methods in the above-described method embodiments.

[0249] In this embodiment, the electronic device, computer-readable storage medium, computer program product or chip are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding methods provided above, and will not be repeated here.

[0250] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0251] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0252] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0253] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0254] Any content in the various embodiments of this application, as well as any content in the same embodiment, can be freely combined. Any combination of the above content is within the scope of this application.

[0255] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0256] The steps of the methods or algorithms described in conjunction with the embodiments of this application can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, read-only optical discs (CD-ROMs), or any other form of storage medium well known in the art. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can reside in an ASIC.

[0257] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer-readable storage media and communication media, wherein communication media include any medium that facilitates the transmission of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0258] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An audio processing method, characterized in that, The method includes: The system acquires a source audio signal, the sound source location of the source audio signal in a virtual scene, and the first location of the user in the virtual scene, wherein the virtual scene is obtained by modeling a real scene. Based on the sound source location and the first location, a first acoustic feature is obtained. The first acoustic feature is obtained by adjusting a second acoustic feature based on conversion information. The second acoustic feature is obtained by extracting acoustic features based on the first location, the virtual scene, and the sound source location. The conversion information is used to describe the conversion relationship between acoustic features at the same location in the virtual scene and the real scene. A spatial audio signal is generated based on the first acoustic feature and the source audio signal.

2. The method according to claim 1, characterized in that, The conversion information is determined by analyzing a set of acoustic features, which includes a third acoustic feature and a fourth acoustic feature. The fourth acoustic feature is the acoustic feature of the second position in the real scene, and the third acoustic feature is the acoustic feature of the second position in the virtual scene.

3. The method according to claim 2, characterized in that, The acoustic feature group can be one or more, the second position can be one or more, and the second position corresponding to the third acoustic feature and the fourth acoustic feature belonging to the same acoustic feature group is the same.

4. The method according to claim 2 or 3, characterized in that, The conversion information includes a conversion function, which is obtained by analyzing the signal processing results of the third acoustic feature and the fourth acoustic feature.

5. The method according to any one of claims 2 to 4, characterized in that, The conversion information includes a feature change rate, which is the rate of change of the third acoustic feature relative to the fourth acoustic feature.

6. The method according to any one of claims 2 to 5, characterized in that, The conversion information includes model output information obtained by inputting the third acoustic feature and the fourth acoustic feature into the model.

7. The method according to any one of claims 1 to 6, characterized in that, The first acoustic feature includes at least one of the following: acoustic response, energy information, acoustic parameters, or acoustic features.

8. The method according to any one of claims 2 to 6, characterized in that, The fourth acoustic feature is obtained by processing the test audio signal received at the second location in the real scene.

9. The method according to any one of claims 2 to 6, characterized in that, The fourth acoustic feature is a user-defined input.

10. The method according to any one of claims 1 to 9, characterized in that, Based on the sound source location and the first location, a first acoustic feature is obtained, including: One or more acoustic feature groups are determined, wherein an acoustic feature group includes a third acoustic feature and a fourth acoustic feature, wherein the fourth acoustic feature is the acoustic feature of the second position in the real scene, and the third acoustic feature is the acoustic feature of the second position in the virtual scene; The conversion information is obtained by analyzing the third and fourth acoustic features in one or more acoustic feature groups; Based on the first location, the virtual scene, and the sound source location, acoustic features are extracted to obtain the second acoustic feature; The second acoustic feature is adjusted according to the conversion information to obtain the first acoustic feature.

11. The method according to any one of claims 1 to 9, characterized in that, Based on the sound source location and the first location, a first acoustic feature is obtained, including: From multiple sets of first acoustic features stored in the database, a first acoustic feature set corresponding to the virtual scene is selected; wherein, the multiple sets of first acoustic features correspond one-to-one with multiple preset virtual scenes, and the multiple sets of first acoustic features correspond one-to-one with multiple sets of second acoustic features. A second acoustic feature set is obtained by extracting acoustic features based on a preset virtual scene, multiple third positions, and multiple preset sound source positions. A third position and a preset sound source position are used to determine a second preset acoustic feature in the second acoustic feature set. A first preset acoustic feature in a first acoustic feature set is obtained by adjusting a second preset acoustic feature in the corresponding second acoustic feature set according to the conversion information. Based on the sound source location and the first location, the first acoustic feature is selected from the first acoustic feature set corresponding to the virtual scene.

12. The method according to any one of claims 1 to 9, characterized in that, Based on the sound source location and the first location, a first acoustic feature is obtained, including: From multiple sets of second acoustic features stored in the database, a second acoustic feature set corresponding to the virtual scene is selected; wherein, the multiple sets of second acoustic features correspond one-to-one with multiple preset virtual scenes, and a second acoustic feature set includes multiple second preset acoustic features. The second acoustic feature set is obtained by acoustic feature extraction based on a preset virtual scene, multiple fourth positions, and multiple preset sound source positions. A fourth position and a preset sound source position are used to determine a second preset acoustic feature in the second acoustic feature set. Based on the sound source location and the first location, the second acoustic feature is selected from the second acoustic feature set corresponding to the virtual scene; Based on the virtual scene, the conversion information is selected from multiple preset conversion information stored in the database, and the multiple preset conversion information corresponds one-to-one with multiple preset virtual scenes; The second acoustic feature is adjusted according to the conversion information to obtain the first acoustic feature.

13. An audio processing apparatus, characterized in that, The device includes: The first acquisition module is used to acquire a source audio signal, the sound source position of the source audio signal in a virtual scene, and the first position of the user in the virtual scene, wherein the virtual scene is obtained by modeling a real scene; The second acquisition module is used to acquire a first acoustic feature based on the sound source location and the first location. The first acoustic feature is obtained by adjusting the second acoustic feature based on conversion information. The second acoustic feature is obtained by extracting acoustic features based on the first location, the virtual scene, and the sound source location. The conversion information is used to describe the conversion relationship between acoustic features at the same location in the virtual scene and the real scene. An audio signal generation module is used to generate a spatial audio signal based on the first acoustic feature and the source audio signal.

14. An augmented reality (AR) device, characterized in that, The AR device includes a display module, an image acquisition module, headphones, and a processor, wherein: The processor is configured to perform the audio processing method as described in any one of claims 1 to 12 above; The headphones are used to play spatial audio signals generated by the audio processing method according to any one of claims 1 to 12.

15. A virtual reality (VR) device, characterized in that, The VR device includes a display module, an image acquisition module, headphones, and a processor, wherein: The processor is configured to perform the audio processing method as described in any one of claims 1 to 12 above; The headphones are used to play spatial audio signals generated by the audio processing method according to any one of claims 1 to 12.

16. An electronic device, characterized in that, include: A memory and a processor, wherein the memory is coupled to the processor; The memory stores program instructions that, when executed by the processor, cause the electronic device to perform the audio processing method as described in any one of claims 1 to 12.

17. A chip, characterized in that, It includes one or more interface circuits and one or more processors; the one or more processors receive or send data through the one or more interface circuits, and when the one or more processors execute computer instructions, the steps of the method as described in any one of claims 1 to 12 are performed.

18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when run on a computer or processor, causes the computer or processor to perform the audio processing method as described in any one of claims 1 to 12.

19. A computer program product, characterized in that, The computer program product includes computer instructions that, when executed by a computer or processor, cause the steps of the method as described in any one of claims 1 to 12 to be performed.

Citation Information

Patent Citations

  • Method and apparatus for generating virtual or augmented reality presentations with 3D audio positioning

    CN109564760A

  • Audio playing method and device, storage medium, client and live broadcast system

    CN115237250A