Audio processing method, audio processing device, computer device, and computer program

The audio processing method in augmented reality devices integrates real and virtual sounds through mixing modes, enhancing auditory fusion and immersion by outputting a mixed signal with a specific effect, addressing the lack of comprehensive audio fusion in existing technologies.

JP2026501197APending Publication Date: 2026-01-14TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025535383
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-08-09
Filing Date
2024-06-21
Publication Date
2026-01-14

AI Technical Summary

Technical Problem

Augmented reality devices primarily focus on visual fusion, lacking comprehensive audio fusion capabilities, which limits the user's sense of immersion in augmented reality scenes.

Method used

An audio processing method that integrates real-world and virtual-world sounds by selecting a mixing mode from multiple options to create a mixed signal, enhancing auditory fusion and immersion.

Benefits of technology

The method achieves deep auditory fusion of real and virtual sounds, enriching the augmented reality experience and improving overall user immersion by outputting a mixed signal with a corresponding mixing effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026501197000001_ABST
    Figure 2026501197000001_ABST
Patent Text Reader

Abstract

[0006] An embodiment of the present application provides an audio processing method and a related device. The method includes the steps of: acquiring a real audio signal and a virtual audio signal in an augmented reality scene; determining a plurality of mixing modes arranged for the augmented reality scene and selecting one mixing mode from the plurality of mixing modes as a target mixing mode, where different mixing modes are for realizing different mixing effects between the real audio signal and the virtual audio signal; and acquiring and outputting a mixed signal by mixing the real audio signal and the virtual audio signal in the target mixing mode. In this way, sound fusion in the augmented reality scene can be realized in an auditory dimension, enriching the fusion ability in the augmented reality scene and improving a user's overall sense of immersion in the augmented reality scene.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority to a Chinese patent application filed with the China Patent Office on August 9, 2023, bearing application number 2023109935634 and entitled "Audio Processing Method and Related Device," the entire contents of which are incorporated herein by reference.

[0002] This application relates to Internet technology, and more particularly to the field of computers, and more particularly to audio processing methods and related devices. [Background technology]

[0003] With the development of augmented reality (AR) technology, the functions of augmented reality devices (e.g., AR glasses, AR headsets, etc.) are becoming increasingly rich and diverse. The visual fusion capabilities of augmented reality devices (e.g., AR glasses, AR headsets, etc.) allow people to immerse themselves in augmented reality scenes created by combining the real and virtual worlds, offering new experiences. For example, when a user wears AR glasses, they can view a composite image of real-world and virtual-world images, and can change the image content of the virtual world through interaction methods such as gestures and voice. Augmented reality applications can also detect and enhance two-dimensional images in the user's environment, such as enhancing product posters at an exhibition. However, these applications only achieve image fusion in augmented reality scenes in the visual dimension, and the fusion capabilities are limited, so the user's sense of immersion in the augmented reality scenes needs to be further improved. Summary of the Invention [Problem to be solved by the invention]

[0004] The embodiments of the present application provide an audio processing method and related devices that can realize the fusion of real-world sounds and virtual-world sounds in an augmented reality scene in an auditory dimension, enrich the fusion capabilities in the augmented reality scene, and improve the user's overall sense of immersion in the augmented reality scene. [Means for solving the problem]

[0005] According to one aspect, an embodiment of the present application provides an audio processing method, the method comprising: acquiring real audio signals and virtual audio signals in an augmented reality scene, wherein the real audio signals refer to audio signals collected in a real environment on which the augmented reality scene is based, and the virtual audio signals refer to audio signals constructed in a virtual environment on which the augmented reality scene is based; determining a plurality of mixing modes arranged for the augmented reality scene, and selecting one mixing mode from the plurality of mixing modes as a target mixing mode, where different mixing modes are for realizing different mixing effects between the real audio signal and the virtual audio signal; obtaining a mixed signal by mixing the real audio signal and the virtual audio signal in a target mixing mode; and outputting the mixed signal.

[0006] According to one aspect, an embodiment of the present application provides an audio processing device, the device comprising: an acquisition unit for acquiring real audio signals and virtual audio signals in an augmented reality scene, wherein the real audio signals refer to audio signals collected in a real environment on which the augmented reality scene is based, and the virtual audio signals refer to audio signals constructed in a virtual environment on which the augmented reality scene is based; a processing unit for determining a plurality of mixing modes arranged for the augmented reality scene, selecting one mixing mode from the plurality of mixing modes as a target mixing mode, and performing a mixing process on the real audio signal and the virtual audio signal in the target mixing mode to obtain a mixing signal, wherein different mixing modes are for realizing different mixing effects between the real audio signal and the virtual audio signal; an output unit for outputting a mixed signal; According to one aspect, an embodiment of the present application provides a computing device, the computing device comprising: a processor suitable for executing a computer program; and a computer-readable storage medium having stored thereon a computer program which, when executed by a processor, causes the audio processing method as described above to be implemented.

[0007] According to one aspect, an embodiment of the present application provides a computer-readable storage medium having stored thereon a computer program that, when loaded by a processor, causes the computer program to perform the audio processing method as described above.

[0008] According to one aspect, an embodiment of the present application provides a computer program product comprising a computer program or computer instructions which, when executed by a processor, implements the audio processing method described above. [Effects of the Invention]

[0009] In an embodiment of the present application, a true audio signal and a virtual audio signal in an augmented reality scene can be obtained. The true audio signal refers to an audio signal collected in a real environment on which the augmented reality scene is based, and the virtual audio signal refers to an audio signal constructed in a virtual environment on which the augmented reality scene is based. A plurality of mixing modes arranged for the augmented reality scene is determined, and one mixing mode from the plurality of mixing modes is selected as a target mixing mode. Here, different mixing modes can be used to achieve different mixing effects between the true audio signal and the virtual audio signal, resulting in different auditory experiences. By determining a plurality of mixing modes for the augmented reality scene, various options are provided for setting the target mixing mode, allowing the target mixing mode to be flexibly set. This can accommodate various scenes requiring mixing of the true audio signal and the virtual audio signal, and can also meet the user's customized mixing needs. A mixed signal is obtained by mixing the true audio signal and the virtual audio signal in the target mixing mode, and the mixed signal is output. The target mixing mode can achieve a corresponding mixing effect between real audio signals and virtual audio signals. By controlling the mixing process between real audio signals and virtual audio signals in the target mixing mode, the fusion of real-world sounds and virtual-world sounds on which the augmented reality scene is based can be achieved in the auditory dimension, thereby increasing the auditory fusion of sounds in the augmented reality scene and enriching the fusion capabilities of the augmented reality scene. The resulting mixing signal also has a mixing effect corresponding to the target mixing mode. Finally, the output of the mixing signal allows the user to more intuitively experience the fusion effect of sounds in the augmented reality scene, enhancing the user's immersion in the augmented reality scene. [Brief explanation of the drawings]

[0010] [Figure 1a]1 is an architecture diagram of an audio processing system provided in an embodiment of the present application. [Figure 1b] FIG. 2 is an architecture diagram of another audio processing system provided in an embodiment of the present application. [Figure 2] FIG. 2 is a flow diagram of an audio processing method provided in an embodiment of the present application. [Figure 3] 1 is a schematic diagram of an audio collection device provided in an embodiment of the present application. [Figure 4] FIG. 2 is a schematic flow diagram of another audio processing method provided in an embodiment of the present application. [Figure 5a] FIG. 1 is a schematic diagram of a scene for setting a target mixing mode provided in an embodiment of the present application. [Figure 5b] FIG. 2 is a schematic diagram of framing of a current audio signal provided in an embodiment of the present application; [Figure 5c] 3A and 3B are schematic diagrams illustrating the relationship between audio frames before and after the audio validity determination process provided in the embodiment of the present application. [Figure 5d] FIG. 2 is a schematic diagram of dividing an audio frame into audio sub-blocks provided in an embodiment of the present application; [Figure 6a] FIG. 2 is a schematic diagram of the flow of volume control provided in an embodiment of the present application. [Figure 6b] FIG. 2 is a schematic diagram of an audio processing flow provided in an embodiment of the present application. [Figure 7] FIG. 1 is a schematic diagram of the configuration of an audio processing device provided in an embodiment of the present application. [Figure 8] FIG. 1 is a schematic diagram of the configuration of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, the configuration of the embodiments of the present application will be described clearly and completely with reference to the drawings of the embodiments of the present application. It is clear that the described embodiments are not all the embodiments but only some of the embodiments of the present application. All other embodiments that a person skilled in the art can obtain from the embodiments of the present application without creative work fall within the scope of protection of the present application.

[0012] This application proposes an audio processing method. The audio processing method selects a target mixing mode from multiple mixing modes arranged for an augmented reality scene, and mixes a real audio signal and a virtual audio signal in the augmented reality scene in the target mixing mode to obtain a mixed signal, thereby achieving deep fusion of real-world sounds and virtual-world sounds in the augmented reality (AR) scene in the auditory plane and enriching the fusion capabilities of the augmented reality scene. The mixed signal can also be output. The mixed signal has a mixing effect corresponding to the target mixing mode. By combining the fusion of images in the visual plane with the fusion of sounds in the auditory plane, the overall atmosphere and immersion of the augmented reality scene can be improved. The multi-dimensional fusion capabilities further enhance the realism of the augmented reality scene.

[0013] In the present embodiment, an augmented reality scene is an interaction scene realized by augmented reality technology. Depending on the content of the interaction, the augmented reality scene may include, but is not limited to, a social scene, a game scene, a choir scene, a live performance scene, etc. Here, augmented reality technology skillfully integrates virtual information with the real world. It broadly utilizes various technological means, such as multimedia, 3D modeling, real-time monitoring and registration, smart interaction, and sensing, to apply simulated computer-generated virtual information, such as text, images, 3D models, music, and videos, to the real world in a complementary manner, thereby achieving "augmentation" of the real world. In an augmented reality scene, a user wears an AR device, such as AR glasses or an AR headset. Through the AR glasses, the user can view the real world and also view virtual images or animations that have been processed and projected onto the lenses of the AR glasses. The user can also actually hear sounds from the real world and the virtual scene from the headset. Such an audio-video experience that integrates the real and virtual worlds can bring the user a different imaginary space and experience.

[0014] An augmented reality scene is created based on a real environment and a virtual environment. The real environment is the actual physical environment in which a user is located. It includes the surrounding scenery, objects, and sounds that the user can see, hear, and feel, and is the basis for people's perception and interaction. The virtual environment is a simulated environment generated by a device. Corresponding technology can immerse a user in a virtual environment, or the virtual environment can augment the real environment. The user can sense and interact with it using an augmented reality device (e.g., AR glasses). The real environment can be understood as being composed of a part of the real world. The real world includes invisible levels of matter in addition to what the user can see, hear, and feel. The virtual environment can be understood as being composed of a part of the virtual world. In a virtual world, various visible, audible, and sensible contents are entirely constructed by the device. For example, sounds in a virtual world are entirely generated by a computer device. In one implementation, in an augmented reality scene, a true audio signal can be formed by collecting audio signals generated from sound sources in the real environment, and a virtual audio signal can be formed by collecting audio signals generated from sound sources in the virtual environment. Here, the audio signals generated from sound sources in the virtual environment can be automatically constructed by a device and can simulate the real environment. In a specific formation process, an analog-to-digital conversion process can be performed on the audio signals (a type of analog signal) to obtain corresponding audio signals (a type of digital signal).

[0015] The mixing mode is a mode for controlling the mixing of real audio signals and virtual audio signals. Different mixing modes can be used to achieve different mixing effects between real audio signals and virtual audio signals, resulting in different auditory experiences. For real audio signals and virtual audio signals in an augmented reality scene, the mixing manner between the two can be controlled by a set target mixing mode, and sounds generated from virtual world sound sources can be mixed with sounds generated from real world sound sources, thereby enhancing the sounds in the real world.

[0016] Based on the above definitions, the principle of the audio processing method proposed in the embodiments of the present application will be described below. Specifically, the general principle of this method is as follows: A true audio signal and a virtual audio signal in an augmented reality scene are obtained, a plurality of mixing modes arranged for the augmented reality scene are determined, and one mixing mode is selected from the plurality of mixing modes as a target mixing mode. The target mixing mode may be selected by a user of the augmented reality device according to their own usage needs. For example, the target mixing mode may be determined according to a manual operation by the user or automatically determined by the device. Providing a plurality of mixing modes provides more diverse selection of mixing effects between the true audio signal and the virtual audio signal. A mixed signal is obtained by mixing the true audio signal and the virtual audio signal in the target mixing mode, and the mixed signal is output.

[0017] In a specific implementation, the above-mentioned method may be executed by a computer device. The computer device may be a terminal or a server. For example, as shown in FIG. 1a, a terminal (e.g., AR glasses) may collect real audio signals and virtual audio signals in an augmented reality scene, select one mixing mode from among multiple mixing modes in response to a user's command, and mix the real audio signals and virtual audio signals in the selected mixing mode. The mixed mixed signal is then output from the terminal. Alternatively, the above-mentioned method may be jointly executed by a terminal and a server. For example, as shown in FIG. 1b, a terminal (e.g., AR glasses or an AR headset) may collect real audio signals and virtual audio signals in an augmented reality scene in real time and send the collected audio signals to a server. The terminal may also determine a target mixing mode in response to a user operation. The server may obtain a mixed signal by mixing the real audio signals and virtual audio signals in the determined target mixing mode, and send the mixed signal to the terminal for output.

[0018] Here, the terminals mentioned above include, but are not limited to, devices such as smartphones, tablet computers, smart wearable devices, smart voice interaction devices, smart home appliances, personal computers, in-vehicle terminals, smart cameras, and augmented reality devices (e.g., AR glasses and AR headsets). This application does not limit the number of terminals. The server may be an independent physical server, a server cluster or a distributed system consisting of multiple physical servers, or a cloud server that provides base cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. This application does not limit the number of servers.

[0019] The audio processing method provided herein relates to artificial intelligence (AI) speech technology. Key speech technologies include automatic speech recognition (ASR), text-to-speech (TTS), and voiceprint recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction. Voice is one of the most promising human-computer interaction methods. Big model technology will revolutionize the development of speech technology. Pre-trained models based on the Transformer (a type of attention-based neural network) architecture, such as WavLM and UniSpeech, have strong generalizability and versatility and can effectively complete various voice processing tasks. In this application, AI models may be used to mix real audio signals and virtual audio signals in augmented reality scenes. AI models include, but are not limited to, neural networks and pre-trained models.

[0020] In addition, the collection and processing of relevant data (e.g., audio signals, geographical location information of the real environment, etc.) in this specification, when applied in practice, should strictly comply with the requirements of relevant local laws and regulations, obtain the informed consent or independent consent of the personal information subject, and carry out subsequent data use and processing actions within the scope of laws and regulations and the permission of the personal information subject.

[0021] Based on the above description, an embodiment of the present application proposes an audio processing method. The audio processing method may be executed by the above-mentioned computer device (terminal or server), or may be executed jointly by the terminal and the server. For ease of explanation, the following description will be given taking the case where the audio processing method is executed by a computer device as an example. Referring to Figure 2, the audio processing method may include the following steps S201 to S203.

[0022] In step S201, a real audio signal and a virtual audio signal in an augmented reality scene are obtained.

[0023] A true audio signal refers to an audio signal collected in a real environment on which an augmented reality scene is based. The real environment on which an augmented reality scene is based is a real physical environment in which a user is located. By collecting sounds in the real environment, a corresponding true audio signal can be obtained. The audio signal collected in the real environment can represent real-world sounds, including, but not limited to, the speech of a speaker in the real environment, sounds played by a speaker in the real environment, and environmental noise in the real environment.

[0024] In one implementation, a true audio signal in an augmented reality scene may be obtained by collecting sounds emitted from sound sources in the real environment using an audio collection device with stereo collection capabilities. The audio collection device may include, but is not limited to, a stereo microphone, a headset with stereo recording capabilities, etc. The audio collection device with stereo collection capabilities can collect sounds from the external real world to obtain a three-dimensional audio signal with a sense of direction. For example, a headset with stereo recording capabilities can collect real-world sounds in real time to obtain a stereo signal. A stereo recording headset may be as shown in FIG. 3. Microphones are located near the left and right ears of the headset to collect left and right audio signals, respectively. Furthermore, the collected audio signals undergo digital signal processing and are mixed with the stereo of the virtual environment. The mixed signal is output by a headset (or a combination of two or more speakers), thereby fully utilizing the spatial acoustic effect of the mixed signal. In this way, the auditory scene of the actual site can be reproduced, and not only can the listener clearly distinguish the direction, perspective, and movement trajectory of different acoustic objects, but also the listener can hear sounds with a more three-dimensional feel and spatial gradation due to the spatial acoustic effect, feel the sensation of being surrounded by sounds from all directions, and have an immersive auditory experience as if they were in the actual environment. In another implementation, a normal audio collection device may be used to collect real audio signals in the augmented reality scene.

[0025] A virtual audio signal refers to an audio signal constructed in a virtual environment on which an augmented reality scene is based. The virtual environment on which an augmented reality scene is based is a simulated environment generated entirely by a device. A corresponding virtual audio signal may be obtained by collecting sounds in the virtual environment. The sounds in the virtual environment are sounds constructed by a device. These sounds simulate the sounds made by objects in a real environment. The collected virtual audio signal may represent sounds in the constructed virtual world, such as the voice of a virtual character speaking in a virtual game scene, the sound of simulated flowing water, the sound of simulated birds singing, and the sound of simulated singing. The sounds in the virtual world may be stereo, which can increase the three-dimensionality of the sound. In other words, an audio signal constructed in a virtual environment can be understood as an audio signal obtained by collecting sounds in the virtual environment.

[0026] The computer device may collect real audio signals and virtual audio signals in the augmented reality scene in real time. The computer device may obtain real audio signals and virtual audio signals previously collected and stored in a database. For example, in an augmented reality scene in which a user plays a multiplayer game using an augmented reality device, the augmented reality device used by the user may collect the user's own speech and other sounds in the real physical environment in which the user is located in real time, and may also collect the speech of other players, thereby obtaining a real audio signal. It may also obtain a virtual audio signal by collecting sounds in the virtual game world. Subsequently, the real audio signal and the virtual audio signal may be mixed in a set target mixing mode. This allows the player to hear sounds with the corresponding mixing effect indicated by the mixing signal, providing a more immersive gaming experience.

[0027] In step S202, a plurality of mixing modes arranged for the augmented reality scene are determined, and one mixing mode is selected from the plurality of mixing modes as a target mixing mode.

[0028] The computing device can determine at least two mixing modes (i.e., multiple mixing modes) configured for the current augmented reality scene. In one implementation, the mixing modes configured for different augmented reality scenes may be different. This can be embodied in differences in the number of mixing modes and the specific mixing modes. For example, the mixing modes configured for AR scene S1 may include five mixing modes, while the mixing modes configured for AR scene S2 may include only three mixing modes. Also, for example, the mixing modes configured for AR scene S1 and the mixing modes configured for AR scene S2 each include three mixing modes, but only one mixing mode is the same. In other implementations, the mixing modes may be the same for different augmented reality scenes. That is, the same multiple mixing modes may be uniformly configured for different augmented reality scenes. In this way, the multiple mixing modes can be used for different augmented reality scenes. For example, five mixing modes may be uniformly configured and used for each augmented reality scene. As can be understood, the multiple mixing modes may be pre-configured for the augmented reality scenes. After arranging the multiple mixing modes for the augmented reality scene, the computer device may store a correspondence between the augmented reality scene and the multiple mixing modes in a storage space. Thus, when performing step S202, the computer device can determine the multiple mixing modes arranged for the augmented reality scene based on the correspondence. Here, the correspondence may include, but is not limited to, a scene identifier (e.g., a scene name) of the augmented reality scene and a mode identifier (e.g., a mode name, a mode symbol, etc.) of the multiple mixing modes arranged for the augmented reality scene.

[0029] The different mixing modes are intended to achieve different mixing effects between real audio signals and virtual audio signals. The mixing effect refers to the auditory perception of sound mixing and can reflect the volume relationship between virtual world sounds (i.e., sounds corresponding to virtual audio signals) and real world sounds (i.e., sounds corresponding to real audio signals). In one implementation, the multiple mixing modes configured for the augmented reality scene may include: (1) a real-off mixing mode (also called a real-world sound-off mode), (2) a virtual-off mixing mode (also called a virtual-world sound-off mode), (3) a strong virtual mixing mode (also called a weak real-world sound mode or weak real mixing mode), (4) a strong real mixing mode (also called a strong real-world sound mode or weak virtual mixing mode), and (5) a real-virtual sound equality mixing mode (also called a real-virtual sound equality mode). These mixing modes achieve different mixing effects. For example, in the real off-mixing mode (or the virtual off-mixing mode), a true audio signal corresponding to a sound in the real world (or a virtual audio signal corresponding to a sound in the virtual world) is not output, thereby achieving a mixing effect of blocking out the sound in the real world (or the sound in the virtual world). In the strong virtual mixing mode (i.e., weak virtual mixing mode), the final mixing effect is that the volume of the sound in the virtual world is greater than the volume of the sound in the real world. That is, the strong virtual mixing mode (weak virtual mixing mode) refers to a mixing mode for increasing the volume of the sound in the virtual world after mixing. In the strong virtual mixing mode (i.e., weak virtual mixing mode), the final mixing effect is that the volume of the sound in the real world is greater than the volume of the sound in the virtual world. That is, the strong virtual mixing mode (weak virtual mixing mode) refers to a mixing mode for increasing the volume of the sound in the real world after mixing. In equal mixing mode, the real-world sounds and virtual-world sounds heard after mixing are at equal levels (i.e., the volume of the real-world sounds and the volume of the virtual-world sounds are the same), and the user perceives the two sounds as being in harmony.In other words, the equal mixing mode is a mixing mode for making the volume of the real world sound and the volume of the virtual world sound equal after mixing.

[0030] In one embodiment, the true audio signal is obtained by collecting real-world sounds and may contain environmental noise that may affect subsequent processing of the true audio signal. Therefore, before performing the mixing process, the true audio signal may be subjected to noise reduction processing. This filters out environmental noise interference in the true audio signal as much as possible to obtain a noise-removed true audio signal. The noise-removed true audio signal may then be used in subsequent processing related to the virtual audio signal. Here, different mixing modes may be configured for the augmented reality scene, such as the above-mentioned real-world sound-off mode, weak real-world sound mode, real-virtual sound equality mode, strong real-world sound mode, and virtual-world sound-off mode. This provides selectable mixing modes to achieve desired mixing effects.

[0031] Specifically, after determining a plurality of mixing modes arranged for the augmented reality scene, one mixing mode may be selected from the plurality of mixing modes. The selected mixing mode may be used as a target mixing mode for mixing the real audio signal and the virtual audio signal. For example, the plurality of mixing modes may include mixing mode R1, mixing mode R2, and mixing mode R3. When mixing mode R3 is selected, mixing mode R3 may be set as the target mixing mode. The target mixing mode may be selected according to a user's needs in the augmented reality scene, or may be selected by automatically analyzing the sound fusion needs of the augmented reality scene.

[0032] In one implementation, the augmented reality scene includes an augmented reality device. The augmented reality device can be used to collect real audio signals and to help a computing device determine a target mixing mode. In one specific implementation, the target mixing mode can be determined by directly operating a physical key on the augmented reality device. In another specific implementation, the augmented reality device can be associated with an augmented reality application. Mode identifiers of multiple mixing modes arranged relative to the AR scene can be displayed in an application interface of the augmented reality application. One mixing mode can be selected as the target mixing mode by selecting the mode identifier in the application interface.

[0033] In step S203, a mixed signal is obtained by mixing the real audio signal and the virtual audio signal in the target mixing mode, and the mixed signal is output.

[0034] Different mixing modes also determine the mixing manner of the real audio signal and the virtual audio signal. The mixing manner specifically relates to the volume adjustment and mixing manner of the audio signal. The mixing signal obtained by performing the mixing process using the mixing manner indicated by the target mixing mode has a mixing effect corresponding to the target mixing mode. The mixing process of the real audio signal and the virtual audio signal can fuse the real audio signal and the virtual audio signal of different tracks into a single mixing signal. The mixing process can correspond to the fusion of real-world sounds and virtual-world sounds, and the corresponding mixing effect can correspond to the fusion effect of real-world sounds and virtual-world sounds.

[0035] In one implementation, after obtaining the mixed signal, the computing device can output the mixed signal in real time in the augmented reality scene. By mixing and outputting the real audio signal and the virtual audio signal in the AR scene, the listener can feel the sound effects associated with the virtual object in the virtual world, providing a more immersive and realistic auditory experience. Based on the above processing flow, the audio processing method provided in this application can not only occlude or unocclude the real-world sound source, but also perform corresponding processing on the real-world audio signal and the virtual-world audio signal, and finally fuse the two audio signals to achieve sound fusion.

[0036] The audio processing method provided in the embodiments of the present application can obtain a true audio signal and a virtual audio signal in an augmented reality scene. The true audio signal refers to an audio signal collected in a real environment on which the augmented reality scene is based, and the virtual audio signal refers to an audio signal constructed in a virtual environment on which the augmented reality scene is based. A plurality of mixing modes arranged for the augmented reality scene is determined, and one mixing mode from the plurality of mixing modes is selected as a target mixing mode. Here, different mixing modes can be used to achieve different mixing effects between the true audio signal and the virtual audio signal, resulting in different auditory experiences. By determining a plurality of mixing modes for the augmented reality scene, various options are provided for setting the target mixing mode, allowing the target mixing mode to be flexibly set. This can accommodate various scenes requiring mixing of the true audio signal and the virtual audio signal, and can also meet the user's customized mixing needs. A mixed signal is obtained by mixing the true audio signal and the virtual audio signal in the target mixing mode, and the mixed signal is output. The target mixing mode can achieve a corresponding mixing effect between real audio signals and virtual audio signals. By controlling the mixing process between real audio signals and virtual audio signals in the target mixing mode, the fusion of real-world sounds and virtual-world sounds on which the augmented reality scene is based can be achieved in the auditory dimension, thereby increasing the auditory fusion of sounds in the augmented reality scene and enriching the fusion capabilities of the augmented reality scene. The resulting mixing signal also has a mixing effect corresponding to the target mixing mode. Finally, the output of the mixing signal allows the user to more intuitively experience the fusion effect of sounds in the augmented reality scene, enhancing the user's immersion in the augmented reality scene.

[0037] Based on the above embodiment of the method shown in Fig. 2, the embodiment of the present application further proposes a more specific audio processing method. In the embodiment of the present application, the audio processing method is mainly described by taking the case where a computer device executes the audio processing method as an example. Referring to Fig. 4, the audio processing method may include the following steps S401 to S404.

[0038] In step S401, a real audio signal and a virtual audio signal in an augmented reality scene are obtained.

[0039] In step S402, a plurality of mixing modes arranged for the augmented reality scene are determined, and one mixing mode is selected from the plurality of mixing modes as a target mixing mode.

[0040] In an embodiment, when selecting one mixing mode from the plurality of mixing modes as a target mixing mode, the computing device may specifically perform the steps of obtaining a mode configuration operation, and selecting one mixing mode from the plurality of mixing modes as a target mixing mode according to the mode configuration operation.

[0041] The mode placement operation acquired by the computing device includes (1) a selection operation for multiple mode identifiers displayed in an augmented reality application, where one mode identifier corresponds to one mixing mode, or (2) a mode selection operation controlled by operating a physical key on the augmented reality device.

[0042] Regarding the mode configuration operation (1), the augmented reality application is an application (APP) for providing interaction operations that can affect an augmented reality scene. The augmented reality application may be installed on a computer device (e.g., a mobile terminal) or on another computer device (e.g., a mobile terminal connected to an AR headset) connected to a computer device that executes an audio processing method. Here, the connection method may be, but is not limited to, a wired connection (e.g., a connection via a data line) or a wireless connection (e.g., a Bluetooth connection). The augmented reality application may provide an application interface that displays mode identifiers corresponding to multiple mixing modes. Different mixing modes correspond to different mode identifiers, and each mode identifier supports selection. A selection operation for the multiple mode identifiers displayed in the augmented reality application may be a selection operation for a certain mode identifier in the application interface provided by the augmented reality application, such as a click operation on a certain mode identifier or a predetermined gesture operation. The selection operation can be used to indicate one mode identifier that represents the mixing mode selected by the user and that the user should use. Furthermore, the computer device can determine a mixing mode corresponding to the mode identifier from among a plurality of mixing modes based on the mode identifier in the selection operation and set it as a target mixing mode. In one possible embodiment, the user can also input a mode identifier into an application interface of an augmented reality application. Furthermore, the computer device can select the corresponding mixing mode as a target mixing mode based on the input mode identifier. Schematically, as shown in the schematic diagram of a scene for setting a target mixing mode shown in Figure 5a, the application interface of the augmented reality application is marked with 501.The application interface displays mode identifiers corresponding to the five mixing modes, and allows the user to check any of the mode identifiers. When the selection of any one of the mode identifiers is confirmed, the computer device can determine the mixing mode corresponding to the selected mode identifier as the target mixing mode.

[0043] In this manner, the augmented reality application can visually display the mode identifiers of the mixing modes to the user, giving the user the initiative to select a target mixing mode for the augmented reality scene. In this way, the user can select a mixing mode according to their own interest points without passively receiving the final mixing signal. This can better meet the user's customization needs for mixing in the augmented reality scene, without being limited to a fixed mixing mode.

[0044] Regarding the mode configuration operation (2), the augmented reality scene includes an augmented reality device. The augmented reality device is a computer device (e.g., an AR headset) for providing an interaction operation that can affect the augmented reality scene. The augmented reality device has a physical key. The physical key can be operated by a user to control the selection of a mixing mode. Optionally, each time the physical key is operated, a mode selection operation can occur, allowing the computer device to select a target mixing mode from multiple mixing modes. Schematically, the required target mixing mode can be manually configured using a button switch (a type of physical key) on the AR headset. Pressing the button switch on the AR headset once represents switching the current mixing mode, followed by obtaining a new mixing mode. In this manner, the mixing mode can be selected using a physical key provided on the augmented reality device, which can provide the user with a more realistic operation feel. As can be understood, if the mode configuration operation (1) is invalid, the mode selection operation using the physical key can be used as a backup method to select the target mixing mode. In an augmented reality scene, the above two mode configuration operations can both realize the determination of the target mixing mode, and can serve as backup methods for each other to ensure the user's normal selection of the mixing mode.

[0045] Although the above mixing mode selection methods are all manual methods provided for the user, in other embodiments, the target mixing mode can be automatically determined without the user's involvement, so that the selected mixing mode can better match the augmented reality scene to achieve a more appropriate mixing effect.When selecting one meeting mode from multiple mixing modes as the target mixing mode, the computer device may specifically perform the following steps (1) and (2) to automatically determine the target mixing mode according to the geographical location information of the real environment.

[0046] In step (1), geographical location information of a real environment on which the augmented reality scene is based is determined, and an environment type to which the real environment belongs is determined based on the geographical location information.

[0047] Geographical location information of a real environment can be used to describe the relative spatial relationships between geographical objects in the real environment. This geographical location information can be expressed in specific latitude and longitude coordinates or specific location names, and is not limited to this specification. For example, the geographical location information of a real environment can be determined by positioning the real environment in which the user is located using a Global Positioning System (GPS) included in the augmented reality device used by the user. For example, the geographical location information determined by GPS may be City A, District C, Street D, No. 58, XX Opera House. Because the geographical location information can describe the characteristics of the real environment in detail, the environment type to which the real environment belongs can be determined by analyzing the geographical location information. Depending on the environment attributes, the environment type of the real environment on which the augmented reality scene is based includes, but is not limited to, an indoor environment and an outdoor environment. Indoor environments include, for example, an opera house, a movie theater, a car driving scene, or a department store, and outdoor environments include, for example, a street, a playground, or a beach. For example, if the geographical location information is xx music hall, it can be determined that the environment type to which the real environment belongs is an indoor environment.

[0048] In step (2), one mixing mode is selected as a target mixing mode from among a plurality of mixing modes according to the sound fusion needs of the determined environment type.

[0049] Here, the sound fusion needs refer to the needs for achieving fusion of real-world sounds and virtual environment sounds in an augmented reality scene by mixing true audio signals with virtual audio signals, and may reflect the volume relationship between virtual world sounds (i.e., sounds corresponding to virtual audio signals) and real-world sounds (i.e., sounds corresponding to true audio signals). Sound fusion needs for different environment types may differ. For example, because indoor environments have less noise interference than outdoor environments, outdoor environments require a higher level of clarity for virtual world sounds and require suppression of real-world sounds. Indoor environments, real-world sounds may require less suppression, and the two may be equally fused. As can be understood, the computer device may pre-record information indicating sound fusion needs for each environment type in a storage space. Thus, when performing step (2), the computer device can determine the sound fusion needs for the corresponding environment type based on the pre-recorded information.

[0050] The sound fusion needs according to the corresponding environment type represent the sound fusion needs according to the real environment. The sound fusion needs can be used to indicate a specific fusion effect, and different mixing modes are intended to achieve different mixing effects. Therefore, based on the sound fusion needs according to the corresponding environment type, one appropriate mixing mode can be selected as the target mixing mode from among the mixing modes. For example, in an outdoor environment, a higher requirement for clarity in the virtual world sound may be met, so the strong virtual world sound mode may be selected for mixing processing. This makes the virtual world sound clearer and easier to hear.

[0051] In the above steps (1) and (2), the environment type to which the real environment belongs is automatically analyzed according to the geographical location information of the real environment, and a target mixing mode is automatically selected according to the sound fusion needs of the environment type. The entire process automatically and intelligently determines the target mixing mode for mixing the real audio signal and the virtual audio signal, and the final mixing effect can match the sound fusion needs of the corresponding environment type, thereby improving the immersive feeling and realism of sound fusion in the augmented reality scene.

[0052] In one embodiment, to enhance the auditory sensory experience in an augmented reality scene, sound effect processing may be performed on the real audio signal and the virtual audio signal before adjusting the volume of the audio signal. Here, sound effect processing includes, but is not limited to, one or more of reverberation processing and equalization processing. Here, reverberation processing can change the feel of the audio by adding a certain percentage of reflected sound and attenuation to the audio signal. Reverberation processing can increase the naturalness and three-dimensionality of the audio, making the audio sound richer and more gradated. Furthermore, the sound after reverberation processing can have a stronger sense of ambience. For example, reverberation can simulate natural effects in real environments such as rooms and halls. Equalization processing can adjust the timbre and spectral balance of the audio by changing the energy distribution of the audio signal at different frequencies, and can make the audio sound brighter, clearer, softer, or more dynamic by boosting or softening the volume of certain frequency bands. For example, if the bass in a song is not strong and needs to be increased, an equalization process can be used to adjust the frequency distribution of the audio and increase the low frequency portion, allowing the user to experience a richer musical experience. In specific implementations, the following methods 1 and 2 may be used to process the sound effects of audio signals in an augmented reality scene.

[0053] In the first method, the scene type of the augmented reality scene is determined, and sound effect processing is performed on the real audio signal and the virtual audio signal in the augmented reality scene according to the sound effect needs of the scene type.

[0054] Specifically, the scene type of an augmented reality scene may be determined based on a service associated with the augmented reality scene. That is, the computing device may determine the service associated with the augmented reality scene and then determine the scene type of the augmented reality scene based on a correspondence between the service and the scene type. For example, if the service associated with the augmented reality scene is a multiplayer online game, the scene type may be a game scene. Also, for example, if the service associated with the augmented reality scene is an online singing social service, the scene type may be a social scene. Different scene types have different sound effect needs. Therefore, sound effect processing may be performed on the acquired real audio signal and virtual audio signal to match the sound effect needs of the scene type. For example, if the scene type of the augmented reality scene is a game scene and the sound effect needs of the game scene require dynamic sound and ambience, reverberation processing may be performed on the real audio signal and virtual audio signal, and equalization processing may be performed with relatively small parameters, or no equalization processing may be performed. This results in higher reverberation and smaller equalization, thereby enhancing the ambient atmosphere of the game scene. As can be understood, the computer device may pre-record the sound effect needs for each scene type, so that after determining the scene type of the augmented reality scene, the computer device can determine the sound effect needs for the scene type of the augmented reality scene based on the pre-recorded information, and perform sound effect processing on each of the real audio signal and the virtual audio signal in the augmented reality scene according to the needs.

[0055] In method 2, the environment type to which the real environment on which the augmented reality scene is based is obtained, and sound effect processing is performed on each of the real audio signal and the virtual audio signal in the augmented reality scene according to the sound effect needs based on the environment type to which the real environment belongs.

[0056] Specifically, the environment type to which the real environment belongs may be determined based on the geographical location information of the real environment. The needs for sound effects for different environment types may differ. For example, when the environment type is an indoor environment, a sound with better three-dimensionality and gradation is required, whereas when the environment type is an outdoor environment, a sound with higher clarity and dynamism is required but the ambience may not be as strong. Therefore, the computer device may not need to perform reverberation processing on the real audio information signal and the virtual audio signal, but may instead perform equalization processing. For example, when the environment type is an indoor environment, a higher requirement for sound ambience may be present but the dynamism of the sound may be low. Therefore, the computer device may not need to perform equalization processing on the real audio signal and the virtual audio signal, but may instead perform reverberation processing to increase the ambience. As can be seen from this, the sound effect processing performed by the computer device can match the needs for sound effects depending on the environment type to which the real environment belongs. This ultimately achieves the sound effects required for the corresponding environment type.

[0057] In the above two methods, the sound effect processing for the audio signal in the augmented reality scene can be automatically determined according to the sound effect needs of the scene type of the augmented reality scene or the environment type to which the real environment on which it is based belongs, and the audio signal after the sound effect processing can be adapted to the corresponding sound effect needs and matched with the augmented reality scene, bringing a better experience.

[0058] In one embodiment, each mixing mode among the plurality of mixing modes has a mode identifier, which may be a mode number or a sequence number and can be used to uniquely identify the mixing mode. Each mode identifier corresponds to a set of equalization parameters required for equalization processing. The set of equalization parameters corresponding to any one mode identifier includes a first volume equalization parameter for the virtual audio signal in the corresponding mixing mode and a second volume equalization parameter for the real audio signal in the corresponding mixing mode. Different mixing modes may have different sets of equalization parameters, which are embodied in different equalization parameters. For example, the second volume equalization parameter in the set of equalization parameters for the strongly realistic mixing mode is larger than the second volume equalization parameter in the set of equalization parameters for the strongly virtual mixing mode. The volume equalization parameters can control the frequency of the audio signal to express different equalization effects for the audio signal. For example, the audio signal can express a low and powerful sound effect, a soft and rich sound effect, or a bright and clear sound effect.

[0059] Based on the set of equalization parameters configured for the mixing mode, the computing device may perform the following operations before mixing the real audio signal and the virtual audio signal in the target mixing mode to obtain a mixed signal: First, obtain a set of equalization parameters corresponding to the mode identifier of the target mixing mode. Since one mode identifier corresponds to one set of equalization parameters, the necessary set of equalization parameters is searched for based on the mode identifier of the target mode. In one implementation, the necessary set of equalization parameters can be obtained by searching a mapping data table. The mapping data table includes multiple mode identifiers and equalization parameters corresponding to each mode identifier. Next, the first volume equalization parameter in the obtained set of equalization parameters is used to equalize the virtual audio signal, and the second volume equalization parameter in the obtained set of equalization parameters is used to equalize the real audio signal. Schematically, if the target mixing mode is the augmented reality mixing mode, the computing device may obtain a set of equalization parameters for the augmented reality mixing mode for a clearer sound, and the first volume equalization parameter in the set of equalization parameters may be greater than the second volume equalization parameter. Thus, after the equalization process, the clarity of the real-world sound is enhanced. In one possible embodiment, an equalized virtual audio signal and an equalized true audio signal can be obtained through the above equalization process. Furthermore, when mixing the virtual audio signal and the true audio signal in the target mixing mode, the computer device may specifically mix the equalized virtual audio signal and the equalized true audio signal in the target mixing mode to obtain a final mixed signal.

[0060] In step S403, the volume of each of the real audio signal and the virtual audio signal is adjusted in the target mixing mode to obtain a real audio signal after volume adjustment and a virtual audio signal after volume adjustment.

[0061] A volume adjustment method for each of the real audio signal and the virtual audio signal can be determined by any one of the mixing modes. Different mixing modes can support different parameters for performing the corresponding volume adjustment. In the corresponding target mixing mode, a volume-adjusted real audio signal and a volume-adjusted virtual audio signal can be obtained by adjusting the volume. In terms of volume, the volume-adjusted real audio signal can be the same as or different from the originally acquired real audio signal, and the volume-adjusted virtual audio signal can be the same as or different from the originally acquired virtual audio signal. For example, when the target mixing mode is a strongly real-world sound mode, the volume of the real audio signal can be adjusted higher, and the volume of the virtual audio signal can be left unchanged or the volume of the virtual audio signal can be adjusted lower.

[0062] The volume of an audio signal refers to the volume of a sound subjectively perceived by a listener, and is also called loudness. The volume can be determined by the amplitude value (also called amplitude, width, or width) of the audio signal and the distance between the listener and the sound source. The larger the amplitude value, the louder the sound, and the smaller the distance between the listener and the sound source, the louder the sound. Therefore, when adjusting the volume of an audio signal, the amplitude value of the audio signal may be adjusted, or the distance between the augmented reality device for collecting real audio signals and the sound source may be adjusted, or the loudness of the audio signal may be directly adjusted. In an embodiment of the present application, a method of adjusting the volume of an audio signal by adjusting the amplitude value of the audio signal is used.

[0063] In one embodiment, the volume of an audio signal changes with a change in the amplitude value of the audio signal. The amplitude value of an audio signal is positively correlated with the volume of the audio signal, and the volume of the audio signal can increase with an increase in the amplitude value of the audio signal and can also decrease with a decrease in the amplitude value of the audio signal. When adjusting the volume of each of the real audio signal and the virtual audio signal in the target mixing mode, the computing device may specifically execute the contents described in (1) to (4) below. Furthermore, the volume adjustment of the real audio signal and the volume adjustment of the virtual audio signal can be realized using any of the logics described in (1) to (4) below. The following content describes the principle of adjusting the volume of an audio signal in an augmented reality scene, using one of the real audio signal and the virtual audio signal as an example.

[0064] (1) Scan the real audio signal and the virtual audio signal, and obtain the amplitude value constraint parameters of the current audio signal in the target mixing mode by using the scanned audio signal as the current audio signal.

[0065] The scanned audio signal may be a real audio signal or a virtual audio signal. If the current audio signal is a real audio signal, the volume of the real audio signal can be adjusted; if the current audio signal is a virtual audio signal, the volume of the virtual audio signal can be adjusted.

[0066] The amplitude value constraint parameter of the current audio signal is a parameter required to constrain the adjustment of the amplitude value of the current audio signal in the target mixing mode. Different mixing modes may have different amplitude value constraint parameters corresponding to the current audio signal. The amplitude value constraint parameter of the real audio signal in the target mixing mode may be the same as or different from the amplitude value constraint parameter of the virtual audio signal in the target mixing mode.

[0067] In one embodiment, the plurality of mixing modes include at least two of a strong real mixing mode, a strong virtual mixing mode, a peer mixing mode, a real off-mixing mode, and a virtual off-mixing mode, and an amplitude value constraint parameter of any one of the audio signals is positively correlated with a volume level of the corresponding audio signal after a volume adjustment, where the larger the amplitude value constraint parameter, the larger the volume level of the corresponding audio signal after a volume adjustment, and the smaller the amplitude value constraint parameter, the smaller the volume level of the corresponding audio signal after a volume adjustment, and where the any one of the audio signals may be a real audio signal or a virtual audio signal.

[0068] Different mixing modes have different magnitude relationships between the amplitude constraint parameters of the real audio signals and the amplitude constraint parameters of the virtual audio signals to achieve different mixing effects. Specifically, the following applies: (1) In the strong reality mixing mode, both the amplitude constraint parameters of the real audio signals and the amplitude constraint parameters of the virtual audio signals are valid, and the amplitude constraint parameters of the real audio signals are greater than the amplitude constraint parameters of the virtual audio signals. (2) In the strong virtual mixing mode, both the amplitude constraint parameters of the real audio signals and the amplitude constraint parameters of the virtual audio signals are valid, and the amplitude constraint parameters of the virtual audio signals are greater than the amplitude constraint parameters of the real audio signals. (3) In the peer mixing mode, both the amplitude constraint parameters of the real audio signals and the amplitude constraint parameters of the virtual audio signals are valid, and the amplitude constraint parameters of the virtual audio signals are equal to the amplitude constraint parameters of the real audio signals. (4) In the reality off-mixing mode, the amplitude constraint parameters of the real audio signals are invalid, and the amplitude constraint parameters of the virtual audio signals are valid. (5) In the virtual off-mixing mode, the amplitude constraint parameter of the real audio signal is a valid value, and the amplitude constraint parameter of the virtual audio signal is an invalid value. The valid value referred to above means a parameter value that enables valid output of an audio signal, for example, a numerical value greater than 0. The invalid value means a parameter value that disables output of an audio signal, for example, 0. If the amplitude constraint parameter of any one audio signal is an invalid value, it represents that the corresponding audio signal is not output.

[0069] For example, the amplitude value constraint parameters of any one of the audio signals may include a target amplitude value (also referred to as a target dB value) and a maximum gain value. In the strong realism mixing mode, the target dB value of the true audio signal is greater than the target dB value of the virtual audio signal, and the maximum gain value of the true audio signal is greater than the maximum gain value of the virtual audio signal, so that the volume of the virtual world sound is lower than the volume of the real world sound after final mixing. In the strong virtual mixing mode, the target dB value of the true audio signal is lower than the target dB value of the virtual audio signal, and the maximum gain value of the true audio signal is lower than the maximum gain value of the virtual audio signal, so that the volume of the virtual world sound is higher than the volume of the real world sound after final mixing. In the equal mixing mode, the target dB value of the true audio signal is equal to the target dB value of the virtual audio signal, and the maximum gain value of the true audio signal is equal to the maximum gain value of the virtual audio signal. This allows the real world sound and the virtual world sound to be heard at equal levels after final mixing, and the user perceives the two sounds as harmonious. In the above mixing modes, the corresponding amplitude constraint parameter values ​​are all valid, and the mixed signal obtained by mixing the audio signals after each volume adjustment can be effectively output to ensure the realization of the corresponding mixing effect. Meanwhile, in the real off-mixing mode (or the virtual off-mixing mode), the amplitude constraint parameter of the real audio signal (or the virtual audio signal) is invalid. For example, when the maximum gain value of the real audio signal (or the virtual audio signal) is 0, the corresponding input audio signal (the input signal corresponding to the real world sound off is the real audio signal, but the input signal corresponding to the virtual world sound off is the virtual audio signal) is multiplied by 0, that is, the corresponding audio signal is not output.

[0070] It can be seen that different mixing modes result in different amplitude value constraint parameters that cause different volume adjustment effects of audio signals in the augmented reality scene, and further result in different mixing effects. A target mixing mode may be selected from the above at least two mixing modes to meet a required mixing effect.

[0071] In one embodiment, the computing device may specifically perform the following steps 1.1 to 1.4 when obtaining the amplitude value constraint parameters of the current audio signal in the target mixing mode.

[0072] In step 1.1, a predetermined mapping table is obtained.

[0073] Here, the predetermined mapping table includes a plurality of mode identifiers and parameter sets mapped to each mode identifier. One mode identifier is used to identify one mixing mode. This mode identifier may be a mode value (e.g., a numerical sequence) or a mode name. The mode identifier can be used to uniquely identify the mixing mode. Different mode identifiers represent different mixing modes. Depending on the different mixing effects to be achieved in different mixing modes, the parameter sets mapped to different mode identifiers may also be different. Specifically, this may be embodied in different values ​​of amplitude constraint parameters included in the parameter sets. The mode identifiers and the parameter sets may have a one-to-one mapping relationship. The parameter set mapped to any one mode identifier includes a first amplitude constraint parameter of a virtual audio signal in the mixing mode identified by the corresponding mode identifier and a second amplitude constraint parameter of a real audio signal in the mixing mode identified by the corresponding mode identifier. The first amplitude constraint parameter and the second amplitude constraint parameter may include the same type of parameter, but the specific values ​​of the parameters may be different. Furthermore, the mixing effect required for the mixing mode can be realized. For a specific format of the predetermined mapping table, please refer to the example shown in Table 1 below.

[0074] [Table 1]

[0075] As described above, when the mixing mode is the real off-mixing mode or the virtual off-mixing mode, the maximum gain value is 0, but the target dB value required for volume control is not set. This is because when the maximum gain value is 0, it is not necessary to set the target dB value. The reason for this is that by directly amplifying the corresponding audio signal with the maximum gain value, the audio signal is not output in the final mixed signal, thereby realizing the sound off represented by the corresponding audio signal. The off effect here is consistent with the sound blocking effect.

[0076] In step 1.2, a predetermined mapping table is searched for the mode identifier of the target mixing mode, thereby obtaining a group of parameters that are mapped to the mode identifier of the target mixing mode.

[0077] In a specific implementation, the computer device may search a predetermined mapping table for a mode identifier that is the same as the mode identifier of the target mixing mode, and then set a parameter set mapped to the searched mode identifier as a parameter set mapped to the mode identifier of the target mixing mode. Thus, the obtained parameter set includes a first amplitude constraint parameter of the virtual audio signal in the target mixing mode and a second amplitude constraint parameter of the real audio signal in the target mixing mode. Furthermore, as shown in the following steps 1.3 and 1.4, the computer device may select a matching amplitude constraint parameter from the parameter set according to whether the current audio signal is a real audio signal or a virtual audio signal.

[0078] In step 1.3, if the current audio signal is a virtual audio signal, a first amplitude value constraint parameter is selected from the obtained parameter group as the amplitude value constraint parameter of the current audio signal in the target mixing mode.

[0079] If the current audio signal is a virtual audio signal, it indicates that the target of the current volume adjustment process is a virtual audio signal, and in order to realize the volume adjustment of the virtual audio signal, the first amplitude value constraint parameter in the acquired parameter group may be set as the amplitude value constraint parameter of the current audio signal in the target mixing mode.

[0080] In step 1.4, if the current audio signal is a true audio signal, a second amplitude value constraint parameter is selected from the obtained parameter group as the amplitude value constraint parameter of the current audio signal in the target mixing mode.

[0081] If the current audio signal is a real audio signal, it indicates that the current volume adjustment processing target is the real audio signal, and in order to realize the volume adjustment of the real audio signal, the second amplitude value constraint parameter in the acquired parameter group may be set as the amplitude value constraint parameter of the current audio signal in the target mixing mode.

[0082] As can be understood, the order in which steps are included in this application does not necessarily limit the order of execution. For example, step 1.3 or step 1.4 may be executed in parallel or may be executed in a reverse order. Furthermore, the terms "first," "second," and the like in this application are used to distinguish between identical or similar items having substantially the same roles and functions. It should be understood that there is no logical or chronological dependency between "first," "second," and "nth," and the quantities and execution orders are not limited.

[0083] In the method for obtaining amplitude value constraint parameters of the current audio signal shown in steps 1.1 to 1.4 above, a parameter set corresponding to each mixing mode can be provided using a predetermined mapping table. The mode identifiers of the mixing modes are unique and can be used to distinguish different mixing modes. Therefore, a parameter set mapped to the mode identifier of the target mixing mode can be quickly and accurately searched for in the predetermined mapping table using the mode identifier of the target mixing mode. Obtaining the amplitude value constraint parameters required for the current audio signal through table lookup improves processing efficiency and ensures the effectiveness and accuracy of subsequent processing.

[0084] (2) Obtaining a level description value of the current audio signal by monitoring the signal level of the current audio signal, the level description value of the current audio signal is for indicating the amplitude value of the current audio signal.

[0085] In order to obtain a purer signal, the current audio signal may be subjected to noise reduction processing to obtain a noise-reduced current audio signal, and then the signal level of the noise-reduced current audio signal may be monitored. The general logic for signal level monitoring may be as follows: The level of the current audio signal can be obtained by signal level monitoring. Here, the level refers to the magnitude of the sample point width value (i.e., the amplitude value of the sampling point, also called the sample point amplitude value) of the current audio signal, and can be described in units of dB or the magnitude of the amplitude value. When using dB as the unit, the sample point width value needs to be converted by the following formula:

number

[0086] Here, dB represents the amplitude value of a certain sampling point in the current audio signal, which can be called level, Sample represents the amplitude value of one sampling point, and |Sample| represents the absolute value of the sample point amplitude value.

[0087] In one embodiment, when executing (2) above, the computer device does not need to frame the current audio signal, but can perform signal level monitoring processing using the entire current audio signal as a unit, and can use the amplitude value of each sampling point contained in the current audio signal as the level description value.

[0088] In another embodiment, since the audio signal is a constantly hopping signal and volume control requires a relatively smooth level description value that can represent the signal amplitude level, a level description value of the current audio signal may be acquired based on the acquired level. The computer device may monitor the signal level of the current audio signal by performing (2) above, and may perform signal level monitoring using an audio frame as the smallest unit when acquiring the level description value of the current audio signal. A specific flow includes the contents shown in steps 2.1 to 2.4 below.

[0089] In step 2.1, a framing process is performed on the current audio signal to obtain K audio frames (K is an integer greater than 1) of the current audio signal.

[0090] In a specific implementation, the current audio signal may be framed using a predetermined frame length. The frame length of any one of the audio frames thus obtained is equal to or less than the predetermined frame length. For example, if the predetermined frame length is 5 milliseconds (ms) and the current audio signal is a 3-second (s) real audio signal, 60 audio frames can be obtained by performing framing processing on it. To ensure the relationship between the framed audio frames, the framed audio frames may be framed with a predetermined frame shift in addition to referring to the predetermined frame length. This frame shift refers to the overlap time (e.g., 2 ms) between two adjacent audio frames. For example, as shown in the schematic diagram of the framing processing of the current audio signal in FIG. 5b, the frame length of each framed audio frame is 20 ms, but the frame shift is 5 ms.

[0091] In step 2.2, K audio frames are scanned to determine the i-th (i∈[1,K]) audio frame currently being scanned.

[0092] In a specific implementation, the K audio frames may be sequentially arranged according to the playback order in the current audio signal. During scanning, the K audio frames may be sequentially scanned according to the playback order of the audio frames in the audio signal. Furthermore, sequential processing of the K audio frames is realized. The playback order here can be understood as the arrangement position of the audio frames in the audio signal and can be expressed by the frame order of the audio frames. The i-th audio frame currently being scanned can be understood as the currently processed audio frame. For example, if the real audio signal is divided into 100 audio frames, with frame orders of 1, 2...100, the real audio signal can be completely presented by sequentially playing the audio frames of the corresponding audio signal according to the frame order. The i-th audio frame is, for example, the 10th audio frame of the K audio frames.

[0093] In step 2.3, the signal level of the i-th audio frame is monitored to obtain the level description value of the i-th audio frame, or an audio validity determination process is performed on the i-th audio frame, and if it is determined that the i-th audio frame is a valid audio frame, the signal level of the i-th audio frame is monitored to obtain the level description value of the i-th audio frame.

[0094] In one specific implementation, the level description value of the i-th audio frame may be obtained by directly monitoring the signal level of the i-th audio frame. In another implementation, to avoid invalid processing, a voice activity determination process may first be performed on the i-th audio frame to determine whether the i-th audio frame is a valid audio frame or an invalid audio frame. If the i-th audio frame is determined to be a valid audio frame, signal level monitoring is performed. Here, a valid audio frame is a valid voice signal, and an invalid audio frame is an unnecessary voice signal or noise signal. The voice activity determination process may specifically use voice activity detection (VAD). VAD recognizes and removes long silent periods from a voice signal stream, separating valid voice signals from unnecessary voice signals or noise signals, thereby making subsequent processing of the audio signal more efficient. A vad result of 1 indicates that the i-th audio frame is a valid audio frame, i.e., a valid voice signal, and a vad result of 0 indicates that the i-th audio frame is an invalid audio frame, i.e., an invalid voice signal. Subsequently, signal level monitoring is performed only on audio frames in the current audio signal whose vad result is 1, and signal level monitoring is not performed on audio frames whose vad result is 0. This method realizes validity determination of the current audio signal, determines valid and invalid audio frames in the current audio signal, and further selects valid audio signals in the current audio signal and performs signal level monitoring on them, ensuring the validity of audio processing and saving processing resources required for subsequent processing.

[0095] In one embodiment, each audio frame includes a plurality of sampling points and an amplitude value of each sampling point. A specific implementation of the step of the computer device obtaining the level description value of the i-th audio frame by monitoring the signal level of the i-th audio signal may include first calculating an average value of the amplitude values ​​of each sampling point in the i-th audio frame to obtain an average signal amplitude value of the i-th audio frame, and then smoothing the average signal amplitude value of the i-th audio frame to obtain the level description value of the i-th audio frame.

[0096] In a specific implementation of the average value calculation, the amplitude values ​​of each sampling point in the i-th audio frame are added to obtain a sum, and then averaged based on this sum and the number of sampling points in the i-th audio frame to obtain the signal amplitude average value of the i-th audio frame (i.e., the average value of the sample point amplitude values ​​within the audio frame, also referred to as the frame amplitude average value). This signal amplitude average value can be expressed in units of amplitude value or dB (decibels). For example, if the i-th audio frame is a 5-millisecond (ms) audio signal frame and the signal amplitude average value of the i-th audio frame is expressed in terms of amplitude value, the specific expression may be as follows:

number

[0097] where:

number

number

[0098] When written in dB units, the specific expression is as follows:

number

[0099] where:

number

number

[0100] After obtaining the average signal amplitude value of the i-th audio frame, in one implementation, the average signal amplitude value of the i-th audio frame may be directly used as the level description value of the i-th audio frame. In another implementation, the level description value of the i-th audio frame may be obtained by smoothing the average signal amplitude value of the i-th audio frame to make the volume change of the audio signal smoother and reduce sudden peak changes.

[0101] In the signal level monitoring realized by the above method, the average value of the amplitude values ​​of the sample points in one audio frame is calculated, and then the calculated average value is directly used as the level description value, or the calculated average value can be smoothed and used as the level description value. The average value can well represent the amplitude value level of the audio signal, and the smoothing process can further improve the smoothness. The obtained level description value is advantageous for supporting subsequent processing of the current audio signal.

[0102] Since the i-th audio frame may or may not have undergone the audio validity determination process, the specific logic for the smoothing process can be referred to in the following a to c.

[0103] a. Find a reference audio frame for the i-th audio frame from the K audio frames.

[0104] The reference audio frame is an audio frame located before the i-th audio frame among the K audio frames. Also, depending on whether the i-th audio frame has undergone active voice detection processing, different types of reference audio frames are required. Specifically, if the i-th audio frame has not undergone active voice detection processing, the reference audio frame includes the (i-1)th audio frame. For example, if the i-th audio frame is the second audio frame, the reference audio frame may include the first audio frame. If the i-th audio frame has undergone active voice detection processing, the reference audio frame includes a valid audio frame located before the i-th audio frame and closest to the i-th audio frame. The i-th audio frame is a valid audio frame determined by the active voice detection processing, and the valid audio frame closest to the i-th audio frame may or may not be adjacent to the i-th audio frame. This depends on whether the adjacent audio frame before the i-th audio frame is a valid audio frame. For example, if the i-th audio frame is the third audio frame, and the first audio frame is a valid audio frame but the second audio frame is an invalid audio frame, the reference audio frame includes the first audio frame. However, regardless of whether the reference audio frame has undergone the voice activity determination process, the reference audio frame is a frame that has undergone the signal level monitoring process. Based on the position of the i-th audio frame among the K audio frames, the computer device may or may not find the reference audio frame. Based on this, the determination of the level description value may be divided into the following two cases, b and c.

[0105] b) If no reference audio frame is found, the average signal amplitude value of the i-th audio frame is used as the level description value of the i-th audio frame. Specifically, if no reference audio frame is found, it indicates that the i-th audio frame is the first audio frame or the first valid audio frame among the K audio frames, and the average signal amplitude value of the i-th audio frame can be directly used as the level description value of the i-th audio frame.

[0106] c. If a reference audio frame is found, the level description value of the reference audio frame is used to smooth the average signal amplitude value of the i-th audio frame to obtain the level description value of the i-th audio frame. Specifically, if a reference audio frame is found, it indicates that the i-th audio frame is not the first audio frame among the K audio frames, and is not the first valid audio frame. Furthermore, the level description value of the reference audio frame can be used to smooth the average signal amplitude value of the i-th audio frame currently being processed to obtain the level description value of the i-th audio frame.

[0107] In one specific implementation, the computer device may obtain the level description value of the i-th audio frame by smoothing the average signal amplitude value of the i-th audio frame using the level description value of the reference audio frame, specifically by the following method: Obtain a first smoothing coefficient and a second smoothing coefficient required for the smoothing process, where the first smoothing coefficient is greater than the second smoothing coefficient. The first smoothing coefficient and the second smoothing coefficient may be set based on empirical values. For example, both the first smoothing coefficient and the second smoothing coefficient are set to values ​​within 1. Exemplarily, the first smoothing coefficient is denoted as a (also referred to as a long-term smoothing coefficient), and the second smoothing coefficient is denoted as b (also referred to as a short-term smoothing coefficient), and both the first smoothing coefficient and the second smoothing coefficient are values ​​less than 1, for example, a=0.95 and b=0.4. In another implementation, the first smoothing coefficient and the second smoothing coefficient may be determined based on the waviness of the i-th audio frame. For example, the first smoothing coefficient may be increased as the undulation of the i-th audio frame indicated by the undulation degree of the i-th audio frame increases.

[0108] The long-term smoothing process is performed on the average signal amplitude value of the i-th audio frame using a first smoothing coefficient and the level description value of the reference audio frame to obtain the average signal amplitude value after the long-term smoothing process. Specifically, long-term smoothing can be understood as smoothing using a large smoothing coefficient. The long-term smoothing process can smooth out changes in the volume of the audio signal over a relatively long time period and ensure consistent amplitude levels between different parts of the entire audio signal. The long-term smoothing process may be performed by weighting and adding the level description value of the reference audio frame and the average signal amplitude value of the i-th audio frame using the first smoothing coefficient. In a specific implementation, the first smoothing coefficient may be first set as the weight of the reference audio frame, and the difference between the reference value and the first smoothing coefficient may be set as the weight of the i-th audio frame. Then, the long-term smoothing process may be performed by weighting and adding the level description value of the reference audio frame and the average signal amplitude value of the i-th audio frame based on the weight of the reference audio frame and the weight of the i-th audio frame to obtain the average signal amplitude value after the long-term smoothing process. Here, the reference value is a value (e.g., 1) that restricts the maximum value of the smoothing coefficient. In a specific process of weighted addition, the average signal amplitude value after long-term smoothing may be obtained by adding the product of the weight of the reference audio frame and the level description value of the reference audio frame and the product of the weight of the i-th audio frame and the average signal amplitude value of the i-th audio frame. The specific calculation formula is as follows:

number

[0109] where:

number

number

number

[0110] Similarly, the signal amplitude average value after short-term smoothing may be obtained by performing short-term smoothing on the signal amplitude average value of the i-th audio frame using a second smoothing coefficient and the level description value of the reference audio frame. Here, short-term smoothing can be understood as smoothing using a small smoothing coefficient. The short-term smoothing can smooth out changes in the volume of the audio signal within a relatively short time period, ensuring that excessive peaks or distortions do not occur in the audio signal. The signal amplitude average value after short-term smoothing may be obtained by weighting and adding the level description value of the reference audio frame and the signal amplitude average value of the i-th audio frame using the second smoothing coefficient. In a specific implementation, the second smoothing coefficient may first be set as the weight of the reference audio frame, and the difference between the reference value and the second smoothing coefficient may be set as the weight of the i-th audio frame. Then, the level description value of the reference audio frame and the signal amplitude average value of the i-th audio frame are weighted and added based on the weight of the reference audio frame and the weight of the i-th audio frame to obtain the signal amplitude average value after short-term smoothing. Here, the reference value is a value (e.g., 1) that restricts the maximum value of the smoothing coefficient. In a specific process of weighted addition, the average signal amplitude value after short-term smoothing may be obtained by adding the product of the weight of the reference audio frame and the level description value of the reference audio frame and the product of the weight of the i-th audio frame and the average signal amplitude value of the i-th audio frame. The specific calculation formula is as follows:

number

[0111] where:

number

number

number

number

[0112] Finally, the level description value of the i-th audio frame is obtained by weighting and adding the average signal amplitude value after the long-term smoothing process and the average signal amplitude value after the short-term smoothing process. Here, the average signal amplitude value after the long-term smoothing process can be abbreviated as the long-term smoothed value, and the average signal amplitude value after the short-term smoothing process can be abbreviated as the short-term smoothed value. Furthermore, a corresponding weighting coefficient can be used during the weighting and addition, and its specific expression is as follows:

number

[0113] where:

number

number

number

[0114] Optionally, the above-mentioned short-term smoothing process may be replaced with a medium- to short-term smoothing process. The medium- to short-term smoothing process can be understood as a smoothing process interposed between the long-term smoothing process and the short-term smoothing process, and the smoothing coefficient used therein is also interposed between the smoothing coefficients used in the above-mentioned two types of smoothing processes. For example, the smoothing coefficient used in the short-term smoothing process is 0.3, the smoothing coefficient used in the long-term smoothing process is 0.9, and the smoothing coefficient used in the medium- to short-term smoothing process is 0.7. In this case, in the specific process of the above-mentioned smoothing process, the medium- to short-term smoothing process can obtain an average signal amplitude value after the medium- to short-term smoothing process, and further, the level description value of the i-th audio frame can be obtained by weighting and adding the average signal amplitude value after the long-term smoothing process.

[0115] The long-term smoothing process and short-term smoothing process (or medium- to short-term smoothing process) for the i-th audio frame described above can ensure that the change in the amplitude value of the audio frame in a long time is smoother, and can also ensure that the change in a short time is more gradual, so that the finally obtained level description value is a relatively smooth value, which is advantageous for the subsequent amplification process of the audio signal.

[0116] In another specific implementation, when smoothing the average signal amplitude value of the i-th audio frame, a smoothing coefficient may be obtained. Based on the i-th audio frame, H audio frames required for long-term smoothing (i.e., from the frame immediately preceding the i-th audio frame to the H frames immediately preceding the i-th audio frame) and F audio frames required for short-term smoothing (i.e., from the frame immediately preceding the i-th audio frame to the F frames immediately preceding the i-th audio frame) may be sequentially selected toward the first frame, where H>F, and both H and F are positive integers. For example, if the i-th audio frame is the fourth audio frame, the immediately preceding frame (i.e., the third audio frame) may be used as the reference audio frame for short-term smoothing, and the previous three frames (including the third, second, and first audio frames) may be used as the reference audio frames for long-term smoothing. Next, the smoothing coefficient and H audio frames are used to perform long-term smoothing of the average signal amplitude value of the i-th audio frame to obtain the average signal amplitude value after the long-term smoothing process. The smoothing coefficient and F audio frames are used to perform short-term smoothing of the average signal amplitude value of the i-th audio frame to obtain the average signal amplitude value after the short-term smoothing process. In this method, the long-term smoothing process can be understood as a smoothing process using a large number of audio frames, and the short-term smoothing process can be understood as a smoothing process using a small number of audio frames. The long-term smoothing process refers to a larger number of previous audio frames to adjust the average signal amplitude value of the current i-th audio frame, while the short-term smoothing process can refer to a smaller number of previous audio frames to adjust the average signal amplitude value of the current i-th audio frame. This improves the smoothness of the audio signal represented by the audio frames.

[0117] In step 2.4, by continuously scanning until all K audio frames are scanned, the level description value of the current audio signal is obtained, and the level description value of the current audio signal includes the level description values of each monitored audio frame.

[0118] When the processing of the i-th audio frame currently being scanned is completed, scanning may continue. Thereby, for the scanned audio frames, processing similar to that of the i-th audio frame is performed. When directly performing signal level monitoring, each monitored audio frame includes K audio frames, and further, the level description value of the current audio signal (that is, including the level description values of K audio frames) is obtained. When performing voice validity discrimination on the audio frame, since each monitored audio frame is a valid audio frame, the level description value of the current audio signal includes the level description values of J (J≦K) valid audio frames. When J<K, it indicates that there are invalid audio frames in the K audio frames of the current audio signal. The invalid audio frames may be discarded, and the discarding method may be deletion or blocking. When J=K, it indicates that all of the K audio frames of the current audio signal are valid audio frames, so the obtained level description value of the current audio signal includes the level description values of K audio frames.

[0119] In the above steps 2.1 to 2.4, the current audio signal is framed, and by directly performing signal level monitoring on each framed audio frame, the level description value of each audio frame is quickly obtained. The level description value may also be obtained by performing signal level monitoring on the valid audio frames after performing the voice validity discrimination process. Thereby, subsequent valid processing is guaranteed, and wasteful consumption of processing resources is avoided.

[0120] (3) Based on the acquired amplitude value constraint parameters and the level description value of the current audio signal, a control process for the gain required for the amplitude value of the current audio signal is performed to acquire the amplitude value gain of the current audio signal.

[0121] In one implementation, the amplitude value constraint parameters of the current audio signal in the target mixing mode (i.e., the acquired amplitude value constraint parameters) include a target amplitude value and an amplitude value gain threshold. The target amplitude value refers to an expected amplitude value (which may be in dB or amplitude magnitude) that should be achieved by the amplitude value adjustment of the current audio signal in the target mixing mode, and is also referred to as a target dB value. The amplitude value gain threshold is an upper limit value for constraining the gain required for volume adjustment of the current audio signal in the target mixing mode, and is also referred to as a maximum gain value because it can be used to limit the gain required to reach the target amplitude value and avoid unreasonable volume adjustment.

[0122] In one specific implementation, based on the above amplitude value constraint parameters, the computing device may perform gain prediction for the entire current audio signal based on the target amplitude value and the level description value of the current audio signal, to obtain a predicted amplitude value gain required for the amplitude value of the current signal to reach the target amplitude value, and select the smaller of the predicted amplitude value gain and the obtained amplitude value gain threshold as the amplitude value gain of the current audio signal.

[0123] In another specific implementation, the current audio signal is divided into K audio frames. Based on the aforementioned signal level monitoring of the current audio signal being performed in units of audio frames, optionally, before the signal level monitoring, it may be determined whether the currently processed audio frame is a valid audio frame. Furthermore, if the audio frame is a valid audio frame, a level description value is obtained. Therefore, the level description value of the current audio signal includes level description values ​​of J audio frames that have undergone signal level monitoring among the K audio frames, where K is an integer greater than 1 and J≦K. Based on this, the computer device can determine, in units of audio frames, a frame amplitude value gain for each audio frame that has undergone signal level monitoring. Specifically, the method includes the following steps 3.1 to 3.4.

[0124] In step 3.1, J audio frames according to the level description values ​​of the current audio signal are scanned to determine the j-th (j∈[1,J]) audio frame currently being scanned.

[0125] Of the K audio frames, the J audio frames for which signal level monitoring has been performed may be arranged according to the playback order. Thus, the computer device scans the J audio frames according to the arrangement positions of the audio frames in the audio signal, and the j-th audio frame to be scanned is one audio frame to be currently processed, and the j-th audio frame has undergone signal level monitoring and has a corresponding level description value. In one embodiment, when J = K, the j-th audio frame can correspond to the i-th audio frame for which the above-described signal level monitoring scan is performed. Exemplarily, after the corresponding level description value is obtained through level monitoring for the first scanned audio frame described above, the necessary gain control processing may be directly performed. In other embodiments, it is determined whether to perform level monitoring by voice validity discrimination processing, and when J < K, the j-th audio frame is a valid audio frame. The j-th audio frame does not necessarily correspond to the i-th audio frame for which the above-described signal level monitoring scan is performed. The J audio frames are a plurality of audio frames in a new arrangement order obtained by selecting and removing invalid audio frames from the K audio frames. Exemplarily, the second audio frame among the K audio frames is actually the first audio frame among the J audio frames. Regarding the relationship between the K audio frames and the J audio frames, and the relationship between the i-th scanned audio and the j-th audio frame, as shown in FIG. 5c, after passing through the voice validity discrimination processing, the first audio frame among the K audio frames is determined as an invalid audio frame, thereby obtaining J audio frames that are valid audio frames.

[0126] In step 3.2, a predicted amplitude value gain is obtained by performing gain prediction using the target amplitude value in the obtained amplitude value constraint parameter and the level description value of the j-th audio frame.

[0127] The predicted amplitude value gain refers to the gain required to adjust the level description value of the jth audio frame to the target amplitude value. In a specific gain prediction process, to ensure the validity of the calculation, the target amplitude value and the level description value of the jth audio frame must be in the same unit. If the unit of the target amplitude value is dB, the unit of the level description value must also be dB. If the unit of the target amplitude value is the magnitude of the amplitude value, the unit of the level description value must also be the magnitude of the amplitude value. If the unit is dB, the difference between the target amplitude value and the level description value of the jth audio frame may be first calculated based on the conversion relationship between the level and the amplitude value, and then the predicted amplitude value gain may be obtained based on the calculated difference. A specific expression may be as follows:

number

[0128] where:

number

number

number

number

[0129] If the units are both amplitude values, the level description value of the jth audio frame can be used as the denominator and the target amplitude value as the numerator to directly calculate the ratio between them, and the calculated ratio can be used as the predicted amplitude value gain. The specific expression is as follows:

number

[0130] where:

number

number

[0131] In step 3.3, the smaller of the amplitude gain threshold in the obtained amplitude constraint parameter and the predicted amplitude gain is determined as the frame amplitude gain required to adjust the amplitude of the j-th audio frame.

[0132] To avoid anomalies caused by excessively large amplitude value adjustment of an audio frame, a final constraint may be imposed by an amplitude value gain threshold. If the calculated predicted amplitude value gain is greater than the amplitude value gain threshold, the amplitude value gain threshold may be used as the frame amplitude value gain required to adjust the amplitude value of the jth audio frame. If the calculated predicted amplitude value gain is less than the amplitude value gain threshold, the predicted amplitude value gain threshold may be used as the frame amplitude value gain required to adjust the amplitude value of the jth audio frame. Of course, if the two are equal, either one of them may be used as the frame amplitude value gain required to adjust the amplitude value of the jth audio frame. The frame amplitude value gain referred to above is the amplitude value gain required to adjust the amplitude value of the audio frame for the audio frame. A specific calculation formula may be as follows:

number

[0133] where:

number

number

[0134] As can be understood, for any of the J audio frames, the corresponding frame amplitude value gain can be determined in the above manner, but the frame amplitude value gains of different audio frames may be the same or different. For example, the frame amplitude value gains of the first audio frame and the second audio frame are both the amplitude value gain threshold. For example, the frame amplitude value gain of the first audio frame is the amplitude value gain threshold, while the frame amplitude value gain of the third audio frame is the predicted amplitude value gain.

[0135] In step 3.4, the amplitude value gain corresponding to the current audio signal is obtained by continuing scanning until all J audio frames have been scanned.

[0136] The computer device may continue to scan the audio frame next to the current audio frame and perform processes similar to steps 3.2 and 3.3 on the next audio frame as a new current audio frame to obtain a frame amplitude value gain for the audio frame. By continuously scanning, a frame amplitude value gain corresponding to each audio frame for which signal level monitoring has been performed can be obtained, and the amplitude value gain of the current audio signal can be further obtained. That is, the amplitude value gain of the current audio signal includes a frame amplitude value gain corresponding to each of the J audio frames. For example, the current audio signal includes 100 frames, 80 of which have undergone signal level monitoring. In this case, the amplitude value gain of the current audio signal may include a frame amplitude value gain corresponding to each of the 80 audio frames. The determined frame amplitude value gains can be used to adjust the amplitude value of the audio frame. Subsequently, the volume of the current audio signal is adjusted.

[0137] In the method shown in steps 3.1 to 3.4 above, an initial amplitude value gain is first predicted according to the target amplitude value and the level description value of the audio frame, and then an amplitude value gain threshold is used to constrain the predicted initial amplitude value gain, thereby realizing the control of the frame amplitude value gain required for the audio frame and ensuring the reasonable adjustment of the amplitude value of the subsequent audio frame.

[0138] (4) Using the amplitude value gain of the current audio signal, the amplitude value of the current audio signal is amplified to adjust the volume of the current audio signal, and a volume-adjusted current audio signal is obtained.

[0139] In one embodiment, the computing device may use the entire current audio signal as a unit. The amplitude value gain of the current audio signal may include one amplitude value gain. Furthermore, the amplitude value gain of the current audio signal may be used to amplify the amplitude value of each sampling point included in the current audio signal, thereby adjusting the volume of the current audio signal and obtaining a volume-adjusted current audio signal. In this manner, the volume-adjusted current audio signal includes each sampling point and the amplified amplitude value of each sampling point. The amplitude value adjustment for each sampling point in the current audio signal is unified, which allows for quick adjustment of the amplitude value of the current audio signal and improves the efficiency of volume adjustment.

[0140] In another embodiment, the current audio signal is framed into K audio frames (K is an integer greater than 1), that is, at least two audio frames can be obtained by framing the current audio signal. The amplitude value gain of the current audio signal includes frame amplitude value gains of J audio frames (J≦K) among the K audio frames for which signal level monitoring has been performed. When performing the content described in (4) above, the computer device may specifically perform the following steps 4.1 to 4.5.

[0141] In step 4.1, J audio frames according to the level description values ​​of the current audio signal are scanned to determine the j-th (j∈[1,J]) audio frame currently being scanned.

[0142] In a specific implementation, the J audio frames may be arranged in a playback order. Thus, the computer device scans the J audio frames in the playback order of the audio frames in the audio signal, with the jth scanned audio frame being the audio frame currently to be processed. For example, the first audio frame of the J audio frames may be used as the starting point for scanning, and the jth audio frame may be scanned in the direction of increasing audio frame position. The jth audio frame has undergone signal level monitoring and has a corresponding level description value. Then, the J audio frames are subjected to necessary gain control according to the content introduced in (3) above to obtain corresponding frame amplitude value gains.

[0143] In one implementation, each audio frame includes multiple sampling points and amplitude values ​​for each sampling point. The computing device may obtain an amplified amplitude value for the jth audio frame by amplifying the amplitude value for each sampling point of the jth audio frame using a frame amplitude gain for the jth audio frame. Specifically, the amplitude value for each sampling point of the jth audio frame may be multiplied by the frame amplitude gain for the jth audio frame. The result is an amplified amplitude value. A volume-adjusted jth audio frame corresponding to the jth audio frame is generated based on each sampling point and its amplified amplitude value. In this method, adjusting the amplitude value for each audio frame is a finer volume adjustment method than adjusting the amplitude value for the entire current audio signal. Appropriate adjustments can be made according to the characteristics of the audio frames to achieve a natural-sounding volume-adjusted current audio signal.

[0144] In other implementations, to avoid noise in the sound due to sudden changes in the audio signal adjustment process, gradual gain adjustments may be made to the currently processed audio frame with smaller granularity (e.g., 1 ms length), as shown in steps 4.2 to 4.4 below.

[0145] In step 4.2, the j-th audio frame is divided into N audio sub-blocks (N is an integer greater than 1), and an amplitude value of each of the N audio sub-blocks is determined.

[0146] Specifically, the jth audio frame may be first divided. The division process divides the jth audio frame into audio sub-blocks (also called sub-blocks or audio sub-frames) with smaller frame lengths. The amplitude value of the audio frame is gradually adjusted for each audio sub-block, thereby enabling smoother changes in the volume of the audio signal. In one implementation, the jth audio frame may be divided using a predetermined sub-block frame length to obtain at least two audio sub-blocks. Each audio sub-block has a frame length equal to or less than the predetermined sub-block frame length. For example, as shown in FIG. 5d, the jth audio frame has a frame length of 5 milliseconds (ms). It may be divided into five audio sub-blocks with a length of 1 ms per block. In another implementation, the jth audio frame may be divided by a predetermined number to obtain a further predetermined number of audio sub-blocks. Furthermore, the frame lengths of each audio sub-block may be equal. For example, a 5-ms audio frame may be divided into five audio sub-blocks using a predetermined number of 5 audio sub-blocks. Then, the amplitude values ​​of the divided audio sub-blocks may be amplified in units of audio sub-blocks, specifically, the amplitude values ​​of each sampling point of the audio sub-blocks are amplified to achieve volume adjustment of the j-th audio frame.

[0147] In step 4.3, the amplitude value of each of the N audio sub-blocks is amplified stepwise based on the frame amplitude value gain of the j-th audio frame to adjust the volume of each audio sub-block.

[0148] In a specific implementation, a block amplitude value gain may be determined for each of the N audio sub-blocks based on the frame amplitude value gain of the jth audio frame. The block amplitude value gain of any one audio sub-block is the amplitude value gain (also called subblock gain, or subgain) required to amplify the amplitude value of the audio sub-block. The amplitude value of each audio sub-block may be gradually amplified based on the block amplitude value gain of each audio sub-block to adjust the volume of each audio sub-block. By gradually amplifying the amplitude value of each audio sub-block, the amplified amplitude value of each audio sub-block divided from the same audio frame may be gradually changed (e.g., gradually increased or gradually decreased). The volume is gradually adjusted with a smaller granularity.

[0149] As can be understood, the "amplitude value gain of the current audio signal", "frame amplitude value gain of an audio frame", and "block amplitude value gain of an audio sub-block" referred to in this application are essentially all gain values ​​required to adjust the amplitude value, but for ease of distinction, they are given different names for audio signals of different lengths. For the complete current audio signal, the corresponding one is the amplitude value gain, for a small audio frame, the corresponding amplitude value gain is called the frame amplitude value gain, and for a smaller audio sub-block, the corresponding amplitude value gain is called the block amplitude value gain.

[0150] In one embodiment, before performing step 4.3, the computer device may determine the sequence position of the jth audio frame among the J audio frames. This sequence position can be used to indicate whether the jth audio frame is the first audio frame among the J audio frames, i.e., whether the jth audio frame is sequenced first. Alternatively, it can be understood as whether the ith audio frame corresponding to the jth audio frame is the first audio frame among the K audio frames for which signal level monitoring has been performed. Based on the difference in sequence position, the following two cases may occur. If the determined sequence position is not the first, the execution of step 4.3 above is triggered. That is, if the jth audio frame is not the first audio frame among the J audio frames, a step of gradually amplifying the amplitude value of each of the N audio sub-blocks based on the frame amplitude value gain of the jth audio frame may be performed to gradually adjust the volume of the jth audio frame. If the determined sequence position is the first, the amplitude value of each of the N audio sub-blocks is respectively amplified using the frame amplitude value gain of the jth audio frame to adjust the volume of each audio sub-block. That is, if the jth audio frame is the first audio frame of J audio frames, the frame amplitude gain of the jth audio frame may be used as the block amplitude gain of each audio sub-block, and the block amplitude gain may be used to amplify the amplitude of each audio sub-block divided from the first audio frame of J audio frames. This achieves volume adjustment for each audio sub-block and volume adjustment for the first audio frame of J audio frames. In this method, a unified block amplitude gain is used to adjust the amplitude of the N audio sub-blocks, allowing for quick volume adjustment of the first audio frame.

[0151] In one implementation, when the computer device adjusts the j-th audio frame in step 4.3, the computer device may specifically perform the following steps 4.3.1 to 4.3.2.

[0152] In step 4.3.1, a block amplitude value gain of each audio sub-block is determined based on the frame amplitude value gain of the j-th audio frame and the arrangement position of each audio sub-block among the N audio sub-blocks.

[0153] The block amplitude value gains of different audio sub-blocks change gradually as the arrangement position of the corresponding audio sub-block increases among the N audio sub-blocks. The gradual change here may be a gradual increase or a gradual decrease. For example, the jth audio frame is divided into five audio sub-blocks. The frame amplitude value gain of the jth audio frame is 5, and the block amplitude value gains of the audio sub-blocks are 3.0, 3.4, 3.8, 4.2, 4.6, and 5, respectively. As can be seen, the block amplitude value gains of the audio sub-blocks increase gradually. Therefore, by using these block amplitude value gains to adjust the amplitude values ​​of the audio sub-blocks, gradual gain adjustment of the audio frame can be achieved through gradual gain adjustment in units of audio sub-blocks, and further, the volume adjustment of the audio frame can be made smoother.

[0154] In one possible embodiment, the jth audio frame in the J audio frames is not at the first position. When performing the above step 4.3.1, the computer device may specifically implement the following scheme:

[0155] (1) Determine the frame amplitude value gain of the j-1th audio frame among the J audio frames, and define the difference between the frame amplitude value gain of the jth audio frame and the frame amplitude value gain of the j-1th audio frame as the frame amplitude value gain difference of the jth audio frame relative to the j-1th audio frame. The j-1th audio frame is the audio frame immediately preceding the jth audio frame among the J audio frames. The difference in frame amplitude value gain between two adjacent audio frames, i.e., the frame amplitude value gain difference, can be expressed as g(j)-g(j-1), where g(j) is the frame amplitude value gain of the jth audio frame, and g(j-1) is the frame amplitude value gain of the j-1th audio frame. (2) For an nth audio subblock among the N audio subblocks, determine a block amplitude value gain change value required for the nth audio subblock relative to the j-1th audio frame based on the arrangement position of the nth audio subblock among the N audio subblocks and the frame amplitude value gain difference. Specifically, the ratio of the arrangement position n of the n-th audio sub-block to the total number N of divided audio sub-blocks may be determined, and the product of the calculated ratio n / N and the frame amplitude value gain difference may be used as the block amplitude value gain change value required for the n-th audio sub-block for the j-1-th audio frame.

number

[0156] Illustratively, the complete expression corresponding to the calculation of the block amplitude value gain according to the above scheme may be as follows:

number

[0157] where:

number

number

number

[0158] According to the above formula, the block amplitude value gain of the last audio block among each audio sub-block divided from the j-1th audio frame is the same as the frame amplitude value gain of the j-1th audio frame. Therefore, the transition of amplitude value adjustment of adjacent audio sub-blocks corresponding to different audio frames is gradual rather than abrupt. It can be seen that for any non-leading audio sub-block among the N audio sub-blocks, the corresponding block amplitude value gain can be obtained by the above method, and the block amplitude value gain can be used to realize gradual gain adjustment of the audio frame.

[0159] When determining the nth (n∈[1,N]) audio sub-block, the computer device may sequentially scan the N audio sub-blocks and determine the currently scanned audio sub-block as the nth audio sub-block. Furthermore, the computer device may sequentially calculate the block amplitude value gains of the N audio sub-blocks. Alternatively, the computer device may calculate the block amplitude value gains of the N audio sub-blocks in parallel. For any of the audio sub-blocks, the block amplitude value gain can be calculated in parallel using the method introduced above. This allows the block amplitude value gains of each of the N audio sub-blocks to be quickly determined.

[0160] According to the above method, the block amplitude value gain of an audio sub-block can be calculated based on the frame amplitude value gain of the currently processed audio frame and the frame amplitude value gain of the immediately preceding audio frame. The obtained block amplitude value gain can be gradually changed as the array position of each audio sub-block increases. By adjusting the audio sub-blocks through gradually changing the block amplitude value gain, the amplitude value within one audio frame can also be gradually changed on an audio sub-block basis, thereby smoothing the amplitude value adjustment within one frame. Based on the connection between the amplitude value gains of adjacent audio frames, the amplitude value between adjacent audio frames can also be smoothly changed.

[0161] In one specific implementation, to realize a gradual change in the block amplitude value gain, the block amplitude value gain of the last audio sub-block among the audio sub-blocks divided from the j-1th audio frame is also used in determining the block amplitude value gain of each audio sub-block in the jth audio frame. In a specific implementation, the frame amplitude value gain of the j-1th audio frame mentioned above may be replaced with the block amplitude value gain of the last audio sub-block among the audio sub-blocks divided from the j-1th audio frame, and the corresponding frame amplitude value gain difference and block amplitude value gain change value may be sequentially obtained for the nth audio sub-block, and the sum of the block amplitude value gain change value and the block amplitude value gain of the last audio sub-block divided from the j-1th audio frame may be determined as the block amplitude value gain of the nth audio sub-block.

[0162] In another embodiment, the execution of step 4.3 above can be triggered regardless of whether the jth audio frame among the J audio frames is first. To achieve gradual amplification of the amplitude value of each audio sub-block in an audio frame, if the audio frame among the J audio frames has the first position in the sequence, i.e., if the jth audio frame among the J audio frames is first in the sequence, the block amplitude value gain of each audio sub-block may be determined based on the sequence position of each audio sub-block among the N audio sub-blocks and the frame amplitude value gain of the jth audio frame. In this way, the block amplitude value gain of each audio sub-block is determined according to the sequence position of the audio sub-block, and the block amplitude value gains of different audio sub-blocks are different. The block amplitude value gains of the N audio sub-blocks change gradually as the sequence position of the audio sub-block increases. This achieves gradual amplitude value adjustment of each audio sub-block among the N audio sub-blocks based on the block amplitude value gain of each audio sub-block. Specifically, for an n-th (n∈[1,N) audio sub-block among N audio sub-blocks, a ratio may be calculated between the array position n of the n-th audio sub-block among the N audio sub-blocks and the total number N of audio sub-blocks divided from the j-th audio frame. Next, this ratio is multiplied by the frame amplitude value gain of the j-th audio frame, and the resulting product is the block amplitude value gain of the n-th audio sub-block. Exemplarily, the expression for the above calculation may be as follows:

number

[0163] where:

number

[0164] As can be seen, the block amplitude value gain of the last audio sub-block among the N audio sub-blocks calculated by the above method is the frame amplitude value gain of the j-th audio frame. Also, the block amplitude value gains of different audio sub-blocks determined by the above method increase stepwise as the arrangement position n of the audio sub-block increases among the N audio sub-blocks. This realizes stepwise adjustment of the amplitude value of the j-th audio frame in units of audio sub-blocks.

[0165] In step 4.3.2, the block amplitude value gain of each audio sub-block is used to amplify the amplitude value of the corresponding audio sub-block.

[0166] The block amplitude value gain of each audio sub-block is used to amplify the amplitude value of the corresponding audio sub-block, thereby obtaining each volume-adjusted audio sub-block. Taking the nth audio sub-block as an example, the nth audio sub-block includes a plurality of sampling points and amplitude values ​​of each sampling point. The computer device may use the determined block amplitude value gain to amplify the amplitude value of the nth audio sub-block, thereby obtaining an amplified amplitude value corresponding to the nth audio sub-block. Specifically, the amplitude value of each sampling point may be multiplied by the block amplitude value gain of the nth audio sub-block. The resulting product is the amplified amplitude value of the corresponding sampling point. That is, the amplified amplitude value corresponding to the nth audio sub-block includes the amplified amplitude value of each sampling point, and the volume-adjusted nth audio sub-block can be obtained by matching the amplified amplitude value of each sampling point with each sampling point. By performing amplification processing using the above method on any of the N audio sub-blocks, it is possible to obtain audio sub-blocks after each volume adjustment, and the volume of each audio sub-block after volume adjustment also changes gradually as the arrangement position of the corresponding audio sub-block in the N audio sub-blocks increases.

[0167] The volume adjustment method for audio frames shown in steps 4.3.1 to 4.3.2 above uses audio sub-blocks as units, and realizes gradual adjustment of the amplitude value of each audio sub-block in one audio frame using the block amplitude value gain of the audio sub-block, and further enables smooth transition of the volume adjustment of the audio frame.

[0168] In step 4.4, after each volume-adjusted audio sub-block is obtained, the volume-adjusted audio sub-blocks are concatenated to obtain the j-th volume-adjusted audio frame.

[0169] Any one of the audio sub-blocks after the volume adjustment is a new audio sub-block that achieves the volume adjustment by the amplitude value gain, and the computer device may obtain the jth audio frame after the volume adjustment by connecting the audio sub-blocks after the volume adjustment according to their arrangement positions.

[0170] In step 4.5, scanning continues until all J audio frames have been scanned to obtain J audio frames after volume adjustment, and the J audio frames after volume adjustment are connected together to obtain the current audio signal after volume adjustment.

[0171] After completing the processing of the j-th audio frame, the computer device may continue to scan the next audio frame until scanning of J audio sub-blocks is completed, and perform stepwise amplification using similar logic to obtain a volume-adjusted audio frame. In this way, a volume-adjusted audio frame corresponding to each audio frame for which signal level monitoring has been performed is obtained, and the volume-adjusted audio frames are further connected according to their arrangement positions to obtain a volume-adjusted current audio signal.

[0172] In one specific implementation, the current audio signal after volume adjustment may be obtained by directly splicing each audio frame after volume adjustment according to its array position. In another specific implementation, if L is less than K, that is, if the number of audio frames for which signal level monitoring has been performed among the K audio frames is less than K, it indicates that the K audio frames include some invalid audio frames for which signal level monitoring has not been performed. The computer device may obtain the current audio signal after volume adjustment by splicing each audio frame after volume adjustment with other audio frames other than the J audio frames among the K audio frames.

[0173] As can be seen, when the current audio signal is processed into divided audio frames, the amplitude values ​​of the audio frames can be gradually adjusted for each audio sub-block divided from each audio frame whose signal level has been monitored, according to the method shown in steps 4.1 to 4.5 above. In a specific implementation, an appropriate block amplitude value gain is determined for each of the N audio sub-blocks based on the amplitude value gain of the jth audio frame, and the corresponding audio sub-block is then amplified based on the block amplitude value gain, thereby achieving gradual amplification of each audio sub-block in the audio frame according to the block amplitude value gain, thereby achieving smooth volume adjustment of the audio frame and smoothing the volume change corresponding to the resulting volume-adjusted audio signal.

[0174] The volume control methods described in (1) to (4) above can be applied to real audio signals or virtual audio signals. In addition, in a specific processing process, the level of the current audio signal being processed is monitored, and a gain control process is performed based on the amplitude constraint parameters of the current audio signal in the target mixing mode and the monitored level description value to determine an appropriate amplitude gain for adjusting the amplitude of the current audio signal. The volume-controlled current audio signal can then be finally obtained by adjusting the amplitude. This process aims to achieve the mixing effect required for the target mixing mode, and the amplitude constraint parameters are used to constrain the adjustment of the amplitude of the current audio signal, making the amplitude adjustment more reasonable, thereby obtaining a desirable volume-controlled audio signal for mixing, which is advantageous for achieving a desirable mixing effect.

[0175] Based on the above introduction, an exemplary flowchart of volume adjustment can be provided, as shown in FIG. 6a. As shown in FIG. 6a, a stereo recording headset can collect real-world stereo in real time to form an audio signal. For example, if the input audio signal is a collected real-world audio signal, the audio signal can be converted into a true audio signal. Then, an audio effectiveness determination process (e.g., VAD determination) is performed on an audio frame basis, and then signal level monitoring of the valid audio frame is performed on the audio frame as the smallest unit. After obtaining a level description value through monitoring, gain control of the valid audio frame can be performed based on a target amplitude value and an amplitude gain threshold corresponding to the target mixing mode setting. The gain control can determine the amplitude gain required to adjust the amplitude value of the valid audio frame. Finally, a volume-adjusted audio frame is obtained by adjusting the sample point amplitude value of the valid audio frame. According to the above flow, adjustment can be performed on any valid audio frame to achieve volume adjustment, and a volume-adjusted true audio signal is finally obtained.

[0176] In step S404, the volume-adjusted real audio signal and the volume-adjusted virtual audio signal are mixed to obtain a mixed signal.

[0177] In one specific implementation, a mixing algorithm (e.g., an average weighting method or an adaptive weighting method) may be used to mix the volume-adjusted real audio signal with the volume-adjusted virtual audio signal to obtain and output a mixed signal, which is the result of fusing real-world sounds with virtual-world sounds. In this way, audio processing in the augmented reality scene is completed, and the enhancement processing of real-world sounds and virtual-world sounds is realized.

[0178] In one possible implementation, each mixing mode is associated with a set of weighting parameters. The set of weighting parameters includes a first weight for the true audio signal in the corresponding mixing mode and a second weight for the virtual audio signal in the corresponding mixing mode. The different sets of weighting parameters corresponding to different mixing modes specifically represent the possibility of mixing real-world sounds and virtual-world sounds with different weights. With the support of the set of weighting parameters, the mixing effect achieved in the corresponding mixing mode can be further enhanced. For example, in the peer mixing mode, the first weight and the second weight in the corresponding set of weighting parameters may be the same. In the strong-reality mixing mode, the first weight is greater than the second weight in the corresponding set of weighting parameters, thereby further enhancing the volume of the true audio signal and making the real-world sounds more audible.

[0179] The computer device may obtain a set of weighting parameters corresponding to a target mixing mode, and use the obtained set of weighting parameters as a target weighting parameter set. The target weighting parameter set includes a first weight for a true audio signal in the target mixing mode and a second weight for a virtual audio signal in the corresponding mixing mode. The computer device may then use the first weight and the second weight in the target parameter set to weight the volume-adjusted true audio signal and the volume-adjusted virtual audio signal, thereby obtaining a mixed signal. In a specific implementation, the computer device may add a product of the first weight and the volume-adjusted true audio signal and a product of the second weight and the volume-adjusted true audio signal. The obtained weighted addition result may be used as a mixed signal.

[0180] In the embodiment of the present application, after determining a target mixing mode for the real audio signals and virtual audio signals in the augmented reality scene, volume adjustment processing for the real audio signals and virtual audio signals can be performed in the target mixing mode, and the volume can be adjusted to reach a volume value required for the mixing effect. Then, the volume-adjusted real audio signals and the volume-adjusted virtual audio signals can be mixed to obtain and output a mixed signal. In the entire process, the target mixing mode can be set manually by a user, or the target mixing mode can be automatically determined by a computing device, and the volume adjustment of the audio signals can be controlled in the target mixing mode. The volume adjustment can achieve a good mixing effect between the real world sounds and the virtual world sounds after mixing, bringing about a corresponding auditory sensory experience.

[0181] Based on the above S401 to S404, an exemplary flowchart as shown in FIG. 6b can be further provided. As shown in FIG. 6b, a stereo recording signal can be acquired. This stereo recording signal is obtained by collecting real-world sounds in real time using a stereo recording headset. Then, a noise-reduced true audio signal can be obtained by filtering environmental noise interference using a noise reduction process. The noise-reduced true audio signal can be used for later processing of virtual world sound sources. The volume of the true sound source and the virtual sound source can be adjusted according to the volume adjustment flow shown in FIG. 6a. Specifically, the volume of the audio signal corresponding to the sound generated from the sound source can be adjusted based on the configured target mixing mode. Finally, the volume-adjusted true audio signal and the volume-adjusted virtual audio signal are mixed to achieve mixing of real-world stereo and virtual-world stereo. This configuration mainly focuses on the fusion of real-world and virtual-world sounds in terms of hearing, specifically, the fusion of real-world and virtual-world audio signals. In this configuration, external real-world sounds can be collected using a stereo recording headset. This sound becomes a sound source with a sense of direction and is subsequently processed with the sound source in the virtual space. Users can set different mixing modes to control the mixing method of real-world sound sources. Finally, the real-world and virtual-world sounds are enhanced, processed, and stereo-mixed for output. For example, after the AR headset collects real-world sounds with a stereo microphone, the real-world sounds can be stereo-mixed with sounds generated from a sound source constructed in the virtual world as sounds generated from independent stereo sources. Users can achieve different mixing effects between real-world and virtual-world sound sources by customizing the mixing mode. This solves the problem that current AR scenes only have visual fusion, but lack deep fusion in terms of audio.

[0182] Next, an audio processing device according to an embodiment of the present invention will be described.

[0183] Please refer to FIG. 7. FIG. 7 is a schematic diagram of the configuration of an audio processing device provided in an embodiment of the present application. This audio processing device may be provided in a computer device provided in an embodiment of the present application. Schematically, the computer device may be the computer device mentioned in the above method embodiment. The audio processing device shown in FIG. 7 may be a computer program (including program code) executed on the computer device. This audio processing device can be used to perform some or all of the steps in the method embodiments shown in FIG. 2 and FIG. 4. This audio processing device may include the following units:

[0184] The acquiring unit 701 acquires a real audio signal and a virtual audio signal in an augmented reality scene, the real audio signal means an audio signal collected in a real environment on which the augmented reality scene is based, and the virtual audio signal means an audio signal constructed in a virtual environment on which the augmented reality scene is based; The processing unit 702 determines a plurality of mixing modes arranged for the augmented reality scene, and selects one mixing mode from the plurality of mixing modes as a target mixing mode, where different mixing modes are for realizing different mixing effects between the real audio signal and the virtual audio signal; The processing unit 702 further mixes the real audio signal and the virtual audio signal in a target mixing mode to obtain a mixed signal; The output unit 703 outputs the mixed signal.

[0185] In one embodiment, when selecting one mixing mode from the plurality of mixing modes as the target mixing mode, the processing unit 702 specifically: A mode arrangement operation is acquired, and one mixing mode is selected as a target mixing mode from among a plurality of mixing modes according to the mode arrangement operation.

[0186] Here, the mode configuration operation includes a selection operation for multiple mode identifiers displayed in the augmented reality application, where one mode identifier corresponds to one mixing mode, or the mode configuration operation includes a mode selection operation by operating and controlling a physical key on the augmented reality device.

[0187] In one embodiment, when selecting one mixing mode from the plurality of mixing modes as the target mixing mode, the processing unit 702 specifically: Determine geographic location information of a real environment on which the augmented reality scene is based, and determine an environment type to which the real environment belongs based on the geographic location information; According to the sound fusion needs of the determined environment type, one mixing mode is selected as a target mixing mode from among a plurality of mixing modes.

[0188] In one embodiment, the processing unit 702 performs a mixing process on the real audio signal and the virtual audio signal in the target mixing mode to obtain the mixed signal, specifically: adjusting the volume of each of the real audio signal and the virtual audio signal in the target mixing mode to obtain a volume-adjusted real audio signal and a volume-adjusted virtual audio signal; A mixed signal is obtained by mixing the volume-adjusted real audio signal and the volume-adjusted virtual audio signal.

[0189] In one embodiment, the volume of the audio signal changes with the change in the amplitude value of the audio signal. The processing unit 702 performs volume adjustment on each of the real audio signal and the virtual audio signal in the target mixing mode to obtain the volume-adjusted real audio signal and the volume-adjusted virtual audio signal, specifically by: Scan the real audio signal and the virtual audio signal, and obtain amplitude value constraint parameters of the current audio signal in the target mixing mode by using the currently scanned audio signal as the current audio signal; By performing signal level monitoring of the current audio signal, a level description value of the current audio signal is obtained, which indicates an amplitude value of the current audio signal; Based on the acquired amplitude value constraint parameter and the level description value of the current audio signal, a gain control process is performed for the amplitude value of the current audio signal to acquire an amplitude value gain of the current audio signal; The amplitude value gain of the current audio signal is used to amplify the amplitude value of the current audio signal, thereby adjusting the volume of the current audio signal and obtaining the volume-adjusted current audio signal.

[0190] In one embodiment, when obtaining the amplitude value constraint parameters of the current audio signal in the target mixing mode, the processing unit 702 specifically: obtaining a predetermined mapping table, the predetermined mapping table including a plurality of mode identifiers and a parameter group mapped to each of the mode identifiers, the parameter group mapped to any one of the mode identifiers including a first amplitude value constraint parameter of the virtual audio signal in the mixing mode identified by the corresponding mode identifier and a second amplitude value constraint parameter of the real audio signal in the mixing mode identified by the corresponding mode identifier; Searching a predetermined mapping table using the mode identifier of the target mixing mode to obtain a parameter group mapped to the mode identifier of the target mixing mode; If the current audio signal is a virtual audio signal, select a first amplitude value constraint parameter from the acquired parameter group as an amplitude value constraint parameter of the current audio signal in the target mixing mode; If the current audio signal is a real audio signal, a second amplitude value constraint parameter is selected from the obtained parameter group as the amplitude value constraint parameter of the current audio signal in the target mixing mode.

[0191] In one embodiment, the plurality of mixing modes include at least two of a strong real mixing mode, a strong virtual mixing mode, a peer mixing mode, a real off-mixing mode, and a virtual off-mixing mode, and an amplitude value constraint parameter of any one of the audio signals is positively correlated with a volume of the corresponding audio signal after volume adjustment; In the strong reality mixing mode, both the amplitude value constraint parameter of the real audio signal and the amplitude value constraint parameter of the virtual audio signal are valid values, and the amplitude value constraint parameter of the real audio signal is greater than the amplitude value constraint parameter of the virtual audio signal; In the strong virtual mixing mode, both the amplitude value constraint parameter of the real audio signal and the amplitude value constraint parameter of the virtual audio signal are valid values, and the amplitude value constraint parameter of the virtual audio signal is greater than the amplitude value constraint parameter of the real audio signal; In the peer mixing mode, both the amplitude value constraint parameter of the real audio signal and the amplitude value constraint parameter of the virtual audio signal are valid values, the amplitude value constraint parameter of the virtual audio signal is equal to the amplitude value constraint parameter of the real audio signal, and In the real off-mixing mode, the amplitude value constraint parameter of the real audio signal is an invalid value, and the amplitude value constraint parameter of the virtual audio signal is an valid value; In the virtual off-mixing mode, the amplitude value constraint parameter of the real audio signal is a valid value, and the amplitude value constraint parameter of the virtual audio signal is an invalid value.

[0192] In one embodiment, the processing unit 702 performs signal level monitoring of the current audio signal to obtain the level description value of the current audio signal, specifically by: By performing framing processing on the current audio signal, K audio frames (K is an integer greater than 1) of the current audio signal are obtained; Scan the K audio frames and determine the i-th (i∈[1,K]) audio frame currently being scanned; Obtaining a level description value of the i-th audio frame by monitoring the signal level of the i-th audio frame, or performing a voice validity determination process on the i-th audio frame, and if it is determined that the i-th audio frame is a valid audio frame, obtaining a level description value of the i-th audio frame by monitoring the signal level of the i-th audio frame; The scanning continues until all K audio frames have been scanned to obtain a level description value of the current audio signal, where the level description value of the current audio signal includes the level description value of each monitored audio frame.

[0193] In one embodiment, each audio frame includes a plurality of sampling points and an amplitude value of each sampling point. When the processing unit 702 obtains the level description value of the i-th audio frame by monitoring the signal level of the i-th audio frame, the processing unit 702 specifically includes: Calculating the average value of the amplitude values ​​of each sampling point in the i-th audio frame to obtain the average signal amplitude value of the i-th audio frame; The level description value of the i-th audio frame is obtained by smoothing the average signal amplitude value of the i-th audio frame.

[0194] In one embodiment, the processing unit 702 smoothes the average signal amplitude value of the i-th audio frame to obtain the level description value of the i-th audio frame, specifically by: A reference audio frame for the i-th audio frame is searched for from the K audio frames. If the i-th audio frame has not been subjected to a voice activation determination process, the reference audio frame includes the (i-1)-th audio frame. If the i-th audio frame has been subjected to a voice activation determination process, the reference audio frame includes a valid audio frame that is located before the i-th audio frame and is closest to the i-th audio frame. If no reference audio frame is found, the average signal amplitude value of the i-th audio frame is taken as the level description value of the i-th audio frame; If a reference audio frame is found, the level description value of the i-th audio frame is obtained by smoothing the average signal amplitude value of the i-th audio frame using the level description value of the reference audio frame.

[0195] In one embodiment, the current audio signal is divided into K audio frames (K is an integer greater than 1), the level description value of the current audio signal includes level description values ​​of J audio frames (J≦K) among the K audio frames that have undergone signal level monitoring, and the acquired amplitude value constraint parameters include a target amplitude value and an amplitude value gain threshold; The processing unit 702 performs a control process for a gain required for the amplitude value of the current audio signal based on the acquired amplitude value constraint parameter and the level description value of the current audio signal, thereby obtaining the amplitude value gain of the current audio signal, specifically: Scanning J audio frames according to the level description values ​​of the current audio signal to determine the jth (j∈[1,J]) audio frame currently being scanned; By performing gain prediction using the target amplitude value in the acquired amplitude value constraint parameter and the level description value of the j-th audio frame, a predicted amplitude value gain is acquired, which indicates a gain required to adjust the level description value of the j-th audio frame to the target amplitude value; determining the smaller one of the amplitude gain threshold in the obtained amplitude constraint parameter and the predicted amplitude gain as the frame amplitude gain required to adjust the amplitude value of the j-th audio frame; By continuing to scan until all of the J audio frames have been scanned, an amplitude value gain of the current audio signal is obtained, and the amplitude value gain of the current audio signal includes a frame amplitude value gain corresponding to each of the J audio frames.

[0196] In one embodiment, the current audio signal is divided into K audio frames (K is an integer greater than 1), and the amplitude value gain of the current audio signal includes frame amplitude value gains of J audio frames (J≦K) among the K audio frames for which signal level monitoring has been performed; The processing unit 702 adjusts the volume of the current audio signal by amplifying the amplitude value of the current audio signal using the amplitude value gain of the current audio signal. When obtaining the current audio signal after volume adjustment, specifically: Scanning J audio frames according to the level description values ​​of the current audio signal to determine the jth (j∈[1,J]) audio frame currently being scanned; Dividing the j-th audio frame into N audio sub-blocks (N is an integer greater than 1), and determining an amplitude value for each of the N audio sub-blocks; Stepwise amplifying the amplitude value of each of the N audio sub-blocks based on the frame amplitude value gain of the j-th audio frame to adjust the volume of each audio sub-block; After obtaining each volume-adjusted audio sub-block, the volume-adjusted audio sub-blocks are spliced ​​together to obtain a volume-adjusted j-th audio frame; Scanning continues until all of the J audio frames have been scanned to obtain J audio frames after volume adjustment, and the J audio frames after volume adjustment are connected together to obtain the current audio signal after volume adjustment.

[0197] In one embodiment, the processing unit 702 further comprises: determining an array position of a j-th audio frame among the J audio frames; If the determined arrangement position is not the first, trigger the execution of a step of gradually amplifying the amplitude value of each of the N audio sub-blocks according to the frame amplitude value gain of the j-th audio frame to adjust the volume of each audio sub-block; If the determined arrangement position is the first, the amplitude value of each of the N audio sub-blocks is amplified using the frame amplitude value gain of the j-th audio frame to adjust the volume of each audio sub-block.

[0198] In one embodiment, when the processing unit 702 gradually amplifies the amplitude value of each of the N audio sub-blocks according to the frame amplitude value gain of the j-th audio frame, specifically: Determine a block amplitude value gain for each audio sub-block according to the frame amplitude value gain of the jth audio frame and the arrangement position of each audio sub-block in the N audio sub-blocks, and the block amplitude value gains of different audio sub-blocks change stepwise with the increase of the arrangement position of the corresponding audio sub-block in the N audio sub-blocks; The block amplitude value gain of each audio sub-block is used to amplify the amplitude value of the corresponding audio sub-block.

[0199] In one embodiment, the j-th audio frame in the J audio frames is not the first audio frame; Specifically, when determining the block amplitude value gain of each audio sub-block based on the frame amplitude value gain of the jth audio frame and the arrangement position of each audio sub-block among the N audio sub-blocks, the processing unit 702 Determine a frame amplitude value gain of a j-1th audio frame among the J audio frames, and define a difference between the frame amplitude value gain of the jth audio frame and the frame amplitude value gain of the j-1th audio frame as a frame amplitude value gain difference of the jth audio frame relative to the j-1th audio frame; For an n-th (n∈[1,N]) audio sub-block among the N audio sub-blocks, determine a block amplitude value gain change value required for the n-th audio sub-block for the j−1-th audio frame based on an arrangement position of the n-th audio sub-block among the N audio sub-blocks and a frame amplitude value gain difference; The block amplitude value gain of the n-th audio sub-block is obtained based on the block amplitude value gain change value and the frame amplitude value gain of the j-1-th audio frame.

[0200] In one embodiment, before obtaining the mixed signal by mixing the real audio signal and the virtual audio signal in the target mixing mode, the processing unit 702 further: Determine a scene type of the augmented reality scene, and perform sound effect processing on the real audio signal and the virtual audio signal in the augmented reality scene according to the sound effect needs of the scene type; or Obtain an environment type to which the real environment on which the augmented reality scene is based belongs, and perform sound effect processing on the real audio signal and the virtual audio signal in the augmented reality scene according to the sound effect needs of the environment type to which the real environment belongs; Here, the sound effect processing includes at least one of reverberation processing and equalization processing.

[0201] In one embodiment, each mixing mode among the plurality of mixing modes has a mode identifier, and each mode identifier corresponds to one set of equalization parameters required for performing equalization processing, and the set of equalization parameters corresponding to any one mode identifier includes a first volume equalization parameter of the virtual audio signal in the corresponding mixing mode and a second volume equalization parameter of the real audio signal in the corresponding mixing mode; and before obtaining the mixed signal by mixing the real audio signal and the virtual audio signal in the target mixing mode, the processing unit 702 further: Obtaining a set of equalization parameters corresponding to a mode identifier of the target mixing mode; The virtual audio signal is equalized using a first volume equalization parameter in the obtained set of equalization parameters, and the real audio signal is equalized using a second volume equalization parameter in the obtained set of equalization parameters.

[0202] In an embodiment of the present application, a true audio signal and a virtual audio signal in an augmented reality scene can be obtained. A plurality of mixing modes arranged for the augmented reality scene are determined, and one mixing mode from the plurality of mixing modes is selected as a target mixing mode. Different mixing modes can be used to achieve different mixing effects between the true audio signal and the virtual audio signal, resulting in different auditory experiences. Determining a plurality of mixing modes for the augmented reality scene provides a variety of options for setting the target mixing mode, allowing the target mixing mode to be flexibly set. This can accommodate various scenes requiring mixing of the true audio signal and the virtual audio signal, and can also meet the user's customized mixing needs. A mixed signal is obtained by mixing the true audio signal and the virtual audio signal in the target mixing mode, and the mixed signal is output. The target mixing mode can achieve a corresponding mixing effect between real audio signals and virtual audio signals. By controlling the mixing process between real audio signals and virtual audio signals in the target mixing mode, the fusion of real-world sounds and virtual-world sounds on which the augmented reality scene is based can be achieved in the auditory dimension, thereby increasing the auditory fusion of sounds in the augmented reality scene and enriching the fusion capabilities of the augmented reality scene. The resulting mixing signal also has a mixing effect corresponding to the target mixing mode. Finally, the output of the mixing signal allows the user to more intuitively experience the fusion effect of sounds in the augmented reality scene, enhancing the user's immersion in the augmented reality scene.

[0203] Next, a computer device provided in an embodiment of the present application will be described.

[0204] In the embodiments of the present application, a schematic diagram of the configuration of a computer device is also provided. See FIG. 8 for the schematic diagram of the configuration of this computer device. This computer device may include a processor 801, an input device 802, an output device 803, and a memory 804. The processor 801, the input device 802, the output device 803, and the memory 804 are connected via a bus. The memory 804 stores a computer program, which includes program instructions, and the processor 801 executes the program instructions stored in the memory 804.

[0205] In one embodiment, the computing device may be a terminal device in the system shown in Figure 1a or a server in the system shown in Figure 1b. In this embodiment, the processor 801 executes executable program code in the memory 804 to: acquiring real audio signals and virtual audio signals in an augmented reality scene, wherein the real audio signals refer to audio signals collected in a real environment on which the augmented reality scene is based, and the virtual audio signals refer to audio signals constructed in a virtual environment on which the augmented reality scene is based; determining a plurality of mixing modes arranged for the augmented reality scene, and selecting one mixing mode from the plurality of mixing modes as a target mixing mode, where different mixing modes are for realizing different mixing effects between the real audio signal and the virtual audio signal; obtaining a mixed signal by mixing the real audio signal and the virtual audio signal in a target mixing mode; and outputting the mixed signal.

[0206] In one embodiment, when the processor 801 selects one mixing mode from the plurality of mixing modes as the target mixing mode, the processor 801 specifically selects one mixing mode from the plurality of mixing modes as the target mixing mode by: Obtain a mode placement operation, and select one mixing mode from the plurality of mixing modes as a target mixing mode according to the mode placement operation; Here, the mode configuration operation includes a selection operation for multiple mode identifiers displayed in the augmented reality application, where one mode identifier corresponds to one mixing mode, or the mode configuration operation includes a mode selection operation by operating and controlling a physical key on the augmented reality device.

[0207] In one embodiment, when the processor 801 selects one mixing mode from the plurality of mixing modes as the target mixing mode, the processor 801 specifically selects one mixing mode from the plurality of mixing modes as the target mixing mode by: Determine geographic location information of a real environment on which the augmented reality scene is based, and determine an environment type to which the real environment belongs based on the geographic location information; According to the sound fusion needs of the determined environment type, one mixing mode is selected as a target mixing mode from among a plurality of mixing modes.

[0208] In one embodiment, the processor 801 performs a mixing process on the real audio signal and the virtual audio signal in the target mixing mode to obtain a mixed signal, specifically: adjusting the volume of each of the real audio signal and the virtual audio signal in the target mixing mode to obtain a volume-adjusted real audio signal and a volume-adjusted virtual audio signal; A mixed signal is obtained by mixing the volume-adjusted real audio signal and the volume-adjusted virtual audio signal.

[0209] In one embodiment, the volume of the audio signal changes with a change in the amplitude value of the audio signal. The processor 801 performs volume adjustment on each of the real audio signal and the virtual audio signal in the target mixing mode to obtain the real audio signal after volume adjustment and the virtual audio signal after volume adjustment, specifically: Scan the real audio signal and the virtual audio signal, and obtain amplitude value constraint parameters of the current audio signal in the target mixing mode by using the currently scanned audio signal as the current audio signal; By performing signal level monitoring of the current audio signal, a level description value of the current audio signal is obtained, which indicates an amplitude value of the current audio signal; Based on the acquired amplitude value constraint parameter and the level description value of the current audio signal, a gain control process is performed for the amplitude value of the current audio signal to acquire an amplitude value gain of the current audio signal; The amplitude value gain of the current audio signal is used to amplify the amplitude value of the current audio signal, thereby adjusting the volume of the current audio signal and obtaining the volume-adjusted current audio signal.

[0210] In one embodiment, when obtaining the amplitude value constraint parameters of the current audio signal in the target mixing mode, the processor 801 specifically: obtaining a predetermined mapping table, the predetermined mapping table including a plurality of mode identifiers and a parameter group mapped to each of the mode identifiers, the parameter group mapped to any one of the mode identifiers including a first amplitude value constraint parameter of the virtual audio signal in the mixing mode identified by the corresponding mode identifier and a second amplitude value constraint parameter of the real audio signal in the mixing mode identified by the corresponding mode identifier; Searching a predetermined mapping table using the mode identifier of the target mixing mode to obtain a parameter group mapped to the mode identifier of the target mixing mode; If the current audio signal is a virtual audio signal, select a first amplitude value constraint parameter from the acquired parameter group as an amplitude value constraint parameter of the current audio signal in the target mixing mode; If the current audio signal is a real audio signal, a second amplitude value constraint parameter is selected from the obtained parameter group as the amplitude value constraint parameter of the current audio signal in the target mixing mode.

[0211] In one embodiment, the plurality of mixing modes include at least two of a strong real mixing mode, a strong virtual mixing mode, a peer mixing mode, a real off-mixing mode, and a virtual off-mixing mode, and an amplitude value constraint parameter of any one of the audio signals is positively correlated with a volume of the corresponding audio signal after volume adjustment; In the strong reality mixing mode, both the amplitude value constraint parameter of the real audio signal and the amplitude value constraint parameter of the virtual audio signal are valid values, and the amplitude value constraint parameter of the real audio signal is greater than the amplitude value constraint parameter of the virtual audio signal; In the strong virtual mixing mode, both the amplitude value constraint parameter of the real audio signal and the amplitude value constraint parameter of the virtual audio signal are valid values, and the amplitude value constraint parameter of the virtual audio signal is greater than the amplitude value constraint parameter of the real audio signal; In the peer mixing mode, both the amplitude value constraint parameter of the real audio signal and the amplitude value constraint parameter of the virtual audio signal are valid values, the amplitude value constraint parameter of the virtual audio signal is equal to the amplitude value constraint parameter of the real audio signal, and In the real off-mixing mode, the amplitude value constraint parameter of the real audio signal is an invalid value, and the amplitude value constraint parameter of the virtual audio signal is an valid value; In the virtual off-mixing mode, the amplitude value constraint parameter of the real audio signal is a valid value, and the amplitude value constraint parameter of the virtual audio signal is an invalid value.

[0212] In one embodiment, the processor 801 performs signal level monitoring of the current audio signal to obtain the level description value of the current audio signal, specifically by: By performing framing processing on the current audio signal, K audio frames (K is an integer greater than 1) of the current audio signal are obtained; Scan the K audio frames and determine the i-th (i∈[1,K]) audio frame currently being scanned; Obtaining a level description value of the i-th audio frame by monitoring the signal level of the i-th audio frame, or performing a voice validity determination process on the i-th audio frame, and if it is determined that the i-th audio frame is a valid audio frame, obtaining a level description value of the i-th audio frame by monitoring the signal level of the i-th audio frame; The scanning continues until all K audio frames have been scanned to obtain a level description value of the current audio signal, where the level description value of the current audio signal includes the level description value of each monitored audio frame.

[0213] In one embodiment, each audio frame includes a plurality of sampling points and an amplitude value of each sampling point. The processor 801 monitors the signal level of the i-th audio frame to obtain the level description value of the i-th audio frame, specifically by: Calculating the average value of the amplitude values ​​of each sampling point in the i-th audio frame to obtain the average signal amplitude value of the i-th audio frame; The level description value of the i-th audio frame is obtained by smoothing the average signal amplitude value of the i-th audio frame.

[0214] In one embodiment, the processor 801 smoothes the average signal amplitude value of the i-th audio frame to obtain the level description value of the i-th audio frame, specifically by: A reference audio frame for the i-th audio frame is searched for from the K audio frames. If the i-th audio frame has not been subjected to a voice activation determination process, the reference audio frame includes the (i-1)-th audio frame. If the i-th audio frame has been subjected to a voice activation determination process, the reference audio frame includes a valid audio frame that is located before the i-th audio frame and is closest to the i-th audio frame. If no reference audio frame is found, the average signal amplitude value of the i-th audio frame is taken as the level description value of the i-th audio frame; If a reference audio frame is found, the level description value of the i-th audio frame is obtained by smoothing the average signal amplitude value of the i-th audio frame using the level description value of the reference audio frame.

[0215] In one embodiment, the current audio signal is divided into K audio frames (K is an integer greater than 1), the level description value of the current audio signal includes level description values ​​of J audio frames (J≦K) among the K audio frames that have undergone signal level monitoring, and the acquired amplitude value constraint parameters include a target amplitude value and an amplitude value gain threshold; The processor 801 performs a control process for a gain required for the amplitude value of the current audio signal based on the acquired amplitude value constraint parameters and the level description value of the current audio signal, and thereby, when acquiring the amplitude value gain of the current audio signal, specifically, Scanning J audio frames according to the level description values ​​of the current audio signal to determine the jth (j∈[1,J]) audio frame currently being scanned; By performing gain prediction using the target amplitude value in the acquired amplitude value constraint parameter and the level description value of the j-th audio frame, a predicted amplitude value gain is acquired, which indicates a gain required to adjust the level description value of the j-th audio frame to the target amplitude value; determining the smaller one of the amplitude gain threshold in the obtained amplitude constraint parameter and the predicted amplitude gain as the frame amplitude gain required to adjust the amplitude value of the j-th audio frame; By continuing to scan until all of the J audio frames have been scanned, an amplitude value gain of the current audio signal is obtained, and the amplitude value gain of the current audio signal includes a frame amplitude value gain corresponding to each of the J audio frames.

[0216] In one embodiment, the current audio signal is divided into K audio frames (K is an integer greater than 1), and the amplitude value gain of the current audio signal includes frame amplitude value gains of J audio frames (J≦K) among the K audio frames for which signal level monitoring has been performed; The processor 801 adjusts the volume of the current audio signal by amplifying the amplitude value of the current audio signal using the amplitude value gain of the current audio signal, and when obtaining the current audio signal after volume adjustment, specifically: Scanning J audio frames according to the level description values ​​of the current audio signal to determine the jth (j∈[1,J]) audio frame currently being scanned; Dividing the j-th audio frame into N audio sub-blocks (N is an integer greater than 1), and determining an amplitude value for each of the N audio sub-blocks; Stepwise amplifying the amplitude value of each of the N audio sub-blocks based on the frame amplitude value gain of the j-th audio frame to adjust the volume of each audio sub-block; After obtaining each volume-adjusted audio sub-block, the volume-adjusted audio sub-blocks are spliced ​​together to obtain a volume-adjusted j-th audio frame; Scanning continues until all of the J audio frames have been scanned to obtain J audio frames after volume adjustment, and the J audio frames after volume adjustment are connected together to obtain the current audio signal after volume adjustment.

[0217] In one embodiment, the processor 801 further comprises: determining an array position of a j-th audio frame among the J audio frames; If the determined arrangement position is not the first, trigger the execution of a step of gradually amplifying the amplitude value of each of the N audio sub-blocks according to the frame amplitude value gain of the j-th audio frame to adjust the volume of each audio sub-block; If the determined arrangement position is the first, the amplitude value of each of the N audio sub-blocks is amplified using the frame amplitude value gain of the j-th audio frame to adjust the volume of each audio sub-block.

[0218] In one embodiment, when the processor 801 amplifies the amplitude value of each of the N audio sub-blocks stepwise based on the frame amplitude value gain of the j-th audio frame, the processor 801 specifically: Determine a block amplitude value gain for each audio sub-block according to the frame amplitude value gain of the jth audio frame and the arrangement position of each audio sub-block in the N audio sub-blocks, and the block amplitude value gains of different audio sub-blocks change stepwise with the increase of the arrangement position of the corresponding audio sub-block in the N audio sub-blocks; The block amplitude value gain of each audio sub-block is used to amplify the amplitude value of the corresponding audio sub-block.

[0219] In one embodiment, the j-th audio frame in the J audio frames is not the first audio frame; Specifically, when determining the block amplitude value gain of each audio sub-block based on the frame amplitude value gain of the j-th audio frame and the arrangement position of each audio sub-block among the N audio sub-blocks, the processor 801 Determine a frame amplitude value gain of a j-1th audio frame among the J audio frames, and define a difference between the frame amplitude value gain of the jth audio frame and the frame amplitude value gain of the j-1th audio frame as a frame amplitude value gain difference of the jth audio frame relative to the j-1th audio frame; For an n-th (n∈[1,N]) audio sub-block among the N audio sub-blocks, determine a block amplitude value gain change value required for the n-th audio sub-block for the j−1-th audio frame based on an arrangement position of the n-th audio sub-block among the N audio sub-blocks and a frame amplitude value gain difference; The block amplitude value gain of the n-th audio sub-block is obtained based on the block amplitude value gain change value and the frame amplitude value gain of the j-1-th audio frame.

[0220] In one embodiment, before obtaining the mixed signal by mixing the real audio signal and the virtual audio signal in the target mixing mode, the processor 801 further Determine a scene type of the augmented reality scene, and perform sound effect processing on the real audio signal and the virtual audio signal in the augmented reality scene according to the sound effect needs of the scene type; or Obtain an environment type to which the real environment on which the augmented reality scene is based belongs, and perform sound effect processing on the real audio signal and the virtual audio signal in the augmented reality scene according to the sound effect needs of the environment type to which the real environment belongs; Here, the sound effect processing includes at least one of reverberation processing and equalization processing.

[0221] In one embodiment, each of the plurality of mixing modes has a mode identifier, and each mode identifier corresponds to one set of equalization parameters required for performing equalization processing, and the set of equalization parameters corresponding to any one of the mode identifiers includes a first volume equalization parameter of a virtual audio signal in the corresponding mixing mode and a second volume equalization parameter of a real audio signal in the corresponding mixing mode; The processor 801 further performs the following steps before obtaining the mixed signal by mixing the real audio signal and the virtual audio signal in the target mixing mode: Obtaining a set of equalization parameters corresponding to a mode identifier of the target mixing mode; The virtual audio signal is equalized using a first volume equalization parameter in the obtained set of equalization parameters, and the real audio signal is equalized using a second volume equalization parameter in the obtained set of equalization parameters.

[0222] In an embodiment of the present application, a true audio signal and a virtual audio signal in an augmented reality scene can be obtained. A plurality of mixing modes arranged for the augmented reality scene are determined, and one mixing mode from the plurality of mixing modes is selected as a target mixing mode. Different mixing modes can be used to achieve different mixing effects between the true audio signal and the virtual audio signal, resulting in different auditory experiences. Determining a plurality of mixing modes for the augmented reality scene provides a variety of options for setting the target mixing mode, allowing the target mixing mode to be flexibly set. This can accommodate various scenes requiring mixing of the true audio signal and the virtual audio signal, and can also meet the user's customized mixing needs. A mixed signal is obtained by mixing the true audio signal and the virtual audio signal in the target mixing mode, and the mixed signal is output. The target mixing mode can achieve a corresponding mixing effect between real audio signals and virtual audio signals. By controlling the mixing process between real audio signals and virtual audio signals in the target mixing mode, the fusion of real-world sounds and virtual-world sounds on which the augmented reality scene is based can be achieved in the auditory dimension, thereby increasing the auditory fusion of sounds in the augmented reality scene and enriching the fusion capabilities of the augmented reality scene. The resulting mixing signal also has a mixing effect corresponding to the target mixing mode. Finally, the output of the mixing signal allows the user to more intuitively experience the fusion effect of sounds in the augmented reality scene, enhancing the user's immersion in the augmented reality scene.

[0223] In addition, an embodiment of the present application also provides a computer-readable storage medium, which stores a computer program including program instructions. When a processor executes the program instructions, the computer program can perform the method in the embodiment corresponding to FIG. 2 and FIG. 4 above. Therefore, further description will be omitted here. For technical details not disclosed in the embodiment of the computer-readable storage medium of the present application, please refer to the description of the method embodiment of the present application. For example, the program instructions may be arranged to be executed on one computer device, on multiple computer devices located at one location, or on multiple computer devices distributed at multiple locations and connected to each other via a communication network.

[0224] According to one aspect of the present application, there is provided a computer program product including a computer program, the computer program being stored in a computer-readable storage medium. A processor of a computing device reads the computer program from the computer-readable storage medium, and when the processor executes the computer program, the computing device can perform the method in the embodiment corresponding to Figures 2 and 4 above. Therefore, further description will be omitted here.

[0225] As will be understood by those skilled in the art, all or part of the flow of the method according to the above embodiments may be executed by instructing relevant hardware via a computer program. The program may be stored in a computer-readable storage medium. When the program is executed, the flow of each method embodiment as described above is executed. Here, the storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), etc.

[0226] The above is only a preferred embodiment of the present application, and of course, it should not limit the scope of the present application. As can be understood by those skilled in the art, any equivalent modification that realizes all or part of the flow of the above embodiment and complies with the claims of the present application still falls within the scope of the present application. [Explanation of symbols]

[0227] 701 Acquired Units 702 Processing Unit 703 Output Unit 801 processor 802 input devices 803 output device 804 memory

Claims

1. 1. A method of audio processing, comprising: acquiring real audio signals and virtual audio signals in an augmented reality scene, wherein the real audio signals refer to audio signals collected in a real environment on which the augmented reality scene is based, and the virtual audio signals refer to audio signals constructed in a virtual environment on which the augmented reality scene is based; determining a plurality of mixing modes arranged for the augmented reality scene, and selecting one mixing mode from the plurality of mixing modes as a target mixing mode, where different mixing modes are for realizing different mixing effects between the real audio signal and the virtual audio signal; obtaining a mixed signal by mixing the true audio signal and the virtual audio signal in the target mixing mode; and outputting the mixed signal.

1. An audio processing method comprising:

2. The step of selecting one mixing mode from the plurality of mixing modes as a target mixing mode includes: obtaining a mode configuration operation; and selecting one mixing mode from the plurality of mixing modes as a target mixing mode according to the mode configuration operation; The mode arrangement operation includes a selection operation for a plurality of mode identifiers displayed in an augmented reality application, where one mode identifier corresponds to one mixing mode, or the mode arrangement operation includes a mode selection operation controlled by operating a physical key on an augmented reality device.

2. The audio processing method according to claim 1.

3. The step of selecting one mixing mode from the plurality of mixing modes as a target mixing mode includes: determining geographic location information of a real environment on which the augmented reality scene is based, and determining an environment type to which the real environment belongs based on the geographic location information; and selecting one mixing mode from the plurality of mixing modes as a target mixing mode according to sound fusion needs based on the determined environment type.

3. An audio processing method according to claim 1 or 2.

4. The step of obtaining a mixed signal by mixing the true audio signal and the virtual audio signal in the target mixing mode includes: adjusting the volume of each of the real audio signal and the virtual audio signal in the target mixing mode to obtain a volume-adjusted real audio signal and a volume-adjusted virtual audio signal; and obtaining a mixed signal by mixing the volume-adjusted true audio signal and the volume-adjusted virtual audio signal.

4. An audio processing method according to claim 1, wherein the audio processing method is a digital audio processing method.

5. The volume of the audio signal changes with a change in the amplitude value of the audio signal, and the step of adjusting the volume of each of the true audio signal and the virtual audio signal in the target mixing mode to obtain a volume-adjusted true audio signal and a volume-adjusted virtual audio signal includes: Scanning the real audio signal and the virtual audio signal, and taking the currently scanned audio signal as a current audio signal, obtaining amplitude value constraint parameters of the current audio signal in the target mixing mode; obtaining a level description value of the current audio signal by performing signal level monitoring of the current audio signal, wherein the level description value of the current audio signal is for indicating an amplitude value of the current audio signal; obtaining an amplitude value gain of the current audio signal by performing a gain control process required for the amplitude value of the current audio signal based on the obtained amplitude value constraint parameter and the level description value of the current audio signal; adjusting the volume of the current audio signal by amplifying the amplitude value of the current audio signal using the amplitude value gain of the current audio signal, and obtaining a volume-adjusted current audio signal; 5. An audio processing method according to claim 1.

6. The step of obtaining an amplitude value constraint parameter of the current audio signal in the target mixing mode includes: obtaining a predetermined mapping table, the predetermined mapping table including a plurality of mode identifiers and parameter groups mapped to each of the mode identifiers, the parameter group mapped to any one of the mode identifiers including a first amplitude value constraint parameter of a virtual audio signal in a mixing mode identified by the corresponding mode identifier and a second amplitude value constraint parameter of a real audio signal in the mixing mode identified by the corresponding mode identifier; searching the predetermined mapping table using the mode identifier of the target mixing mode to obtain a parameter group mapped to the mode identifier of the target mixing mode; If the current audio signal is the virtual audio signal, selecting a first amplitude constraint parameter from the acquired parameter group as an amplitude constraint parameter of the current audio signal in the target mixing mode; If the current audio signal is the true audio signal, selecting a second amplitude value constraint parameter from the acquired parameter group as the amplitude value constraint parameter of the current audio signal in the target mixing mode.

6. An audio processing method according to any one of claims 1 to 5.

7. The plurality of mixing modes include at least two of a strong real mixing mode, a strong virtual mixing mode, a peer mixing mode, a real off-mixing mode, and a virtual off-mixing mode, and an amplitude value constraint parameter of any one of the audio signals is positively correlated with a volume of the corresponding audio signal after volume adjustment; In the strong reality mixing mode, both the amplitude value constraint parameter of the real audio signal and the amplitude value constraint parameter of the virtual audio signal are valid values, and the amplitude value constraint parameter of the real audio signal is greater than the amplitude value constraint parameter of the virtual audio signal; In the strong virtual mixing mode, both the amplitude value constraint parameter of the real audio signal and the amplitude value constraint parameter of the virtual audio signal are valid values, and the amplitude value constraint parameter of the virtual audio signal is greater than the amplitude value constraint parameter of the real audio signal; In the peer mixing mode, both the amplitude value constraint parameter of the real audio signal and the amplitude value constraint parameter of the virtual audio signal are valid values, and the amplitude value constraint parameter of the virtual audio signal is equal to the amplitude value constraint parameter of the real audio signal; In the real off-mixing mode, the amplitude value constraint parameter of the real audio signal is an invalid value, and the amplitude value constraint parameter of the virtual audio signal is an valid value; In the virtual off-mixing mode, the amplitude value constraint parameter of the real audio signal is a valid value, and the amplitude value constraint parameter of the virtual audio signal is an invalid value.

7. Audio processing method according to any one of claims 1 to 6.

8. the step of obtaining a level description value of the current audio signal by monitoring the signal level of the current audio signal, performing framing processing of the current audio signal to obtain K audio frames of the current audio signal, where K is an integer greater than 1; scanning the K audio frames to determine a currently scanned i-th (i∈[1,K]) audio frame; a step of monitoring the signal level of the i-th audio frame to obtain a level description value of the i-th audio frame, or performing a voice validity determination process on the i-th audio frame, and if it is determined that the i-th audio frame is a valid audio frame, monitoring the signal level of the i-th audio frame to obtain a level description value of the i-th audio frame; obtaining level description values ​​of the current audio signal by continuing scanning until all of the K audio frames have been scanned, wherein the level description values ​​of the current audio signal include level description values ​​of each monitored audio frame.

8. Audio processing method according to any one of claims 1 to 7.

9. Each audio frame includes a plurality of sampling points and an amplitude value of each sampling point, and the step of obtaining a level description value of the i-th audio frame by monitoring a signal level of the i-th audio frame includes: calculating an average value of amplitude values ​​of each sampling point in the i-th audio frame to obtain an average signal amplitude value of the i-th audio frame; and obtaining a level description value of the i-th audio frame by smoothing the average signal amplitude value of the i-th audio frame.

9. Audio processing method according to any one of claims 1 to 8.

10. the step of obtaining a level description value of the i-th audio frame by smoothing an average signal amplitude value of the i-th audio frame, a step of searching for a reference audio frame of the i-th audio frame from the K audio frames, in which if the i-th audio frame has not been subjected to a voice activity determination process, the reference audio frame includes the (i-1)-th audio frame, and if the i-th audio frame has been subjected to a voice activity determination process, the reference audio frame includes a valid audio frame that is located before the i-th audio frame and is closest to the i-th audio frame; If the reference audio frame is not found, the signal amplitude average value of the i-th audio frame is set as the level description value of the i-th audio frame; If the reference audio frame is found, smoothing the average signal amplitude value of the i-th audio frame using the level description value of the reference audio frame to obtain the level description value of the i-th audio frame.

10. Audio processing method according to any one of claims 1 to 9.

11. The current audio signal is divided into K audio frames (K is an integer greater than 1), the level description values ​​of the current audio signal include level description values ​​of J audio frames (J≦K) among the K audio frames for which signal level monitoring has been performed, and the acquired amplitude value constraint parameters include a target amplitude value and an amplitude value gain threshold; the step of acquiring an amplitude value gain of the current audio signal by performing a gain control process required for the amplitude value of the current audio signal based on the acquired amplitude value constraint parameter and the level description value of the current audio signal, scanning J audio frames according to the level description values ​​of the current audio signal to determine a j-th (j∈[1, J]) audio frame to be currently scanned; a step of obtaining a predicted amplitude value gain by performing gain prediction using a target amplitude value in the obtained amplitude value constraint parameters and a level description value of the j-th audio frame, wherein the predicted amplitude value gain means a gain required to adjust the level description value of the j-th audio frame to the target amplitude value; determining a smaller one of an amplitude value gain threshold in the obtained amplitude value constraint parameter and the predicted amplitude value gain as a frame amplitude value gain required to adjust the amplitude value of the j-th audio frame; and obtaining an amplitude value gain of the current audio signal by continuing to scan until all of the J audio frames have been scanned, wherein the amplitude value gain of the current audio signal includes a frame amplitude value gain corresponding to each of the J audio frames.

11. Audio processing method according to any one of claims 1 to 10.

12. the current audio signal is divided into K audio frames (K is an integer greater than 1), and the amplitude value gain of the current audio signal includes frame amplitude value gains of J audio frames (J≦K) among the K audio frames, the J audio frames being subjected to signal level monitoring; the step of adjusting the volume of the current audio signal by amplifying the amplitude value of the current audio signal using the amplitude value gain of the current audio signal and obtaining a volume-adjusted current audio signal, scanning J audio frames according to the level description values ​​of the current audio signal to determine a j-th (j∈[1, J]) audio frame to be currently scanned; Dividing the j-th audio frame into N audio sub-blocks (N is an integer greater than 1), and determining an amplitude value of each of the N audio sub-blocks; amplifying the amplitude value of each of the N audio sub-blocks in a stepwise manner based on the frame amplitude value gain of the j-th audio frame to adjust the volume of each of the audio sub-blocks; After obtaining each volume-adjusted audio sub-block, the volume-adjusted audio sub-blocks are spliced ​​together to obtain a volume-adjusted j-th audio frame; and continuing to scan until all of the J audio frames have been scanned to obtain J audio frames after volume adjustment, and splicing the J audio frames after volume adjustment to obtain a current audio signal after volume adjustment.

12. Audio processing method according to any one of claims 1 to 11.

13. determining an arrangement position of the j-th audio frame among the J audio frames; If the determined arrangement position is not the first, triggering the execution of the step of amplifying the amplitude value of each of the N audio sub-blocks stepwise based on the frame amplitude value gain of the j-th audio frame to adjust the volume of each of the audio sub-blocks; If the determined arrangement position is the first, amplifying the amplitude value of each of the N audio sub-blocks using a frame amplitude value gain of the j-th audio frame to adjust the volume of each of the N audio sub-blocks.

13. Audio processing method according to any one of claims 1 to 12.

14. the step of amplifying the amplitude value of each of the N audio sub-blocks in stages based on the frame amplitude value gain of the j-th audio frame, determining a block amplitude value gain for each audio sub-block based on a frame amplitude value gain of the jth audio frame and an arrangement position of each audio sub-block in the N audio sub-blocks, wherein the block amplitude value gains of different audio sub-blocks change stepwise with an increase in the arrangement position of the corresponding audio sub-block in the N audio sub-blocks; amplifying the amplitude value of the corresponding audio sub-block using a block amplitude value gain of the corresponding audio sub-block; 14. Audio processing method according to any one of claims 1 to 13.

15. The j-th audio frame is not at the first position in the J audio frames; determining a block amplitude value gain for each audio sub-block based on a frame amplitude value gain of the j-th audio frame and an arrangement position of each audio sub-block among the N audio sub-blocks, determining a frame amplitude value gain of a j-1th audio frame among the J audio frames, and setting a difference between the frame amplitude value gain of the jth audio frame and the frame amplitude value gain of the j-1th audio frame as a frame amplitude value gain difference of the jth audio frame relative to the j-1th audio frame; For an n-th (n∈[1,N]) audio sub-block among the N audio sub-blocks, determining a block amplitude value gain change value required for the n-th audio sub-block for the j-1-th audio frame based on an arrangement position of the n-th audio sub-block among the N audio sub-blocks and the frame amplitude value gain difference; obtaining a block amplitude value gain of the n-th audio sub-block based on the block amplitude value gain change value and a frame amplitude value gain of the j-1-th audio frame; 15. Audio processing method according to any one of claims 1 to 14.

16. before the step of obtaining a mixed signal by mixing the true audio signal and the virtual audio signal in the target mixing mode; determining a scene type of the augmented reality scene, and performing sound effect processing on each of the true audio signal and the virtual audio signal in the augmented reality scene according to the sound effect needs of the scene type; or Further comprising: obtaining an environment type to which a real environment on which the augmented reality scene is based belongs; and performing sound effect processing on each of the true audio signal and the virtual audio signal in the augmented reality scene according to the sound effect needs of the environment type to which the real environment belongs; The sound effect processing includes at least one of reverberation processing and equalization processing.

16. Audio processing method according to any one of claims 1 to 15.

17. Each of the plurality of mixing modes has a mode identifier, and each mode identifier corresponds to one set of equalization parameters required for performing equalization processing, and the set of equalization parameters corresponding to any one of the mode identifiers includes a first volume equalization parameter of a virtual audio signal in the corresponding mixing mode and a second volume equalization parameter of a real audio signal in the corresponding mixing mode; before the step of obtaining a mixed signal by mixing the true audio signal and the virtual audio signal in the target mixing mode; obtaining a set of equalization parameters corresponding to a mode identifier of the target mixing mode; performing equalization processing on the virtual audio signal using a first volume equalization parameter in the obtained set of equalization parameters, and performing equalization processing on the real audio signal using a second volume equalization parameter in the obtained set of equalization parameters, 17. Audio processing method according to any one of claims 1 to 16.

18. 1. An audio processing device comprising: an acquisition unit for acquiring real audio signals and virtual audio signals in an augmented reality scene, wherein the real audio signals refer to audio signals collected in a real environment on which the augmented reality scene is based, and the virtual audio signals refer to audio signals constructed in a virtual environment on which the augmented reality scene is based; a processing unit for determining a plurality of mixing modes arranged for the augmented reality scene, selecting one mixing mode from the plurality of mixing modes as a target mixing mode, and performing a mixing process on the real audio signal and the virtual audio signal in the target mixing mode to obtain a mixing signal, wherein different mixing modes are for realizing different mixing effects between the real audio signal and the virtual audio signal; an output unit for outputting the mixed signal.

1. An audio processing device comprising:

19. 1. A computing device comprising: a processor suitable for executing a computer program; a computer-readable storage medium having stored thereon a computer program that, when executed by the processor, causes the audio processing method of any one of claims 1 to 17 to be performed.

1. A computing device comprising:

20. 18. A computer-readable storage medium having stored thereon a computer program that, when executed by a processor, causes the computer to perform the audio processing method of any one of claims 1 to 17.

21. 18. A computer program product comprising a computer program or computer instructions which, when executed by a processor, causes the computer program or computer instructions to implement the audio processing method according to any one of claims 1 to 17.

Citation Information

Patent Citations

  • Automatic sound volume control device and voice communication equipment employing same

    JP2009171208A

  • Audio processing device

    JP2018085663A

  • Audio output device, audio output method, program, and audio system

    WO2017134973A1

  • Inserting imagery from a real environment into a virtual environment

    WO2020051490A1