Audio signal processing method and device, equipment and storage medium

By arranging different audio and image positions in the sound field space for the audio signals of multiple applications, the audio and image overlap problem when multiple applications play audio concurrently is solved, and the effect of users being able to clearly listen to the sounds of each application is achieved.

CN119946543APending Publication Date: 2025-05-06GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311452247.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-02
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In scenes where multiple applications play audio concurrently, it is difficult for users to hear the sounds of different applications clearly, resulting in the problem of overlapping sound and image.

Method used

By obtaining the audio signals created by multiple applications and determining their audio images in the sound field space, the audio signals are output according to these positions, thereby arranging different audio images for the audio signals of different applications.

Benefits of technology

It effectively avoids overlapping audio and image, ensures that users can clearly listen to audio content from multiple applications, and improves users' listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946543A_ABST
    Figure CN119946543A_ABST
Patent Text Reader

Abstract

The invention provides an audio signal processing method and device, equipment and a storage medium. The method comprises the following steps: acquiring audio signals created by a plurality of applications; determining sound image positions of the audio signals created by the plurality of applications in a sound field space; wherein the sound image positions of the audio signals of different applications are different; and outputting the audio signals created by the plurality of applications according to the sound image positions respectively corresponding to the audio signals of the plurality of applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to audio technology, and is related to but not limited to audio signal processing methods and devices, equipment, and storage media. Background Art

[0002] With the development of the digital intelligent era and electronic devices, there are more and more types of electronic devices, and the functions of each electronic device are becoming more and more abundant. Users can use different functions of an electronic device at the same time. For example, taking a mobile phone as an example, in one possible scenario, the user can open a game application to play a game, and at the same time, the user can also open a music application to listen to music, and during this process, the phone may receive an incoming call and ring. Therefore, in a scenario where multiple applications play audio concurrently, whether the user can hear the sound played by each application is the key to affecting the user's hearing. Summary of the invention

[0003] In view of this, the audio signal processing method, device, equipment, and storage medium provided in the present application can overcome the problem of sound and image overlap in the scenario where the sounds of multiple different applications are played simultaneously, so that in the scenario where multiple applications play audio concurrently, the user can hear the sound played by each application clearly, thereby improving the user's hearing experience.

[0004] According to one aspect of an embodiment of the present application, a method for processing an audio signal is provided, comprising: obtaining audio signals created by multiple applications; determining the sound and image positions of the audio signals created by the multiple applications in a sound field space; wherein the sound and image positions of the audio signals of different applications are different; and outputting the audio signals created by the multiple applications according to the sound and image positions respectively corresponding to the audio signals of the multiple applications.

[0005] According to one aspect of an embodiment of the present application, there is provided an audio signal processing device, comprising: an acquisition module, configured to acquire audio signals created by multiple applications; a determination module, configured to determine the sound and image positions of the audio signals created by the multiple applications in the sound field space, respectively; wherein the sound and image positions of the audio signals of different applications are different; and an audio output module, configured to output the audio signals created by the multiple applications according to the sound and image positions respectively corresponding to the audio signals of the multiple applications.

[0006] According to one aspect of an embodiment of the present application, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, the method described in the embodiment of the present application is implemented.

[0007] According to one aspect of an embodiment of the present application, an audio system is provided, comprising the electronic device and a listening device described in the embodiment of the present application; wherein the listening device is used to listen to an audio signal played by the electronic device.

[0008] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method provided in the embodiment of the present application is implemented.

[0009] In an embodiment of the present application, when there are multiple applications that have created audio signals, different sound and image positions are arranged for the audio signals of different applications, so as to overcome the problem of sound and image overlap in the scenario where the sounds of multiple different applications are played simultaneously in the scenario where the multiple applications play the created audio signals concurrently, thereby avoiding the problem that the listener cannot clearly listen to the playback content from multiple applications due to sound and image overlap.

[0010] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and are used together with the specification to illustrate the technical solution of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0012] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0013] Figure 1 Schematic diagram of the implementation process of the audio signal processing method provided in the embodiment of the present application Figure 1 ;

[0014] Figure 2 Schematic diagram of the implementation process of the audio signal processing method provided in the embodiment of the present application Figure 2 ;

[0015] Figure 3 A schematic diagram of an implementation framework of the audio signal processing method provided in an embodiment of the present application;

[0016] Figure 4 A schematic diagram of an example of a holographic audio algorithm provided in an embodiment of the present application;

[0017] Figure 5 A schematic diagram of the positions of sound and images in a sound field space for multiple applications provided in an embodiment of the present application;

[0018] Figure 6 A schematic diagram of the structure of an audio signal processing device provided in an embodiment of the present application;

[0019] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0020] Figure 8 A schematic diagram of the structure of an audio system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the specific technical solution of the present application will be further described in detail below in conjunction with the drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0023] In the following description, reference is made to “some embodiments”, “this embodiment”, “embodiments of the present application” and examples, etc., which describe a subset of all possible embodiments, but it can be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0024] The descriptions of “first, second, third” etc. that appear in the embodiments of the present application are only for illustration and distinction of the description objects. There is no distinction of order, nor do they indicate any special limitation on the number of devices in the embodiments of the present application, and cannot constitute any limitation on the embodiments of the present application.

[0025] In the embodiment of the present application, the audio signal may also be understood as audio data, track data, etc. The sound field space may also be understood as a virtual sound field space.

[0026] The embodiment of the present application provides an audio signal processing method, which is applied to an electronic device. During the implementation process, the electronic device can be various types of devices with audio signal capabilities, for example, the electronic device can include a mobile phone, a tablet computer, a television, a projector, etc. The function implemented by the method can be implemented by calling a program code by a processor in the electronic device. Of course, the program code can be stored in a computer storage medium. It can be seen that the electronic device at least includes a processor and a storage medium.

[0027] Figure 1 Schematic diagram of the implementation process of the audio signal processing method provided in the embodiment of the present application Figure 1 ,like Figure 1 As shown, the method may include the following steps 101 to 103:

[0028] Step 101, obtaining audio signals created by multiple applications;

[0029] Step 102: determine the sound and image positions of the audio signals created by the multiple applications in the sound field space; wherein the sound and image positions of the audio signals of different applications are different.

[0030] In some embodiments, the electronic device may determine the sound image position of the corresponding audio signal from pre-configured mapping information of the application identifier and the sound image position according to the application identifier of the application. Further, in some embodiments, the electronic device may determine the sound image position of the corresponding audio signal from pre-configured mapping information of the application identifier, the type of the audio signal and the sound image position according to the application identifier of the application and the type of the audio signal (such as the left channel signal or the right channel signal).

[0031] In other embodiments, the electronic device may also determine the sound and image position of an audio signal created by an application in the sound field space by: receiving a control signal based on UI interface input, the control signal including an application identifier and a corresponding sound and image position; and determining the sound and image position of the audio signal of the application with the corresponding application identifier based on the control signal.

[0032] In one possible implementation, the UI interface can display a visualized sound field space, and the user can drag an application to a desired position in the sound field space in the UI interface, thereby generating a control signal including a corresponding sound and image position and an application identifier of the application based on the desired position.

[0033] Of course, in the embodiment of the present application, there is no limitation on the method of inputting control signals based on the UI interface. The electronic device supports users to manually drag any application to any position in the visualized sound field space, and to specify the sound and image position of a certain application by inputting text or voice.

[0034] In some further embodiments, the electronic device may also implement step 102 by using steps 202 and 203 described in the following embodiments, that is, to determine the sound image positions of the audio signals created by the multiple applications in the sound field space.

[0035] Step 103: output the audio signals created by the multiple applications according to the sound and image positions respectively corresponding to the audio signals of the multiple applications.

[0036] In some embodiments, the electronic device may implement step 103 as follows: according to the sound and image position of the audio signal of the application, the audio signal is rendered and processed; the audio signals rendered and processed by the multiple applications are superimposed and output, for example, the left channel data in the audio signals rendered and processed by the multiple applications are superimposed (such as linear superposition or nonlinear superposition), and the right channel data in the audio signals rendered and processed by the multiple applications are superimposed (such as linear superposition or nonlinear superposition) and then output dual-channel data; in this way, the sounds of the audio signals of the multiple applications are perceived to be emitted from different positions. Among them, being perceived to be emitted from different positions can be understood as being perceived to be emitted from different directions or angles.

[0037] In an embodiment of the present application, when there are multiple applications that have created audio signals, different sound and image positions are arranged for the audio signals of different applications, so as to overcome the problem of sound and image overlap in the scenario where the sounds of multiple different applications are played simultaneously in the scenario where the multiple applications play the created audio signals concurrently, thereby avoiding the problem that the listener cannot clearly listen to the playback content from multiple applications due to sound and image overlap.

[0038] The present application further provides an audio signal processing method. Figure 2 Schematic diagram of the implementation process of the audio signal processing method provided in the embodiment of the present application Figure 2 ,like Figure 2 As shown, the method includes the following steps 201 to 205:

[0039] Step 201, obtaining audio signals created by multiple applications;

[0040] Step 202, determining the scenario types to which the multiple applications belong respectively;

[0041] Step 203, determining the sound image position of the audio signal created by the application in the sound field space according to the scene type to which the application belongs;

[0042] Step 204, rendering the audio signal according to the sound image position of the audio signal of the application;

[0043] Step 205: superimpose and output the audio signals rendered by the multiple applications.

[0044] The following describes further optional implementations and related terms of each of the above steps.

[0045] In step 201, audio signals created by multiple applications are acquired.

[0046] In the embodiment of the present application, the audio signal can also be understood as data recorded by a track. It can be understood that an audio signal created by an application may include left channel data and / or right channel data.

[0047] In some embodiments, the multiple applications are applications that are allowed to participate in audio spatialization. The electronic device can determine whether to execute the audio signal processing method described in the embodiment of the present application for the application that currently creates the audio signal based on a pre-configured application whitelist.

[0048] In step 202, the scenario types to which the multiple applications respectively belong are determined.

[0049] It can be understood that the so-called application scene type refers to the playback scene type or application type to which the application belongs. The scene type may be game, music, audiobook, navigation, notification, voice, incoming call ringtone, alarm ringtone or video scene type.

[0050] In some embodiments, the electronic device may determine the scenario type to which the application belongs by: determining an application identifier of the application; and obtaining, based on the application identifier of the application, a scenario type recorded corresponding to the application identifier of the application from a pre-configured application whitelist.

[0051] Exemplarily, the application identifier recorded in the pre-configured application whitelist may be an application package name or other information used to uniquely identify the application.

[0052] In an embodiment of the present application, the electronic device supports users to customize application identifiers in an application whitelist and modify application identifiers in the application whitelist at any time.

[0053] In step 203, the sound image position of the audio signal of the application in the sound field space is determined according to the scene type to which the application belongs.

[0054] In the embodiment of the present application, the method for implementing step 203 is not limited. In short, the basis for determining the sound image position of the audio signal of the application in the sound field space at least includes the scene type to which the application belongs. For example, step 203 can be implemented by the following embodiment 1 or embodiment 2.

[0055] Among them, in embodiment 1, the electronic device can obtain the corresponding sound and image position from the holographic audio strategy parameters according to the scene type of the application; wherein the holographic audio strategy parameters include the sound and image position corresponding to at least one scene type.

[0056] In some embodiments, the scene type and the corresponding sound and image position in the holographic audio strategy parameters can be predefined. The predefined holographic audio strategy parameters can be saved in the electronic device in the form of an XML file. The electronic device can load the preconfigured XML file to obtain the holographic audio strategy parameters, which include mapping information between the scene type and the sound and image position. The electronic device can obtain the corresponding sound and image position from the mapping information according to the scene type of the application.

[0057] And / or, in some embodiments, the sound and image position corresponding to the scene type in the holographic audio strategy parameters may also include the sound and image position configuration according to the corresponding scene type previously determined to be used for the audio signal, that is, the sound and image position configuration according to the historical use of the scene type.

[0058] Of course, in a possible implementation, the electronic device also supports users or developers to configure the sound and image positions corresponding to each scene type. That is, in some embodiments, the electronic device receives configuration information, the configuration information includes specifying the sound and image position of at least one scene type; and, according to the configuration information, configures the holographic audio strategy parameters; in this way, the user or developer can customize the sound and image position of any scene type by inputting the configuration information. In particular, for users, the electronic device supports users to configure the sound and image positions corresponding to each scene type, so that the hearing sense of spatial audio can be customized to meet the personalized hearing needs of users.

[0059] In the embodiment of the present application, there is no limitation on the method for the electronic device to receive configuration information. It can be based on the UI interface of the electronic device to receive configuration information input by the user, or it can be based on the communication module of the electronic device to receive configuration information sent remotely by the developer.

[0060] In another possible implementation, the electronic device also supports users or developers to change the sound and image position corresponding to the scene type in the holographic audio strategy parameters anytime and anywhere. That is, in some embodiments, the electronic device receives an update signal, the update signal is used to indicate the update of the sound and image position of at least one specific scene type, the update signal includes the identification information of the at least one specific scene type and the new sound and image position; and according to the update signal, the sound and image position of the at least one specific scene type in the holographic audio strategy parameters is updated. The sound and image position of the specific scene type in the holographic audio strategy parameters is updated to the new sound and image position corresponding to the update signal.

[0061] In the embodiment of the present application, there is no limitation on the method for the electronic device to receive the update signal. It can be based on the UI interface of the electronic device to receive the update signal input by the user, or it can be based on the communication module of the electronic device to receive the update signal sent remotely by the developer.

[0062] In the embodiment of the present application, there is no limitation on the circumstances under which the electronic device obtains the corresponding sound and image position from the holographic audio strategy parameter according to the scene type of the application. In any case, the electronic device obtains the corresponding sound and image position from the holographic audio strategy parameter based on the scene type of the application.

[0063] In other embodiments, the electronic device may also obtain the corresponding sound and image position from the holographic audio strategy parameters according to the scene type of the application during the initialization phase or when the sound and image state of the audio signal of the application is in a non-transition phase; wherein the non-transition phase refers to the sound and image position of the audio signal of the same audio track in different time periods (such as different frames) remains unchanged.

[0064] It can be understood that the so-called initialization phase refers to the situation where the electronic device is turned on and the multiple applications are started. The so-called time period can be understood as a unit data amount of a cycle of reading, writing / processing an audio signal, for example, the unit data amount is a frame of audio signal or other unit data amount.

[0065] Among them, in embodiment 2, the electronic device can also determine the sound and image position of the audio signal in the sound field space according to the scene type to which the application belongs and the order in which the audio signal is created.

[0066] Similarly, in the embodiments of the present application, there is no restriction on the circumstances under which the electronic device determines the sound and image position of the audio signal according to the scene type of the application and the order in which the audio signal is created. Regardless of the circumstances, the electronic device can use this method to determine the sound and image position of the audio signal. In other embodiments, the electronic device can also determine the sound and image position of the audio signal in the sound field space according to the scene type to which the application belongs and the order in which the audio signal is created when the sound and image state of the audio signal of the application is in a transitional stage; wherein the transitional stage refers to the change in the sound and image position of the audio signal in different time periods (e.g., different frames) of the same audio track.

[0067] Furthermore, in some embodiments, the electronic device may assign a corresponding transition state to the audio signal according to the scene type to which the application belongs and the order in which the audio signal is created; and determine the sound and image position of the audio signal according to the transition state of the audio signal.

[0068] For example, depending on the scene type and the creation order of the audio signal, the transition state of the audio signal can be start, fade in, fade out, center, or home. That is to say, different scene types and different creation orders of the audio signal give different transition states to the audio signal. And the sound and image positions corresponding to different transition states are different. For example, if the scene type is a music scene and the creation order is the last one created, the transition state of the audio signal of this scene type is center, and the corresponding sound and image position is the sound and image position corresponding to the center.

[0069] In a possible implementation, for an audio signal belonging to a center position, the sound image position thereof may be determined according to the number of channels of the audio signal.

[0070] In step 204, the audio signal is rendered according to the sound image position of the applied audio signal.

[0071] In some embodiments, the electronic device can input the audio signals of the multiple applications and the corresponding sound and image positions into a spatial audio rendering algorithm, so that the spatial audio rendering algorithm performs spatial processing on the corresponding audio signals according to the sound and image positions, and outputs an audio signal with the sound and image positions. For example, the spatial audio rendering algorithm selects a corresponding head-related transfer function (HRTF) according to the input sound and image positions, and renders the corresponding audio signal through the HRTF function. Among them, the parameters of the HRTF function corresponding to different sound and image positions are different.

[0072] In step 205, the audio signals rendered by the multiple applications are superimposed and then output.

[0073] In some embodiments, the electronic device may implement step 205 by superimposing left channel data of audio signals rendered by multiple applications and superimposing right channel data of audio signals rendered by multiple applications to obtain dual channel data for output.

[0074] In some embodiments, the plurality of applications include a first application and a second application; wherein,

[0075] When the priority of the first application is higher than the priority of the second application, the sound image position of the first audio signal of the first application is the center position in the sound field space, and the sound image position of the first audio signal of the second application is the scene position in the sound field space; wherein the center position is different from the scene position.

[0076] Further, in some embodiments, the priority of the first application may refer to the priority of the scenario type to which the first application belongs, and the priority of the second application may refer to the priority of the scenario type to which the second application belongs.

[0077] In some embodiments, the initial value in the preconfigured holographic audio strategy parameter may define the corresponding sound and image position according to the priority of the scene type.

[0078] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.

[0079] Figure 3 A schematic diagram of an implementation framework of the audio signal processing method provided in the embodiment of the present application, such as Figure 3 As shown, where:

[0080] 1. Start the spatial information management service (SpService) process when booting up the computer.

[0081] The SpService process is used for spatial information management services. It is started automatically at boot time through init.rc and resides as a native service.

[0082] a. During the initialization of the SpService process, the SpService process parses the XML file to load the application whitelist and holographic audio policy parameters;

[0083] b. Start the Audioserver process when booting. Here, the Audioserver process obtains the SpService service AIDL binder object to facilitate subsequent cross-process access to SpService.

[0084] c. When creating a track, you need to add the holographic audio information of the track. At this time, access SpService to obtain MetaAudioInfo (holographic audio stream information) for subsequent use. Holographic audio information includes the following information: application package name, scene type Scene (such as games, music, audiobooks, navigation, notifications, voice, incoming ringtones, alarm ringtones or videos, etc.), audio and video three-dimensional position pos (x, y, z) and fade-in / fade-out type (fadeInStyle / fadeOutStyle).

[0085] 2. Add holographic algorithm instances to each necessary playback thread:

[0086] For the playback threads corresponding to each channel (Priamry / Fast, Deepbuffer, Lhdc and Voiprx) of the AudioServer process, when the mixer playback threads such as MixerThread and SpatializerThread are started, the holographic rendering algorithm instance Normal OplusMetaAudio is added when the mixer AudioMixer is created (normal holographic rendering algorithm instance, general channel for media playback, moderate buffer, low latency requirement, suitable for channels such as priamry, deepbuffer and Lhdc); the AudioMixer held by fastMixer used in the Fast channel also loads FastOplusMetaAudio (lightweight, small buffer, high latency requirement); this type of voip_rx goes through the direct channel DirectOutputThread, but the thread has no mixer, and directly creates DirecOplusMetaAudio (based on NormalOplusMetaAudio, necessary format conversion, upmixing channels and resampling are added before process) for holographic rendering mixing.

[0087] 3. AudioMixer transformed into a 3D mixer

[0088] The mixer adds a holographic audio algorithm instance, that is, an audio signal processing method, Figure 4 The spatial sound algorithm library in is OplusMetaAudio.

[0089] Figure 4 A schematic diagram of an example of a holographic audio algorithm provided in an embodiment of the present application, such as Figure 4 As shown, during the mix process, the data of each track and the holographic audio stream information (obtained from spservice through audioFlinger) are obtained from the playback thread corresponding to the track. These data and holographic audio stream information are collected and passed to OplusMetaAudio for holographic audio algorithm processing.

[0090] 3. Oplus MetaAudio audio and video processing:

[0091] OplusMetaAudio is an important module for audio and video processing, which is used to realize audio and video control and rendering. The audio tracks created by its application will be marked as a specific scene (spatial information annotation) by the system, and then the audio and video control algorithm will assign different transition states (start, fadein, fadeout, center, home, etc.) to each track according to the scene to which the track belongs, the creation order and other information, and treat each track as a separate sound object, and calculate the position of each channel in the track according to the different states of the track. Finally, the spatial audio algorithm uses the audio data and position information in the track as input, and renders each channel in each track to the specified position in the virtual sound field space.

[0092] Summary: Figure 3 and Figure 4 As shown, the holographic audio algorithm is designed to provide spatial rendering and audio-visual arrangement capabilities for multiple audio streams. Since different audio stream types may go to different channels for Android audio playback, and audio-visual arrangement needs to be done for each track played at the same time, a similar mixing action is required to get all tracks of the same mixer and add holographic audio stream information. The holographic audio algorithm performs spatial rendering and audio-visual arrangement based on this, and the holographic audio algorithm needs to be added to AudioMixer. Since the mixer of each channel is separate, the holographic audio algorithm needs to be added to the AudioMixer mixer of the main thread. The input of OplusMetaAudio is all sound source track data, and the output is two-channel data after 3D mixing.

[0093] In the embodiment of the present application, Figure 3 The holographic audio framework shown, that is, the implementation framework of the audio signal processing method, can distribute different types of sound sources in different spatial sound images (such as Figure 5 As shown), there will be no problem of sound and image overlap when two different types of sounds are played at the same time, resulting in unclear hearing and loss of important content, providing a holographic immersive experience.

[0094] For example, Figure 5 A schematic diagram of the positions of sound images in a sound field space for multiple applications provided in the embodiments of the present application, such as Figure 5 As shown, audio source data such as calls, incoming calls, navigation voice, alarms, music, and notifications have different sound and image positions in the sound field space 50; among them, 501 is the position of the listening device headphones in the sound field space 50.

[0095] In the embodiments of the present application, it is proposed to develop from the spatialization of a single sound source to the spatialization of multiple audio streams; and provide users with the ability to customize the sound and image position of a specific sound source, that is, users can independently adjust the sound and image position of each application suitable for their own scenes through the UI interface; and provide the ability to render multiple sound sources and uniformly distribute sound and image; and when multiple audio streams of different scenes are concurrent, the concept of center position (center) preemption is introduced. When the audio stream of a high-priority scene starts to play, the audio stream of a low-priority scene slowly fades out from the center position to the sound and image position where it should be; and two modes are provided for sound and image arrangement, one is the intelligent mode, that is, the sound and image arrangement of different scenes is realized by configuring the xml file, and the other is that users can customize the sound and image position of different applications in the sound field space.

[0096] It should be noted that although the steps of the method in the present application are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all the steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps, etc.; or, steps in different embodiments may be combined into a new technical solution.

[0097] Based on the foregoing embodiments, an embodiment of the present application provides an audio signal processing device, which includes the modules included and the units included in the modules, which can be implemented by a processor; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be an AI acceleration engine (such as NPU, etc.), GPU, central processing unit (CPU), microprocessor (MPU), digital signal processor (DSP) or field programmable gate array (FPGA), etc.

[0098] Figure 6 A schematic diagram of the structure of an audio signal processing device provided in an embodiment of the present application is shown in FIG. Figure 6 As shown, the audio signal processing device 60 includes:

[0099] An acquisition module 601 is configured to acquire audio signals created by multiple applications;

[0100] The determination module 602 is configured to determine the sound image positions of the audio signals created by the multiple applications in the sound field space; wherein the sound image positions of the audio signals of different applications are different;

[0101] The audio output module 603 is configured to output the audio signals created by the multiple applications according to the sound and image positions respectively corresponding to the audio signals of the multiple applications.

[0102] In some embodiments, outputting the audio signals created by the multiple applications according to the sound and image positions respectively corresponding to the audio signals of the multiple applications includes: rendering the audio signals according to the sound and image positions of the audio signals of the applications; and superimposing and outputting the rendered audio signals of the multiple applications.

[0103] In some embodiments, determining the sound and image positions of the audio signals of the multiple applications in the sound field space respectively includes: determining the scene types to which the multiple applications respectively belong; and determining the sound and image positions of the audio signals of the applications in the sound field space according to the scene types to which the applications belong.

[0104] Further, in some embodiments, determining the scenario type to which the application belongs includes: determining an application identifier of the application; and obtaining, based on the application identifier of the application, from a pre-configured application whitelist a recorded scenario type corresponding to the application identifier of the application.

[0105] Further, in some embodiments, determining the sound and image position of the audio signal of the application in the sound field space according to the scene type to which the application belongs includes: obtaining the sound and image position corresponding to the scene type of the application from a holographic audio strategy parameter according to the scene type of the application; wherein the holographic audio strategy parameter includes the sound and image position corresponding to at least one scene type.

[0106] In some embodiments, the sound and image positions corresponding to at least one scene type in the holographic audio strategy parameters are predefined, and / or are configured according to the sound and image positions of the corresponding scene type used by the audio signal previously determined.

[0107] In some embodiments, the audio signal processing device 60 also includes a receiving module and a configuration module; wherein the receiving module is configured to receive configuration information, wherein the configuration information includes specifying the sound and image position of at least one scene type; and the configuration module is configured to configure the holographic audio strategy parameters according to the configuration information.

[0108] In some embodiments, the audio signal processing device 60 also includes a receiving module and an updating module; wherein the receiving module is configured to receive an update signal, the update signal is used to indicate an update of the sound and image position of at least one specific scene type, the update signal includes identification information of the at least one specific scene type and a new sound and image position; the updating module is configured to update the sound and image position of the at least one specific scene type in the holographic audio strategy parameters according to the update signal.

[0109] In some embodiments, the determination module 602 is configured to obtain the corresponding sound and image position from the holographic audio strategy parameters according to the scene type of the application during the initialization phase or when the sound and image state of the audio signal of the application is in a non-transition phase; wherein the non-transition phase refers to the sound and image position of the audio signal of the same audio track in different time periods remains unchanged.

[0110] In some embodiments, determining the sound and image position of the audio signal of the application in the sound field space according to the scene type to which the application belongs includes: determining the sound and image position of the audio signal in the sound field space according to the scene type to which the application belongs and the order in which the audio signals are created.

[0111] Further, in some embodiments, the determination module 602 is configured to: when the sound and image state of the audio signal of the application is in a transition stage, determine the sound and image position of the audio signal in the sound field space according to the scene type to which the application belongs and the order in which the audio signal is created; wherein the transition stage refers to the change in the sound and image position of the audio signal in different time periods of the same audio track.

[0112] In some embodiments, determining the sound and image position of the audio signal in the sound field space according to the scene type to which the application belongs and the order in which the audio signal is created includes: assigning the audio signal a corresponding transition state according to the scene type to which the application belongs and the order in which the audio signal is created; and determining the sound and image position of the audio signal according to the transition state of the audio signal.

[0113] In some embodiments, the multiple applications include a first application and a second application; wherein, when the priority of the first application is higher than the priority of the second application, the sound image position of the audio signal of the first application is a center position in the sound field space, and the sound image position of the audio signal of the second application is a scene position in the sound field space; wherein the center position is different from the scene position.

[0114] In some embodiments, the sound image position of the audio signal belonging to the center position is determined based on the number of channels of the audio signal belonging to the center position.

[0115] In some embodiments, determining the sound and image positions of the audio signals created by the application in the sound field space includes: determining the sound and image positions of the audio signals of the application from pre-configured mapping information between the application identifier and the sound and image position according to the application identifier of the application.

[0116] In some embodiments, determining the sound and image position of an audio signal created by an application in a sound field space includes: receiving a control signal based on a UI interface input, the control signal including an application identifier and a corresponding sound and image position; and determining the sound and image position of the audio signal of the application of the corresponding application identifier according to the control signal.

[0117] The description of the above device embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of the present application, please refer to the description of the method embodiment of the present application for understanding.

[0118] It should be noted that the division of modules in the embodiments of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, each functional unit in each embodiment of the present application may be integrated into a processing unit, or may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit. It may also be implemented in the form of a combination of software and hardware.

[0119] It should be noted that in the embodiment of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application can be essentially or partly embodied in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium, including several instructions to enable an electronic device to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.

[0120] An embodiment of the present application provides an electronic device, Figure 7 A schematic diagram of a hardware entity of an electronic device provided in an embodiment of the present application, such as Figure 7 As shown, the electronic device 70 includes a memory 701 and a processor 702. The memory 701 stores a computer program that can be run on the processor 702. When the processor 702 executes the program, the steps in the method provided in the above embodiment are implemented.

[0121] It should be noted that the memory 701 is configured to store instructions and applications executable by the processor 702, and can also cache data to be processed or processed by the processor 702 and various modules in the electronic device 70 (for example, image data, audio data, voice communication data, and video communication data), which can be implemented through flash memory (FLASH) or random access memory (Random Access Memory, RAM).

[0122] The present application embodiment provides an audio system, Figure 8 A schematic diagram of the structure of the audio system provided in the embodiment of the present application is shown in FIG. Figure 8 As shown, the audio system 80 includes the electronic device 70 and a listening device 801 ; wherein the listening device 801 is used to listen to the audio signal played by the electronic device 80 .

[0123] An embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the method provided in the above embodiment are implemented.

[0124] An embodiment of the present application provides a computer program product including instructions, which, when executed on a computer, enables the computer to execute the steps of the method provided in the above method embodiment.

[0125] It should be noted here that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium, storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0126] It should be understood that "one embodiment" or "an embodiment" or "some embodiments" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in one embodiment" or "in some embodiments" appearing throughout the specification may not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application. The above-mentioned sequence numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments. The above description of each embodiment tends to emphasize the differences between the various embodiments, and the same or similar aspects can be referenced to each other. For the sake of brevity, this article will not repeat them.

[0127] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there may be three relationships. For example, object A and / or object B can represent three situations: object A exists alone, object A and object B exist at the same time, and object B exists alone.

[0128] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0129] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as: multiple modules or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be electrical, mechanical or other forms.

[0130] The modules described above as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules; they may be located in one place or distributed on multiple network units; some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment.

[0131] In addition, all functional modules in the embodiments of the present application may be integrated into one processing unit, or each module may be a separate unit, or two or more modules may be integrated into one unit; the above-mentioned integrated modules may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0132] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, etc., various media that can store program codes.

[0133] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can essentially or in other words, the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling an electronic device to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0134] The methods disclosed in several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0135] The features disclosed in several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0136] The features disclosed in several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0137] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. An audio signal processing method, characterized in that: The method comprises: Get audio signals created by multiple applications; Determining the sound image positions of the audio signals created by the multiple applications in the sound field space respectively; wherein the sound image positions of the audio signals of different applications are different; The audio signals created by the multiple applications are output according to the sound image positions respectively corresponding to the audio signals of the multiple applications.

2. The method according to claim 1, characterized in that The step of outputting the audio signals created by the multiple applications according to the sound and image positions respectively corresponding to the audio signals of the multiple applications comprises: Rendering the audio signal according to the sound and image position of the audio signal of the application; The audio signals rendered by the multiple applications are superimposed and then output.

3. The method according to claim 1 or 2, wherein: The determining of the sound image positions of the audio signals of the plurality of applications in the sound field space respectively comprises: Determining the scenario types to which the multiple applications respectively belong; According to the scene type to which the application belongs, a sound image position of the audio signal of the application in the sound field space is determined.

4. The method according to claim 3, characterized in that Determine the scenario type to which the application belongs, including: Determining an application identifier of the application; According to the application identifier of the application, a recorded scenario type corresponding to the application identifier of the application is obtained from a pre-configured application whitelist.

5. The method according to claim 3, characterized in that: The step of determining the sound image position of the audio signal of the application in the sound field space according to the scene type to which the application belongs includes: According to the scene type of the application, the sound and image position corresponding to the scene type of the application is obtained from the holographic audio strategy parameters; wherein the holographic audio strategy parameters include the sound and image position corresponding to at least one scene type.

6. The method according to claim 5, characterized in that The sound and image positions corresponding to at least one scene type in the holographic audio strategy parameters are predefined and / or are configured according to the sound and image positions of the corresponding scene type used by the audio signal previously determined.

7. The method according to claim 5, characterized in that The method further comprises: receiving configuration information, the configuration information including specifying an audio and video position of at least one scene type; According to the configuration information, configure the holographic audio strategy parameters.

8. The method according to claim 5, characterized in that The method further comprises: receiving an update signal, the update signal being used to indicate updating of the sound and image position of at least one specific scene type, the update signal comprising identification information of the at least one specific scene type and a new sound and image position; According to the update signal, the sound and image position of the at least one specific scene type in the holographic audio strategy parameters is updated.

9. The method according to claim 5, characterized in that During the initialization phase or when the sound and image state of the audio signal of the application is in a non-transition phase, the corresponding sound and image position is obtained from the holographic audio strategy parameters according to the scene type of the application; wherein the non-transition phase refers to the sound and image position of the audio signal of the same audio track in different time periods remains unchanged.

10. The method according to claim 3, characterized in that: The step of determining the sound image position of the audio signal of the application in the sound field space according to the scene type to which the application belongs includes: According to the scene type to which the application belongs and the order in which the audio signal is created, the sound image position of the audio signal in the sound field space is determined.

11. The method according to claim 10, characterized in that When the sound and image state of the audio signal of the application is in a transition stage, the sound and image position of the audio signal in the sound field space is determined according to the scene type to which the application belongs and the order in which the audio signal is created; wherein the transition stage refers to the change in the sound and image position of the audio signal in different time periods of the same audio track.

12. The method according to claim 11, characterized in that The step of determining the sound image position of the audio signal in the sound field space according to the scene type to which the application belongs and the order in which the audio signal is created includes: assigning a corresponding transition state to the audio signal according to the scene type to which the application belongs and the order in which the audio signal is created; The sound image position of the audio signal is determined according to the transition state of the audio signal.

13. The method according to claim 1, characterized in that The multiple applications include a first application and a second application; wherein, When the priority of the first application is higher than the priority of the second application, the sound image position of the audio signal of the first application is the center position in the sound field space, and the sound image position of the audio signal of the second application is the scene position in the sound field space; wherein the center position is different from the scene position.

14. The method according to claim 13, characterized in that The sound image position of the audio signal belonging to the center position is determined based on the number of channels of the audio signal belonging to the center position.

15. The method according to claim 1, characterized in that Determining the sound image positions of the audio signals created by the application in the sound field space, respectively, includes: According to the application identifier of the application, the sound image position of the audio signal of the application is determined from pre-configured mapping information between the application identifier and the sound image position.

16. The method according to claim 1, characterized in that Determine the image position of the audio signal created by the application in the sound field space, including: Receiving a control signal based on a UI interface input, the control signal including an application identifier and a corresponding sound and image position; According to the control signal, the sound image position of the audio signal of the application corresponding to the application identifier is determined.

17. An audio signal processing device, characterized in that: include: An acquisition module configured to acquire audio signals created by multiple applications; A determination module, configured to determine the sound image positions of the audio signals created by the multiple applications in the sound field space; wherein the sound image positions of the audio signals of different applications are different; The audio output module is configured to output the audio signals created by the multiple applications according to the sound and image positions respectively corresponding to the audio signals of the multiple applications.

18. An electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 16 is implemented.

19. An audio system comprising the electronic device and the listening device according to claim 16; wherein: The listening device is used to listen to the audio signal played by the electronic device.

20. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 16 is implemented.