Sound source data processing method and device, equipment and storage medium
By determining the audio-image position of the sound source data in the sound field space and performing corresponding processing, the problem of sound interference in multi-source scenes is solved, and the sound separation effect and immersive audio experience are achieved.
Patent Information
- Application Number
- CN202311452168.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-02
- Publication Date
- 2025-05-06
AI Technical Summary
When mixing sounds in multi-sound source scenarios, the sounds interfere with each other seriously, resulting in a decline in the audio experience. Especially in scenes where songs and game sounds are difficult to distinguish between songs and game sounds.
By determining the sound image position of at least one sound source data in the sound field space, determining the second sound source data based on the position and sound source data, and performing mixing and playing, the sound separation effect is achieved in a multi-sound source scene.
This makes the sound source data played after mixing have a spatial sense, improves the sound separation, and users can hear different types of sounds clearly in scenes where multiple sound sources are superimposed and playback, and enjoy an immersive audio experience.
Smart Images

Figure CN119946541A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to audio processing technology, and is related to but not limited to sound source data processing methods and devices, equipment, and storage media. Background Art
[0002] The related mixing methods mainly use the classic one-dimensional linear superposition method to mix multiple sound source data. However, this method interferes with each other when multiple sound source data are played concurrently, which greatly reduces the audio experience. For example, when listening to music and playing games at the same time, the song and game sounds overlap, making it difficult to hear both types of sounds. Summary of the invention
[0003] In view of this, the sound source data processing method, device, equipment, and storage medium provided in the present application enable at least one sound source data played after mixing to have a sense of space. Especially for multi-sound source scenarios, the method can make the final played sound have a sense of sound separation, so that in the scenario where multiple sound sources are superimposed and played, the user can hear different types of sounds clearly and enjoy the immersive audio experience under multiple sound sources.
[0004] According to one aspect of an embodiment of the present application, a method for processing sound source data is provided, including: determining the sound and image position of at least one first sound source data in the sound field space; determining second sound source data based on the sound and image position of the first sound source data and the first sound source data; and mixing the second sound source data of the at least one first sound source data and then playing it.
[0005] According to one aspect of an embodiment of the present application, a sound source data processing device is provided, including: a first determination module, configured to determine the sound and image position of at least one first sound source data in the sound field space; a second determination module, configured to determine the second sound source data based on the sound and image position of the first sound source data and the first sound source data; and a mixing and playing module, configured to mix and play the second sound source data of the at least one first sound source data.
[0006] According to one aspect of an embodiment of the present application, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program executable on the processor, and when the processor executes the program, the method described in the embodiment of the present application is implemented.
[0007] According to one aspect of an embodiment of the present application, an audio system is provided, comprising the electronic device and a listening device described in the embodiment of the present application; wherein the listening device is used to listen to the mixed audio source data played by the electronic device.
[0008] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method provided in the embodiment of the present application is implemented.
[0009] In an embodiment of the present application, before mixing the at least one first sound source data, different sound and image positions are arranged for the one or more first sound source data; based on this, the second sound source data is determined according to the sound and image position of the first sound source data and the first sound source data; finally, the second sound source data of the at least one sound source data is mixed and played. Since the sound source data for mixing is determined based on the sound and image position and the first sound source data, the at least one sound source data played after mixing has a sense of space. Especially for multi-sound source scenes, this method can make the sound that is finally played have a sense of sound separation, so that in the scene where multiple sound sources are superimposed and played, the user can hear different types of sounds and enjoy the immersive audio experience under multiple sound sources.
[0010] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and are used together with the specification to illustrate the technical solution of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0012] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.
[0013] Figure 1 Schematic diagram of the implementation process of the sound source data processing method provided in the embodiment of the present application Figure 1 ;
[0014] Figure 2 This is a schematic diagram of the original 3D mixing framework;
[0015] Figure 3 A schematic diagram of a modified original 3D mixing framework provided in an embodiment of the present application;
[0016] Figure 4 A schematic diagram of a flow chart of a method for determining the position of sound and image provided in an embodiment of the present application;
[0017] Figure 5 A schematic diagram of the relationship between the center position and the scene position provided in an embodiment of the present application;
[0018] Figure 6 A schematic diagram of a method implementation flow for implementing step 103 provided in an embodiment of the present application;
[0019] Figure 7 Schematic diagram of the implementation process of the sound source data processing method provided in the embodiment of the present application Figure 2 ;
[0020] Figure 8 Schematic diagram of the implementation process of the sound source data processing method provided in the embodiment of the present application Figure 3 ;
[0021] Fig. 9 Schematic diagram of the implementation process of the sound source data processing method provided in the embodiment of the present application Figure 4 ;
[0022] Fig.10 A schematic diagram of the implementation process of the scenario selection strategy provided in the embodiment of the present application;
[0023] Fig.11 A schematic diagram of the timing of track 3D information marking provided in an embodiment of the present application;
[0024] Fig.12 A schematic diagram of a 3D audio mixing framework provided in an embodiment of the present application;
[0025] Fig.13 A schematic diagram of the AudioMixer modification process provided in the embodiment of the present application;
[0026] Fig.14 A schematic diagram of DirectOutput transformation processing provided in an embodiment of the present application;
[0027] Fig.15 A schematic diagram of DirectOutput transformation processing provided in an embodiment of the present application;
[0028] Fig.16 A schematic diagram of the structure of a sound source data processing device provided in an embodiment of the present application;
[0029] Fig.17 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0030] Fig.18 A schematic diagram of the structure of an audio system provided in an embodiment of the present application. DETAILED DESCRIPTION
[0031] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the specific technical solution of the present application will be further described in detail below in conjunction with the drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.
[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0033] In the following description, reference is made to “some embodiments”, “this embodiment”, “embodiments of the present application” and examples, etc., which describe a subset of all possible embodiments, but it can be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0034] The descriptions of “first, second, third” etc. that appear in the embodiments of the present application are only for illustration and distinction of the description objects. There is no distinction of order, nor do they indicate any special limitation on the number of devices in the embodiments of the present application, and cannot constitute any limitation on the embodiments of the present application.
[0035] The "sound source data" that appears in the embodiments of the present application can be understood as "track data".
[0036] Currently, there is a lack of information about 3D mixing in the field of audio systems. The main mixing method is to mix multiple sound source scenes using the classic one-dimensional linear superposition method. In terms of spatial performance, audio systems mainly focus on the spatial audio technology and experience of a single sound source.
[0037] When multiple audio sources are playing, the mobile audio system will continuously process the data written by different applications in a loop. The mixing method of the relevant Android system is usually: first, set the mixer's volume, sampling rate, audio format, number of channels, destination cache and other parameters in the preparation stage of each loop. Furthermore, within the mixer, select the functions that meet the requirements according to the various states of the audio source to perform digital audio re-sampling, volume adjustment and superposition processing. The processing functions pre-defined by the Android system here are mainly for different audio source quantity and format scenarios, such as ordinary resampling scenarios, one-channel audio stream dual-channel non-resampling scenarios, etc. In terms of functions, they only have simple mixing capabilities and no effect rendering related capabilities.
[0038] In terms of spatial performance, Dolby Atmos and Android spatial audio solutions are mainly aimed at multi-channel sound sources. They take advantage of the natural data of multi-channels and restore the spatial information of the sound source through the corresponding sound rendering algorithm after mixing, thereby achieving an immersive spatial audio experience. Subsequently, various manufacturers further expanded the application scenarios of spatial audio, and realized the spatial audio experience of dual-channel data by mixing the specified sound source into 12 channels, and there is no longer an application condition that the sound source must be multi-channel.
[0039] From this analysis, we can see that the above solutions mainly focus on completing the original function of mixing and improving the spatial audio experience of a single sound source, but there are still some problems:
[0040] 1. The simple linear superposition mixing method makes the data of multiple sound sources mixed together, making it difficult to separate the data and interfering with each other, which greatly reduces the audio experience. For example, when listening to music and playing games at the same time, the music and game sounds overlap, making it difficult to hear both types of sounds clearly.
[0041] Second, the related spatial audio solutions only consider the single sound source scenario. Since the mixed data cannot be separated, the sound separation effect cannot be achieved in the multi-sound source scenario. The sounds interfere with each other, resulting in the inability of the spatial audio effect to work.
[0042] 3. The rendering processing of related spatial audio solutions conflicts with the effect processing of traditional scene sound effects such as music and cinema, resulting in the inability to customize traditional sound effects.
[0043] Based on the above analysis, an embodiment of the present application provides a method for processing sound source data, which is applied to an electronic device. During the implementation, the electronic device can be various types of devices with audio processing, for example, the electronic device can include a mobile phone, a tablet computer, a desktop computer, a television or a projector, etc. The function implemented by the method can be implemented by calling a program code by a processor in the electronic device. Of course, the program code can be stored in a computer storage medium. It can be seen that the electronic device at least includes a processor and a storage medium.
[0044] Figure 1 Schematic diagram of the implementation process of the sound source data processing method provided in the embodiment of the present application Figure 1 ,like Figure 1 As shown, the method may include the following steps 101 to 103:
[0045] Step 101, determining the sound image position of at least one first sound source data in the sound field space;
[0046] Step 102, determining second sound source data according to the sound image position of the first sound source data and the first sound source data;
[0047] Step 103: Mix the second sound source data of the at least one first sound source data and then play it.
[0048] It can be understood that in the embodiment of the present application, since the sound source data for mixing is determined based on the sound and image position and the first sound source data, at least one sound source data played after mixing has a sense of space. Especially for multi-sound source scenes, this method can make the final played sound have a sense of sound separation, so that in the scene where multiple sound sources are superimposed and played, the user can hear different types of sounds and enjoy the immersive audio experience under multiple sound sources.
[0049] The following describes further optional implementations and related terms of each of the above steps.
[0050] In an embodiment of the present application, the at least one first sound source data may be one first sound source data, or may be two or more first sound source data. Before the electronic device executes step 101, that is, before executing the step of determining the sound image position of at least one first sound source data in the sound field space, it is also necessary to obtain the first sound source data first. In a possible implementation, the electronic device may obtain the first sound source data through the following embodiment 1, embodiment 2, or embodiment 3.
[0051] Among them, in Example 1, the electronic device can obtain the at least one first sound source data in this way: obtain at least one third sound source data created; and process the third sound source data according to the resampling parameters and volume processing parameters to obtain the first sound source data.
[0052] Furthermore, in some embodiments, when the sound scene of the third sound source data is a non-voice call scene, the electronic device processes the third sound source data according to resampling parameters and volume processing parameters to obtain the first sound source data.
[0053] Exemplarily, in some embodiments, the electronic device may obtain the first sound source data by: resampling and volume processing the third sound source data according to the resampling parameters and the volume processing parameters to obtain fourth sound source data; and superimposing the fourth sound source data with an empty cache to obtain the first sound source data.
[0054] In the embodiment of the present application, the sound scene of the sound source data can be understood as the sound type of the sound source data or the application type of the application that created the sound source data. For example, the sound scene may be music, video, audiobook, ringing, game or voice call, etc. Among them, music, video, audiobook, ringing, and game can be classified as non-voice call scenes.
[0055] It can be understood that the reason for superimposing the resampled and audio-processed third sound source data (ie, the fourth sound source data) with the empty buffer is to achieve minimal modification of the original 3D mixing framework. Figure 2 As shown, the native 3D mixing framework processes the sound source data in three steps: first, resampling, then volume processing; finally, linearly superimposing the volume-processed sound source data to output dual-channel data.
[0056] In the embodiment of the present application, Figure 3 As shown, after the third sound source data (i.e., sound source data 1 or sound source data 2 output by the decoder) is resampled and volume processed, the obtained fourth sound source data is superimposed on the empty cache, so that what is obtained in the empty cache is the sound source data after pseudo-mixing (i.e., the first sound source data) after resampled by the native 3D mixing framework and volume processing, and then the first sound source data is track separated, the sound and image position is determined, and the second sound source data is determined. Finally, the sound and image rendering and 3D mixing output of the first sound source data are completed with as little modification as possible and compatibility with the native framework.
[0057] In an embodiment of the present application, the at least one third sound source data may be one sound source data or multiple sound source data. One sound source data or multiple sound source data may be understood as one track data or multiple track data. When the at least one third sound source data is multiple track data, the multiple sound source data may be created by one application or multiple applications.
[0058] It can be understood that the reason why the electronic device needs to process the third sound source data according to the resampling parameters and the volume processing parameters is that the sampling rates of different third sound source data may be different, but the sampling rate of the playback thread used to play the sound source data is fixed, so it is necessary to first process the at least one third sound source data belonging to the same type of playback thread into sound source data with a unified format, that is, the data format supported by the corresponding playback thread. Therefore, in some embodiments, different third sound source data use the same resampling parameter, and the size of the resampling parameter is the same as the sampling rate supported by the corresponding playback thread.
[0059] In practical applications, Embodiment 1 may be applicable to the MixerThread playback thread of the original 3D mixing framework. Embodiment 1 may be a modification process of the MixerThread playback thread of the original 3D mixing framework.
[0060] In Example 2, before determining the sound and image position of the first sound source data in the sound field space, the method also includes: acquiring at least one created third sound source data; and when the sound scene of the third sound source data is a voice call scene, resampling and converting the data type of the third sound source data, or converting the data type of the third sound source data to obtain the first sound source data.
[0061] Further, in some embodiments, when the first sampling rate of the third sound source data is less than or greater than the second sampling rate, and the first data type of the third sound source data is not the second data type, the electronic device resamples and converts the data type of the third sound source data according to the second sampling rate and the second data type to obtain the first sound source data. Accordingly, in step 103, the second sound source data of the at least one first sound source data is mixed and played, including: according to the first data type and the first sampling rate, performing reverse data type conversion and reverse resampling on the second sound source data to obtain fifth sound source data; and playing the fifth sound source data.
[0062] In other embodiments, when the first sampling rate of the third sound source data is equal to the second sampling rate, and the first data type of the third sound source data is not the second data type, the third sound source data is converted into the first sound source data according to the second data type to obtain the first sound source data. Accordingly, in step 103, the second sound source data of the at least one first sound source data is mixed and then played, including: according to the first data type, the second sound source data is reversely converted into the fifth sound source data to obtain the fifth sound source data; and the fifth sound source data is played.
[0063] Further, in some embodiments, the playing of the fifth sound source data includes: playing the fifth sound source data through a playback thread adapted to the voice call scenario.
[0064] In practical applications, Embodiment 2 may be applicable to the DirectOutputThread playback thread of the native 3D mixing framework. Embodiment 2 may be a modification process of the DirectOutputThread playback thread of the native 3D mixing framework that may be used in voice call-type sound scenes.
[0065] Exemplarily, the second sampling rate and the second data type are working parameters supported by the DirectOutputThread playback thread. For example, the second sampling rate supported by the playback thread is 48000 Hz, and the supported second data type is a floating point data type.
[0066] It is understandable that since DirectOutputThread only processes one active track per playback loop, each 3D mix only needs to process this one track, which is simpler than the logic of MixerThread that needs to collect, classify and copy multiple tracks. However, since the rendering algorithm only supports receiving data formats with a sampling rate of 48000Hz and floating-point data types, resampling and format conversion logic are added before actual processing. After the data is prepared, it is sent to the rendering algorithm, and after rendering, it is reversed and the format is reversed to the format of the original track. Finally, the data is replaced and written to the underlying device to complete a cycle of the direct type playback thread.
[0067] In Embodiment 3, the at least one first sound source data is sound source data that does not need to be spatialized based on the corresponding destination cache address. After acquiring the at least one third sound source data created, the electronic device can identify whether the sound source data is sound source data that does not need to be spatialized based on the destination cache address of the third sound source data; if so, perform the operation described in Embodiment 1 on the third sound source data to obtain the first sound source data.
[0068] Exemplarily, in some embodiments, the sound source data that does not require spatialization includes non-multi-channel sound source data.
[0069] In practical applications, Embodiment 3 can be applied to the SpatializerThread playback thread of the original 3D mixing framework, and the above Embodiment 3 can be a modification of the SpatializerThread playback thread of the original 3D mixing framework. The SpatializerThread playback thread is a subclass of MixerThread, which is used to play multi-channel spatial sound source data such as 7.1.4. This thread is designed to allow the playback of sound source data that does not require spatialization, such as non-multi-channel sound source data.
[0070] In an embodiment of the present application, this type of playback thread is modified, that is, the sound source data processing method provided by the embodiment of the present application is also executed for the sound source data that is considered not to need to be spatialized in the original 3D mixing framework, so that it also has the corresponding sound and image position in the playback listening experience.
[0071] After the first sound source data is obtained based on the above-mentioned embodiment 1, embodiment 2 or embodiment 3, the corresponding sound image position can be set / assigned to the first sound source data.
[0072] That is, in step 101, the sound image position of at least one first sound source data in the sound field space is determined.
[0073] In the embodiment of the present application, the electronic device can determine the sound image position of the first sound source data in various ways. The electronic device can determine the sound image position of the first sound source data based on a control signal input from a UI interface, or can search for the sound image position of the first sound source data based on a first mapping table.
[0074] Specifically, in some embodiments, the sound and image position of the first sound source data can be determined as follows: receiving a control signal based on UI interface input, the control signal including an application identifier and a corresponding sound and image position; and determining the sound and image position of the first sound source data of the application of the corresponding application identifier based on the control signal.
[0075] In one possible implementation, the UI interface can display a visualized sound field space, and the user can drag an application to a desired position in the sound field space in the UI interface, thereby generating a control signal including a corresponding sound and image position and an application identifier of the application based on the desired position.
[0076] Of course, in the embodiment of the present application, there is no limitation on the method of inputting control signals based on the UI interface. The electronic device supports users to manually drag any application to any position in the visualized sound field space, and to specify the sound and image position of a certain application by inputting text or voice.
[0077] In other embodiments, for the case where the at least one first sound source data is one track data or multiple tracks data, the method for determining the sound image position may be the same or different. Figure 4 The embodiment shown in the figure determines the sound image position of each sound source data; in some embodiments, the electronic device can implement step 101 as follows: when the at least one first sound source data is one, the sound image position of the first sound source data belongs to the center position. When the at least one first sound source data is multiple, the sound image position of the first sound source data is adopted. Figure 4 The embodiment shown determines the sound image position of each sound source data, such as Figure 4 As shown, in some embodiments, the electronic device can determine the sound image position of the first sound source data through the following steps 401 and 402:
[0078] Step 401: determine the sound scene to which the first sound source data belongs.
[0079] In some embodiments, the electronic device may implement step 401 as follows: obtaining the application identifier and stream type of the first sound source data; and determining the sound scene to which the first sound source data belongs according to the application identifier and stream type of the first sound source data.
[0080] It can be understood that when an application creates the original data corresponding to the first sound source data (i.e., the third sound source data), it is also necessary to specify the application identifier (such as the application package name) and stream type (such as media stream, ringtone, notification tone, voice call or alarm, etc.) of the application to which the sound source data belongs.
[0081] After obtaining the application identifier and stream type of the first sound source data, the electronic device may further, in some embodiments, determine the sound scene to which the first sound source data belongs based on these parameters by: determining the corresponding default sound scene according to the application identifier of the first sound source data; and when the stream type of the first sound source data does not belong to the first stream type, the sound scene to which the first sound source data belongs is the default sound scene; when the stream type of the first sound source data belongs to the first stream type, the sound scene to which the first sound source data belongs is the stream type of the first sound source data.
[0082] In other embodiments, when the stream type of the first sound source data does not belong to the first stream type and the sampling rate of the first sound source data is not equal to the third sampling rate, the sound scene to which the first sound source data belongs is the default sound scene; wherein the third sampling rate is the sampling rate of the sound source data of human voice; when the stream type of the first sound source data does not belong to the first stream type and the sampling rate of the first sound source data is equal to the third sampling rate, the sound scene to which the first sound source data belongs is a voice call; wherein the third sampling rate is the sampling rate of the sound source data of human voice.
[0083] It should be noted that in the embodiment of the present application, there is no limitation on the first flow type, which may be any predefined one or more flow types.
[0084] Exemplarily, in some embodiments, the first stream type includes at least one of the following: a ringing tone, a notification tone, a voice call, and an alarm.
[0085] Regarding determining the corresponding default sound scene based on the application identifier of the first sound source data, in one possible implementation, the electronic device can load a pre-configured second mapping table, which records the default sound scenes corresponding to at least one application identifier; then, based on the application identifier of the first sound source data, the corresponding default sound scene is found from the second mapping table.
[0086] It can be understood that the so-called default sound scene refers to a predefined sound scene. The reason why the electronic device determines the sound scene to which the first sound source data ultimately belongs based on the above embodiment is that: for some applications (such as instant messaging applications), it actually supports both media streaming (i.e., video streaming) and voice calls and ringtones. However, in the predefined second mapping table, the configured default sound scene is video. Therefore, determining the sound scene to which the first sound source data ultimately belongs based on the above embodiment is beneficial to obtaining a more accurate sound scene for the first sound source data, thereby arranging a more appropriate sound and image position for it.
[0087] In some embodiments, after determining the sound scene to which the first sound source data belongs, the method further includes: generating first attribute information of the first sound source data; wherein the first attribute information is used to record the scene identifier of the sound scene to which the first sound source data belongs, so that when executing step 402, the sound image position of the sound source data is determined based on the first attribute information of the first sound source data.
[0088] Step 402: Determine the sound image position of the first sound source data according to the sound scene to which the first sound source data belongs.
[0089] In some embodiments, when there are multiple pieces of at least one first sound source data, the electronic device determines the sound and image position of the first sound source data according to the sound scene to which the first sound source data belongs.
[0090] Furthermore, in some embodiments, when the sound and image state of the first sound source data is in a non-transition stage and there are multiple first sound source data, the sound and image position of the first sound source data is determined according to the sound scene to which the first sound source data belongs; wherein the non-transition stage refers to the sound and image position of the first sound source data in different time periods of the same audio track remains unchanged.
[0091] Exemplarily, in some embodiments, the electronic device can determine the sound and image position of the first sound source data from the first mapping table according to the sound scene to which the first sound source data belongs; for example, based on the first attribute information of the first sound source data, the sound and image position corresponding to the scene identifier carried by the first attribute information is searched from the first mapping table. The first mapping table records the sound and image positions corresponding to at least one sound scene.
[0092] Furthermore, in the first mapping table, different sound scenes correspond to different sound image positions.
[0093] In some embodiments, the sound and image positions corresponding to at least one sound scene in the first mapping table are predefined. In a possible implementation, the electronic device can save the predefined first mapping table locally in the form of an XML file, and the electronic device can load the preconfigured XML file to obtain the mapping information between the scene identifier and the sound and image position of at least one sound scene.
[0094] And / or, it is configured according to the sound image position of the corresponding sound scene used by the second sound source data previously determined, that is, it is configured according to the sound image position used in the sound scene history.
[0095] Of course, in a possible implementation, the electronic device also supports users or developers to configure the first mapping table by themselves. That is, in some embodiments, the method further includes: receiving configuration information, the configuration information including specifying the sound and image position of at least one sound scene; and configuring the sound and image position of the corresponding sound scene in the first mapping table according to the configuration information; in this way, the user or developer can customize the sound and image position of any sound scene through the configuration information.
[0096] In the embodiment of the present application, there is no limitation on the method for the electronic device to receive configuration information. It can be based on the UI interface of the electronic device to receive configuration information input by the user, or it can be based on the communication module of the electronic device to receive configuration information sent remotely by the developer.
[0097] In another possible implementation, the electronic device also supports users or developers to change the sound and image position corresponding to the sound scene in the first mapping table at any time and anywhere. That is, in some embodiments, the method further includes: receiving an update signal, the update signal is used to indicate the update of the sound and image position of at least one specific sound scene, the update signal includes the identification information of the at least one specific sound scene and the new sound and image position; and updating the sound and image position of the at least one specific sound scene in the first mapping table according to the update signal. That is, the sound and image position of the specific scene type in the first mapping table is updated to the new sound and image position corresponding to the update signal.
[0098] In the embodiment of the present application, there is no limitation on the method for the electronic device to receive the update signal. It can be based on the UI interface of the electronic device to receive the update signal input by the user, or it can be based on the communication module of the electronic device to receive the update signal sent remotely by the developer.
[0099] The above-mentioned embodiment of determining the sound image position can be used as a method for determining the sound image position when the sound image state of the first sound source data is in a non-transition stage.
[0100] When the sound and image state of the first sound source data is in a transition stage, in some embodiments, the electronic device may implement step 101 as follows: when the sound and image state of the first sound source data is in a transition stage, the sound and image position of the first sound source data is determined according to the sound and image position and the sound and image position update step of the first sound source data played in the previous time period (such as the previous frame); wherein the transition stage refers to the change in the sound and image position of the first sound source data in different time periods (such as different frames) of the same audio track.
[0101] It should be noted that, in the embodiment of the present application, the time period can be understood as a unit data amount of a cycle of reading, writing / processing the first sound source data, for example, the unit data amount is a frame of sound source data or other unit data amount.
[0102] Further, in some embodiments, when the sound and image state of the first sound source data is in a transitional stage, the sound and image position of the first sound source data played in the initial time period (such as the initial frame) is equal to the first spatial coordinate belonging to the scene position. In one possible implementation, the first spatial coordinate may be the sound and image position corresponding to the sound scene of the sound source data in the first mapping table. In another possible implementation, the first spatial coordinate may also be a predefined scene position. The first spatial coordinate is the same as or different from the sound and image position corresponding to the sound scene to which the first sound source data belongs.
[0103] In some embodiments, at the end of the transition phase of the first sound source data, the sound image position of the first sound source data played in the current time period (such as the current frame) is equal to the sound image position corresponding to the sound scene to which the first sound source data belongs.
[0104] In some embodiments, the method further includes: generating second attribute information of the first sound source data, the second attribute information being used to record whether the sound and image state of the first sound source data is a transition stage or a non-transition stage. That is, before determining the sound and image position of the first sound source data, the electronic device may first mark the attribute information of the first sound source data, such as marking the second attribute information of the first sound source data, so that the next module determines the sound and image state of the first sound source data at this time according to the second attribute information, and thus determines the sound and image position of the first sound source data using a corresponding method based on the sound and image state.
[0105] In some embodiments, the method further includes: when the transition phase of the first sound source data ends, updating the second attribute information of the first sound source data to indicate that the sound and image state is in a non-transition phase.
[0106] In some embodiments, the transition stage refers to the gradual transition of the sound and image position from the first spatial coordinate belonging to the scene position to the center position (i.e., fading in to the center position) according to the sound and image position update step when the first sound source data is played at different time periods (e.g., different frames), and / or the gradual transition from the center position to the sound and image position corresponding to the sound scene to which the first sound source data belongs (i.e., fading out from the center position to the final position) according to the sound and image position update step.
[0107] That is to say, in the embodiment of the present application, there is no limitation on the transition method of the first sound source data. It can be fade-in from the first spatial coordinate to the center position, or fade-out from the center position to the sound image position corresponding to the sound scene, or fade-in from the first spatial coordinate to the center position, and then fade-out from the center position to the sound image position corresponding to the sound scene.
[0108] The last transition method mentioned above achieves a fade-in and fade-out sound effect in terms of auditory perception, that is, the sound first goes from far to near, then from near to far, and finally returns to the sound and image position corresponding to the sound scene to which the first sound source data belongs.
[0109] Of course, the electronic device may not distinguish between the transition stage and the non-transition stage, that is, instead of using the above-mentioned method for determining the sound image position in the transition stage or the non-transition stage, the electronic device may use the following method to determine the sound image position of the first sound source data:
[0110] That is, in some embodiments, the electronic device may also implement step 402 in the following manner: determining the first sound source data occupying the center position according to the sound scene of the at least one first sound source data.
[0111] Furthermore, in a possible implementation, the electronic device may determine the first sound source data with the highest priority from the at least one first sound source data based on the priority of each sound scene defined in advance; and set the sound image position of the first sound source data with the highest priority as the center position. For example, the electronic device may determine the spatial coordinates (i.e., the sound image position) of the center position in the sound field space based on the number of channels of the first sound source data with the highest priority.
[0112] Furthermore, in some embodiments, for multiple first sound source data of the same priority, the first sound source data that occupies the center position may be determined in the order of creation time from near to far from the current moment. That is, for multiple first sound source data of the same priority, the first sound source data whose creation time is closest to the current moment occupies the center position first.
[0113] In the embodiment of the present application, there is no limitation on the way the first sound source data occupies the center position, and the sound image position can be directly at the center position, or it can be gradually transitioned from the first space coordinate to the center position according to the sound image position update step.
[0114] Regarding the method for determining the sound image position of the first sound source data belonging to the center position, in some embodiments, the electronic device may determine the sound image position of the first sound source data based on the number of channels of the first sound source data. The sound image position of the first sound source data belonging to the scene position may be the sound image position corresponding to the corresponding sound scene recorded in the first mapping table.
[0115] For example, Figure 5 A schematic diagram of the relationship between the center position and the scene position provided in the embodiment of the present application, such as Figure 5 As shown, the center position may be the middle position 502 in front of the listening device 501 in the sound field space 50, and the rest of the elliptical areas are scene positions. For another example, the center position is a position 1 meter away from the listening device 501 in a direction of 45 degrees to the left and right.
[0116] In some embodiments, the method also includes: after determining the sound and image position of the first sound source data in the sound field space, generating third attribute information, wherein the third attribute information is used to record the sound and image position of the first sound source data; so that the module executing step 102 can obtain the sound and image position of the first sound source data according to the third attribute information.
[0117] In some embodiments, the at least one first sound source data satisfies a validity condition; wherein the validity condition includes: the application identifier of the application to which the at least one first sound source data belongs is in the application whitelist; and / or the sampling rate of the at least one first sound source data or the corresponding third sound source data is at least one supported sampling rate.
[0118] In an embodiment of the present application, the application whitelist may be pre-configured, and the electronic device supports users to customize the application identifiers in the application whitelist and modify the application identifiers in the application whitelist at any time.
[0119] In some embodiments, before determining the sound and image position of at least one first sound source data in the sound field space, the method also includes: based on determining that the at least one first sound source data satisfies the validity condition, generating fifth attribute information, the fifth attribute information is equal to a first value; wherein the first value is used to indicate that the at least one first sound source data is played using the sound source processing method.
[0120] In some embodiments, the electronic device determines the sound image position of at least one first sound source data in the sound field space when the fifth attribute information of the at least one first sound source data is equal to the first value.
[0121] In step 102, second sound source data is determined according to the sound image position of the first sound source data and the first sound source data.
[0122] In some embodiments, the electronic device may render the first sound source data according to the sound and image position of the first sound source data (such as according to the third attribute information) to obtain the second sound source data.
[0123] In one possible implementation, the electronic device may input the first sound source data and the third attribute information for recording the corresponding sound and image position into the spatial audio rendering algorithm, so that the spatial audio rendering algorithm performs spatial processing on the corresponding first sound source data according to the sound and image position, and outputs the second sound source data with the sound and image position. For example, the spatial audio rendering algorithm selects the corresponding head-related transfer function (HRTF) according to the input sound and image position, and renders the corresponding first sound source data through the HRTF function. Among them, the parameters of the HRTF function corresponding to different sound and image positions are different.
[0124] In step 103, the second sound source data of the at least one first sound source data is mixed and then played.
[0125] Figure 6 A schematic diagram of a method for implementing step 103 provided in an embodiment of the present application is shown in FIG. Figure 6 As shown, in some embodiments, the electronic device can implement step 103 through the following steps 601 and 602:
[0126] Step 601 : superimpose the second sound source data of the at least one first sound source data to obtain multi-channel data to be played.
[0127] It should be noted that the multi-channel data to be played outputted in step 601 may be dual-channel data, or may be three-channel audio data, or more-channel audio data.
[0128] The superposition processing described in step 601 can be understood as an implementation of the mixing described in step 103. For example, in one possible implementation, the electronic device can superimpose (such as linear superposition or nonlinear superposition) the left channel data in the second sound source data of the at least one first sound source data to obtain the left channel data, and superimpose (such as linear superposition or nonlinear superposition) the right channel data in the second sound source data of the at least one first sound source data to obtain the right channel data.
[0129] Step 602: Play the multi-channel data to be played through a playing thread.
[0130] In some embodiments, the at least one first audio source data belongs to the same playback thread type.
[0131] In some embodiments, before mixing and playing the second sound source data of the at least one first sound source data, the method further includes: determining the playback thread to which the at least one first sound source data belongs; and generating fourth attribute information based on the playback thread to which the first sound source data belongs, the fourth attribute information being used to record the playback thread to which the first sound source data belongs; in this way, the corresponding playback thread can be selected based on the fourth attribute information to play the mixed multi-channel data.
[0132] It can be understood that the processing framework / processing logic for the original sound source data belonging to different playback threads is different.
[0133] The present application embodiment provides a method for processing sound source data, which is applicable to a MixerThread playback thread. Figure 7 Schematic diagram of the implementation process of the sound source data processing method provided in the embodiment of the present application Figure 2 ,like Figure 7 As shown, the method includes the following steps 701 to 707:
[0134] Step 701, obtaining a plurality of created third sound source data;
[0135] Step 702, when the playback thread of the plurality of third sound source data is the first playback thread (such as the MixerThread playback thread), resampling and volume processing are performed on the third sound source data according to the resampling parameter and the volume processing parameter to obtain fourth sound source data;
[0136] Step 703, superimposing the fourth sound source data with the empty buffer to obtain the first sound source data;
[0137] Step 704, marking the first sound source data with attribute information; wherein the marked attribute information includes first attribute information, second attribute information, third attribute information, fourth attribute information and fifth attribute information;
[0138] Step 705, rendering the first sound source data according to the attribute information of the first sound source data to obtain second sound source data;
[0139] Step 706, superimposing the second sound source data corresponding to the plurality of third sound source data to obtain dual-channel data;
[0140] Step 707: output the two-channel data.
[0141] The present application embodiment provides a method for processing sound source data, which is applicable to the DirectOutputThread playback thread. Figure 8 Schematic diagram of the implementation process of the sound source data processing method provided in the embodiment of the present application Figure 3 ,like Figure 8 As shown, the method includes the following steps 801 to 806:
[0142] Step 801, obtaining at least one third sound source data created;
[0143] Step 802, when the playback thread of the third sound source data is the second playback thread (such as DirectOutputThread playback thread), resample and convert the data type of the third sound source data to obtain the first sound source data;
[0144] Step 803, determining the sound image position of the first sound source data of at least one third sound source data in the sound field space;
[0145] Step 804, determining second sound source data according to the sound image position of the first sound source data and the first sound source data;
[0146] Step 805, performing inverse data type conversion and inverse resampling on the second sound source data to obtain fifth sound source data;
[0147] Step 806, playing the fifth sound source data.
[0148] The present application embodiment provides a method for processing sound source data, which is applicable to a SpatializerThread playback thread. Fig. 9 Schematic diagram of the implementation process of the sound source data processing method provided in the embodiment of the present application Figure 4 ,like Fig. 9 As shown, the method includes the following steps 901 to 909:
[0149] Step 901, obtaining at least one third sound source data created;
[0150] Step 902, when the playback thread of the third sound source data is the third playback thread (such as the SpatializerThread playback thread), identify whether the third sound source data is sound source data that does not need to be spatialized according to the destination cache address of the third sound source data; if yes, execute step 904; otherwise, execute step 903;
[0151] Step 903, outputting the third sound source data;
[0152] Step 904, resampling and volume processing the third sound source data according to the resampling parameter and the volume processing parameter to obtain fourth sound source data;
[0153] Step 905, superimposing the fourth sound source data with the empty buffer to obtain the first sound source data;
[0154] Step 906, marking the first sound source data with attribute information; wherein the marked attribute information includes first attribute information, second attribute information, third attribute information, fourth attribute information and fifth attribute information;
[0155] Step 907, rendering the first sound source data according to the attribute information of the first sound source data to obtain second sound source data;
[0156] Step 908, superimposing the second sound source data corresponding to the plurality of third sound source data to obtain dual-channel data;
[0157] Step 909: output the two-channel data.
[0158] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.
[0159] The main problem with related audio systems is that they lack consideration for the multi-source sound scene experience and are unable to solve the problems of mutual interference and lack of spatial sense in multi-source sound scenes.
[0160] The embodiment of the present application aims to solve this problem. First, when the sound source applies to create a playback, the 3D mixing mark and sound scene affiliation are added to the playback through the incoming stream type and application category (such as application identifier) and other information; secondly, during mixing, the sound source data after resampling and volume processing is sent to the audio-visual control module, and the module arranges the sound source data according to the aforementioned mark and sound scene. The sound and image position is divided into a center position and a scene position, and a certain sound source data can be switched between the two by fading in and out; thirdly, these sound and image information are added to the calculation in the rendering process of upmixing into multi-channels, and then converted back to dual-channel data after achieving the spatial effect of sound separation; finally, the sound and image rendered data is copied back to the mixer main cache for linear superposition, and the output sound at this time is 3D mixing data with a sense of space.
[0161] In an embodiment of the present application, a 3D mixing method is provided for an Android system device when multiple types of sounds are played coexisting. The method enables the system to have the ability to classify, spatially arrange and mix different sounds by expanding and modifying the original mixer. Combined with the sound and image arrangement and rendering algorithm, the mutual interference of sounds in multi-source scenes can be greatly alleviated, and the sound separation is improved while having a good spatial audio effect.
[0162] In the embodiment of the present application, a configurable pre-mixing rendering process is added on the basis of the existing source mixer. In the actual data processing stage, the sound source data after resampling and volume processing is obtained by mixing the empty buffer with the single sound source data, and then the data is transmitted to the audio-visual arrangement and rendering module, and the spatial arrangement and effect processing are performed in combination with the sound source classification information. The data processed by the algorithm is copied back to the main buffer of the mixer, so that different spatial rendering effects can be added to each audio stream instead of one-dimensional linear superposition, and finally 3D mixing is completed. Among them, it is necessary to perform extended processing on each audio channel, including sampling rate compatibility of the direct channel, and the collection of the sound source that is not spatialized in the spatializer channel for 3D mixing.
[0163] The embodiment of the present application is applied to the audio system module of the mobile terminal. For the multi-source playback scenario, the method can enable the system mixer to have the pre-rendering capability to distinguish the sound and image information, and can ultimately achieve a spatial audio experience with good sound separation in combination with the spatial rendering algorithm.
[0164] Common audio playback scenarios include music, video, audiobooks, voice, games, navigation, etc. Users may have different usage requirements. For example, when a user wears headphones and uses network communication software to make a voice call while playing a game, the 3D mixer can render the game in the center and the voice in the left front scene, achieving a spatial effect of sound separation without interference.
[0165] The embodiments of this application mainly involve the following technical contents:
[0166] 1. Scene selection strategy and 3D mixing information marking:
[0167] In the Android system, an AudioTrack class is created when each sound source is played. Before 3D mixing, these tracks need to be classified into scenes and marked with various information so that different sound and image positions can be distributed during subsequent mixing.
[0168] Fig.10 A schematic diagram of the implementation flow of the scenario selection strategy provided in the embodiment of the present application is shown in FIG. Fig.10 As shown, the selection strategy includes the following steps 1001 to 1010:
[0169] Step 1001, AudioTrack creation (i.e. creation of third audio source data);
[0170] Step 1002, determine whether the application that creates AudioTrack is in the application whitelist; if so, execute step 1003; otherwise, bypass (byPass);
[0171] Step 1003, obtaining a default scene type (ie, a default sound scene) from the application whitelist;
[0172] Step 1004, determine whether the stream type (stream) is equal to Ring (i.e., ring tone); if so, execute step 1005; otherwise, execute step 1006;
[0173] Step 1005, determining that the final scene type is a ringing tone scene;
[0174] Step 1006, determine whether the stream type (stream) is equal to Notification (i.e., notification sound); if so, execute step 1007; otherwise, execute step 1008;
[0175] Step 1007, determining that the final scene type is a notification sound scene;
[0176] Step 1008, determining whether the stream type (stream) is equal to VoiceCall (i.e., voice call) or whether the sampleRate (i.e., sampling rate) is equal to 16K (i.e., an example of the third sampling rate); if so, executing step 1009; otherwise, executing step 1010;
[0177] Step 1009, determining that the final scenario type is a voice call scenario;
[0178] Step 1010 , determining that the final scene type is the default scene type determined in step 1003 .
[0179] Exemplarily, in an embodiment of the present application, the scene type (i.e., the sound scene to which the sound source data belongs) is divided into scenes such as music, video, audiobook, ringtone, game, and voice. When selecting a scene type, the application whitelist supported by the function is first read. The application whitelist defines the default scene type for each application, but an application does not necessarily play only the sound of one scene. For example, a music application can also play videos such as MVs. For this purpose, scene correction is performed based on the stream type and sampling rate information of the track. Here, since the minimum sampling rate requirement for media sound to fully restore the sound information is 40k, and the sampling rate requirement for human voice is much lower due to its low frequency of occurrence, when the stream type of the sound source is media stream and the sampling rate is 16kHz, which is common for human voice, the scene type selects the voice call scene; wherein, the voice call scene can also be understood as the voice connection scene.
[0180] Fig.11 A schematic diagram of the timing of the track 3D information marking provided in the embodiment of the present application is shown in FIG. Fig.11As shown in the figure, the application package name (InfoForPkgName) and scene member variables (used to record the scene type) are added to the original Track class. When the track is created, it will request the spatial information management service (SpecailizerPLService) in its constructor (track contructor) to obtain the application package name, default scene type and corresponding audio and video position, and then output the actual scene value of the track according to the selection strategy. For example Fig.11 As shown, in a possible implementation, trackcontructor calls the SpecailizerPLService module via addPkgNameForTrack, getMetaAudioTrackInfoForPkgName, and getMetaInfoByPkgName to obtain the application package name, default scene type, and corresponding sound and image position. The SpecailizerPLService module returns the application package name, default scene type, and corresponding sound and image position of the track to the AudioFlingerExtimpl module, which calls the MetaAudioUtils module via string2MetaAudioFormat to convert the string representing the application package name, default scene type, and corresponding sound and image position into data in the target format, and then returns it to the track module, so that the track module can mark the attribute information of the track data based on this information. The Android audio system has multiple playback threads such as MixerThread and DirectOutputThread to cope with different playback scenarios. All playback threads (Threads) work in a cyclic reading and writing data mode. The method will call the newly added function in the data preparation (preparetrack) stage of each thread cycle to set the cache size, playback thread flag, and other information for subsequent modules to use.
[0181] 2. Mixer framework transformation:
[0182] Fig.12 A schematic diagram of a 3D mixing framework provided in an embodiment of the present application, such as Fig.12 As shown, the Android system's native mixer data processing has three steps (such as Fig.12), first of all, resampling is performed. Because the sampling rates of different sound source data are likely to be different, but the sampling rate of the playback thread is fixed, the mixer needs to resample the data to a unified format first, and then perform volume processing and superposition mixing to complete the output. Due to the addition of the spatialization calculation and rendering module, the data needs to be classified and collected inside the mixer and sent to the newly added module, so the mixer needs to be modified. That is, in an embodiment of the present application, after the sound source data output by application 1 and application 2 are resampled and volume processed respectively, they are sent to the spatial processing module for track separation, spatial information marking and spatial rendering, and finally 3D mixed data is output.
[0183] A new spatial track data class is defined in the spatial processing module. This class mainly has the following attributes: the third attribute information is used to indicate whether the track is in the center or scene position; the sound image state (i.e., the second attribute information) indicates whether the track is in the process of changing the sound image position; the spatial position (i.e., the third attribute information) indicates the spatial coordinates of the track; the sound object array defines the basic unit of data spatial processing, and each channel of the track is considered as a sound object; the rendering state (i.e., the fifth attribute information) indicates whether the track's current loop data needs to be rendered by a spatial algorithm; the mixing type (i.e., the fourth attribute information) indicates the type of playback thread to which the track belongs. The spatial track corresponds one-to-one to the AudioTrack. The original track needs to be collected and filtered, and the attribute information and audio data are accurately given to the spatial track.
[0184] Fig.13 The following is a schematic diagram of the AudioMixer modification process provided in the embodiment of the present application, as shown in FIG. Fig.13 As shown in the figure, normal-channel indicates that the original sound source data created by the AudioTrack module of the application is dual-channel data. When the sound starts playing, the audio system playback thread loops to process the data. Each loop (threadLoop) includes three processes: data preparation (prepareTrack), mixing processing (threadLoop_mix) and data writing (threadLoop_write), corresponding to Fig.13 The right half of the audio system; Threadloop_mix is the AudioMixer (i.e., mixer) in the left half. As mentioned above, the audio system selects different processing functions (hook functions) to actually perform the mixing function according to the number and format of the audio sources. Fig.13The process module includes resampling of the sound source data, volume processing, and superposition processing with an empty buffer. The process here applies for a pointer array for each hook function. The array is applied before traversing the track during the mixing process to collect the data of each hook, so that the track will be divided into different sound source groups (trackGroup) by type for difference processing. When a single track is mixed in a loop, the original logic is replaced. In the process module, the empty buffer pointed to by the applied pointer array is mixed with the track. In this way, the empty buffer is the track data after the original logic is resampled and the volume processing is completed. Then the pointer array is passed to the spatial processing module (that is, it is finally passed to the spatial sound effect algorithm library), so that the classified track information and data corresponding to the sound source are obtained. Spatial processing mainly performs spatial information marking and effect algorithm rendering. The processed data will be directly copied (Return data) to the main buffer output by the mixer (AudioMixer) to replace the original data and write to the underlying device (audioHAL play) through the Threadloop_write module, and finally complete a threadLoop cycle.
[0185] Fig.14 The DirectOutput transformation processing diagram provided in the embodiment of the present application is as follows: Fig.14As shown in the figure, in addition to the most commonly used MixerThread playback thread, the DirectOutputThread playback thread that may be used for call-type sounds (i.e., voice call scenarios) also needs to be modified. Unlike the resident MixerThread, this thread is dynamically created when used, and a new direct type spatial algorithm instance is initialized when it is created. Since DirectOutputThread only processes one active track per playback loop, each 3D mixing only needs to process this track, which is simpler than the logic of collecting, classifying and copying multiple tracks required by MixerThread. However, since the spatial rendering algorithm can only receive data in a sampling rate of 48000Hz and a floating point data type, resampling (resample48k) and format conversion (convertFormat) logic are added before actual processing. After the data is ready, it is sent to the spatial sound algorithm library for rendering through the process module. After rendering, it is resampled (revertResample) and the format is reversed (revertFormat) to the format of the original track. Finally, the data is replaced and written to the underlying device (audioHAL play) through the Threadloop_write module, thus completing a cycle of the direct type playback thread. It should be noted that, Fig.14 The DirectOutputThreadExt module shown in is the processMetaAudio module in the right half, which includes resample48k, convertFormat, process, spatial sound algorithm library, revertResample and revertFormat.
[0186] In addition, there is the SpatializerThread playback thread, which is a subclass of MixerThread and is a playback thread added to the original spatial audio framework. It is used to play multi-channel spatial audio data such as 7.1.4. This thread is designed to allow the playback of audio sources that do not require spatialization, such as non-multi-channel audio sources. When such tracks that do not actually require spatialization are identified by detecting the track destination cache address during the data preparation phase of the thread loop, they will be collected and reused in the MixerThread manner, and these special tracks will be mixed in 3D.
[0187] 3. Spatial information marking and processing:
[0188] Fig.15 The DirectOutput transformation processing diagram provided in the embodiment of the present application is as follows: Fig.15As shown in the figure, in the spatial processing module of 3D mixing, the concepts of center position and scene position are added. There is only one center position, and the non-center sound and image positions are all scene positions. When there is only one sound source, it will occupy the center position regardless of its scene. When there are multiple sound sources, the priority is defined for each scene from high to low: calls, ringing, alarm clock, games, videos, audiobooks, and music. Tracks of high-priority scenes will occupy the center position, and tracks of scenes with the same priority will occupy the center position in the order of starting broadcast, with later broadcasts taking priority. If a track occupies the center position, it will fade in from the original scene position to the center position, and the track of the original center position will fade out to its scene position.
[0189] After the track information and data group are sent to the spatial processing module, the validity check will be performed first, and then the sound and image position of each sound source will be set according to the priority strategy, and the track at the center position will be selected. Then all sound objects are traversed and the spatial position of the sound object is updated. Here, the spatial position is a three-dimensional spatial coordinate centered on the human head. Its input has three parts. One is to define the default position for each scene, the second is to allow users to customize the spatial position of each scene, and the third is that if the track is in a transition stage such as fade in and fade out, the spatial position will be gradually adjusted according to the step size. After the spatial position is assigned, the rendering flag will continue to be updated. When the track is not in the center position, it is necessary to combine the position coordinates for algorithm rendering. No rendering is performed in the center position, and the effect is equivalent to the original mixer. When all tracks are traversed, their spatial positions and various flags have been accurately updated. The audio data is further sent to the algorithm rendering module for digital processing. The data retrieved is data with a sense of spatial separation. Finally, the data of each track is superimposed to complete the 3D mixing effect.
[0190] It can be understood that based on the above technical solution, the expansion and transformation of the Android system mixer framework can be completed, so that the mixer can separate the sound source data into groups between the volume processing and superposition mixing stages, laying the foundation for the spatial processing of the data. A new spatial marker and rendering module is added, which can calculate the spatial position of each sound source in combination with the default position, custom position and transition state step. Furthermore, the effect rendering is completed by combining the spatial information and the sound source data, and the final 3D mixed data is output. The mixed data has a spatial separation effect, which can bring a good immersive audio experience under multiple sound sources.
[0191] It should be noted that the above solution is implemented based on the Android system, and its linear mixer is expanded to a 3D mixer. At present, the mainstream mobile open source system that can be deeply customized is mainly Android, but its core idea can be extended to other systems. For multi-source playback scenarios, the sound image position and spatial position are added to each sound source during mixing, and the spatialized 3D mixing data is obtained by effect rendering.
[0192] It should be noted that although the steps of the method in the present application are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all the steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps, etc.; or, steps in different embodiments may be combined into a new technical solution.
[0193] Based on the foregoing embodiments, an embodiment of the present application provides a sound source data processing device, which includes the modules included and the units included in the modules, which can be implemented by a processor; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be an AI acceleration engine (such as NPU, etc.), GPU, central processing unit (CPU), microprocessor (MPU), digital signal processor (DSP) or field programmable gate array (FPGA), etc.
[0194] Fig.16 A schematic diagram of the structure of the sound source data processing device provided in the embodiment of the present application, such as Fig.16 As shown, the sound source data processing device 160 includes:
[0195] A first determination module 1601 is configured to determine a sound image position of at least one first sound source data in a sound field space;
[0196] A second determination module 1602 is configured to determine second sound source data according to the sound image position of the first sound source data and the first sound source data;
[0197] The mixing and playing module 1603 is configured to mix the second sound source data of the at least one first sound source data and then play it.
[0198] In some embodiments, the first determination module 1601 is further configured to: before determining the sound and image position of the first sound source data in the sound field space, obtain at least one third sound source data created; and process the third sound source data according to the resampling parameters and volume processing parameters to obtain the first sound source data.
[0199] In some embodiments, the first determination module 1601 is configured to: when the sound scene of the third sound source data is a non-voice call scene, process the third sound source data according to resampling parameters and volume processing parameters to obtain the first sound source data.
[0200] In some embodiments, the processing of the third sound source data according to the resampling parameters and the volume processing parameters to obtain the first sound source data includes: resampling and volume processing the third sound source data according to the resampling parameters and the volume processing parameters to obtain fourth sound source data; and superimposing the fourth sound source data with an empty cache to obtain the first sound source data.
[0201] In some embodiments, different third audio source data use the same resampling parameters.
[0202] In some embodiments, the first determination module 1601 is configured to: before determining the sound and image position of the first sound source data in the sound field space, obtain at least one third sound source data created; and when the sound scene of the third sound source data is a voice call scene, resample and convert the data type of the third sound source data, or, convert the data type of the third sound source data, to obtain the first sound source data.
[0203] In some embodiments, the third sound source data is resampled and the data type is converted to obtain the first sound source data, including: when the first sampling rate of the third sound source data is less than or greater than the second sampling rate, and the first data type of the third sound source data is not the second data type, the third sound source data is resampled and the data type is converted according to the second sampling rate and the second data type to obtain the first sound source data.
[0204] In some embodiments, the data type of the third sound source data is converted to obtain the first sound source data, including: when the first sampling rate of the third sound source data is equal to the second sampling rate, and the first data type of the third sound source data is not the second data type, according to the second data type, the data type of the third sound source data is converted to obtain the first sound source data.
[0205] In some embodiments, the mixing and playing of the second sound source data of the at least one first sound source data includes: performing inverse data type conversion and inverse resampling on the second sound source data according to the first data type and the first sampling rate to obtain fifth sound source data; and playing the fifth sound source data.
[0206] In some embodiments, mixing and then playing the second sound source data of the at least one first sound source data includes: performing a reverse data type conversion on the second sound source data according to the first data type to obtain fifth sound source data; and playing the fifth sound source data.
[0207] In some embodiments, the playing of the fifth sound source data includes: playing the fifth sound source data through a playback thread adapted to the voice call scenario.
[0208] In some embodiments, the at least one third sound source data includes sound source data that does not require spatialization and is identified based on a corresponding destination cache address.
[0209] In some embodiments, the sound source data that does not require spatialization includes non-multi-channel sound source data.
[0210] In some embodiments, determining the sound image position of the first sound source data in the sound field space includes: determining the sound image position of the first sound source data according to a sound scene to which the first sound source data belongs.
[0211] In some embodiments, determining the sound and image position of the first sound source data according to the sound scene to which the first sound source data belongs includes: determining the sound and image position of the first sound source data from a first mapping table according to the sound scene to which the first sound source data belongs; wherein the first mapping table records the sound and image positions corresponding to at least one sound scene.
[0212] In some embodiments, in the first mapping table, different sound scenes correspond to different sound image positions.
[0213] In some embodiments, the sound image position corresponding to at least one sound scene in the first mapping table is predefined and / or is configured according to the sound image position of the corresponding sound scene used by the second sound source data previously determined.
[0214] In some embodiments, the sound source data processing device 160 also includes a receiving module and a configuration module; wherein the receiving module is configured to receive configuration information, wherein the configuration information includes specifying the sound and image position of at least one sound scene; and the configuration module is configured to configure the sound and image position of the corresponding sound scene in the first mapping table according to the configuration information.
[0215] In some embodiments, the sound source data processing device 160 also includes a receiving module and an updating module; wherein the receiving module is configured to receive an update signal, the update signal is used to indicate an update of the sound and image position of at least one specific sound scene, and the update signal includes identification information and a new sound and image position of the at least one specific sound scene; the updating module is configured to update the sound and image position of the at least one specific sound scene in the first mapping table according to the update signal.
[0216] In some embodiments, determining the sound and image position of the first sound source data in the sound field space includes: receiving a control signal based on UI interface input, the control signal including an application identifier and a corresponding sound and image position; and determining the sound and image position of the first sound source data of the application of the corresponding application identifier based on the control signal.
[0217] In some embodiments, the first determination module 1601 is further configured to obtain the application identifier and stream type of the first sound source data before determining the sound and image position of the first sound source data according to the sound scene to which the first sound source data belongs; and determine the sound scene to which the first sound source data belongs according to the application identifier and stream type of the first sound source data.
[0218] In some embodiments, determining the sound scene to which the first sound source data belongs based on the application identifier and stream type of the first sound source data includes: determining a corresponding default sound scene based on the application identifier of the first sound source data; and in a case where the stream type of the first sound source data does not belong to the first stream type, the sound scene to which the first sound source data belongs is the default sound scene.
[0219] In some embodiments, when the stream type of the first sound source data does not belong to the first stream type and the sampling rate of the first sound source data is not equal to the third sampling rate, the sound scene to which the first sound source data belongs is the default sound scene; wherein the third sampling rate is the sampling rate of the sound source data of human voice.
[0220] In some embodiments, when the stream type of the first sound source data does not belong to the first stream type and the sampling rate of the first sound source data is equal to a third sampling rate, the sound scene to which the first sound source data belongs is a voice call; wherein the third sampling rate is the sampling rate of the sound source data of human voice.
[0221] In some embodiments, when the stream type of the first sound source data belongs to the first stream type, the sound scene to which the first sound source data belongs is the stream type of the first sound source data.
[0222] In some embodiments, the first stream type includes at least one of the following: a ringing tone, a notification tone, a voice call, and an alarm.
[0223] In some embodiments, the sound source data processing device 160 also includes a generation module, which is configured to generate first attribute information of the first sound source data after determining the sound scene to which the first sound source data belongs; wherein the first attribute information is used to record the scene identifier of the sound scene to which the first sound source data belongs.
[0224] Based on this, in some embodiments, determining the sound and image position of the first sound source data from a first mapping table according to the sound scene to which the first sound source data belongs includes: searching the sound and image position corresponding to the scene identifier from the first mapping table according to the first attribute information of the first sound source data.
[0225] In some embodiments, the first determination module 1601 is configured to: when there are multiple pieces of at least one first sound source data, determine the sound and image position of the first sound source data according to the sound scene to which the first sound source data belongs.
[0226] In some embodiments, determining the sound image position of at least one first sound source data in the sound field space includes: when the at least one first sound source data is one, the sound image position of the first sound source data belongs to the center position.
[0227] In some embodiments, the first determination module 1601 is configured to: when the sound and image state of the first sound source data is in a non-transition stage and the at least one first sound source data is multiple, determine the sound and image position of the first sound source data according to the sound scene to which the first sound source data belongs; wherein the non-transition stage refers to the sound and image position of the first sound source data in different time periods of the same audio track remains unchanged.
[0228] In some embodiments, determining the sound and image position of the first sound source data in the sound field space includes: when the sound and image state of the first sound source data is in a transition stage, determining the sound and image position of the first sound source data according to the sound and image position and the sound and image position update step of the first sound source data played in the previous time period; wherein the transition stage refers to the change of the sound and image position of the first sound source data in different time periods of the same audio track.
[0229] In some embodiments, when the sound image state of the first sound source data is in a transitional stage, the sound image position of the first sound source data played in the initial time period is equal to the first spatial coordinate.
[0230] In some embodiments, at the end of the transition phase of the first sound source data, the sound and image position of the first sound source data played in the current time period is equal to the sound and image position corresponding to the sound scene to which the first sound source data belongs; wherein the first spatial coordinate is the same or different from the sound and image position corresponding to the sound scene to which the first sound source data belongs.
[0231] In some embodiments, the sound source data processing device 160 further includes a generation module configured to generate second attribute information of the first sound source data, wherein the second attribute information is used to record whether the sound and image state of the first sound source data is a transition stage or a non-transition stage.
[0232] In some embodiments, the generating module is further configured to: when the transition phase of the first sound source data ends, update the second attribute information of the first sound source data to indicate that the sound and image state is in a non-transition phase.
[0233] In some embodiments, the transition stage refers to the gradual transition of the sound and image position from the first spatial coordinate belonging to the scene position to the center position according to the sound and image position update step when the first sound source data is played at different time periods, and / or the gradual transition from the center position to the sound and image position corresponding to the sound scene to which the first sound source data belongs according to the sound and image position update step.
[0234] In some embodiments, the first determination module 1601 is further configured to: determine the first sound source data occupying the center position according to the sound scene of the at least one first sound source data.
[0235] In some embodiments, for multiple first sound source data of the same priority, the first sound source data that occupies the center position is determined in order of creation time from near to far from the current moment.
[0236] In some embodiments, determining the second sound source data based on the sound and image position of the first sound source data and the first sound source data includes: rendering the first sound source data according to the sound and image position of the first sound source data to obtain the second sound source data.
[0237] In some embodiments, the sound source data processing device 160 further includes a generating module configured to generate third attribute information after the first determining module 1601 determines the sound image position of the first sound source data in the sound field space, wherein the third attribute information is used to record the sound image position of the first sound source data;
[0238] The rendering of the first sound source data according to the sound and image position of the first sound source data to obtain the second sound source data includes: rendering the first sound source data according to the third attribute information to obtain the second sound source data.
[0239] In some embodiments, the first determination module 1601 is configured to determine the sound image position of the first sound source data belonging to the center position based on the number of channels of the first sound source data belonging to the center position.
[0240] In some embodiments, mixing and then playing the second sound source data of the at least one first sound source data includes: superimposing the second sound source data of the at least one first sound source data to obtain multi-channel data to be played; and playing the multi-channel data to be played through a playback thread.
[0241] In some embodiments, the at least one first audio source data belongs to the same playback thread type.
[0242] In some embodiments, the sound source data processing device 160 also includes a generation module, which is configured to determine the playback thread to which the at least one first sound source data belongs before mixing the second sound source data of the at least one first sound source data and playing it; and generate fourth attribute information based on the playback thread to which the first sound source data belongs, and the fourth attribute information is used to record the playback thread to which the first sound source data belongs.
[0243] In some embodiments, the at least one first sound source data satisfies a validity condition; wherein the validity condition includes: the application identifier of the application to which the at least one first sound source data belongs is in the application whitelist; and / or the sampling rate of the at least one first sound source data is at least one supported sampling rate.
[0244] In some embodiments, the sound source data processing device 160 also includes a generation module, which is configured to generate fifth attribute information based on determining that the at least one first sound source data satisfies the validity condition before determining the sound and image position of the at least one first sound source data in the sound field space, and the fifth attribute information is equal to a first value; wherein the first value is used to indicate that the at least one first sound source data is played using the sound source processing method.
[0245] In some embodiments, determining the sound and image position of at least one first sound source data in the sound field space includes: determining the sound and image position of at least one first sound source data in the sound field space when the fifth attribute information of the at least one first sound source data is equal to the first value.
[0246] The description of the above device embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of the present application, please refer to the description of the method embodiment of the present application for understanding.
[0247] It should be noted that the division of modules in the embodiments of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, each functional unit in each embodiment of the present application may be integrated into a processing unit, or may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit. It may also be implemented in the form of a combination of software and hardware.
[0248] It should be noted that in the embodiment of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application can be essentially or partly embodied in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium, including several instructions to enable an electronic device to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.
[0249] An embodiment of the present application provides an electronic device, Fig.17 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application, such as Fig.17 As shown, the electronic device 170 includes a memory 1701 and a processor 1702. The memory 1701 stores a computer program that can be run on the processor 1702. When the processor 1702 executes the program, the steps in the method provided in the above embodiment are implemented.
[0250] It should be noted that the memory 1701 is configured to store instructions and applications executable by the processor 1702, and can also cache data to be processed or already processed by the processor 1702 and various modules in the electronic device 170 (for example, image data, audio data, voice communication data, and video communication data), which can be implemented through flash memory (FLASH) or random access memory (Random Access Memory, RAM).
[0251] The present application embodiment provides an audio system, Fig.18 A schematic diagram of the structure of the audio system provided in the embodiment of the present application is shown in FIG. Fig.18 As shown, the audio system 180 includes an electronic device 170 and a listening device 1801 ; wherein the listening device 1801 is used to listen to the mixed audio source data played by the electronic device 170 .
[0252] An embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the method provided in the above embodiment are implemented.
[0253] An embodiment of the present application provides a computer program product including instructions, which, when executed on a computer, enables the computer to execute the steps of the method provided in the above method embodiment.
[0254] It should be noted here that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium, storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.
[0255] It should be understood that "one embodiment" or "an embodiment" or "some embodiments" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in one embodiment" or "in some embodiments" appearing throughout the specification may not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application. The above-mentioned sequence numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments. The above description of each embodiment tends to emphasize the differences between the various embodiments, and the same or similar aspects can be referenced to each other. For the sake of brevity, this article will not repeat them.
[0256] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there may be three relationships. For example, object A and / or object B can represent three situations: object A exists alone, object A and object B exist at the same time, and object B exists alone.
[0257] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0258] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as: multiple modules or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be electrical, mechanical or other forms.
[0259] The modules described above as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules; they may be located in one place or distributed on multiple network units; some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment.
[0260] In addition, all functional modules in the embodiments of the present application may be integrated into one processing unit, or each module may be a separate unit, or two or more modules may be integrated into one unit; the above-mentioned integrated modules may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0261] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, etc., various media that can store program codes.
[0262] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can essentially or in other words, the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling an electronic device to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0263] The methods disclosed in several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0264] The features disclosed in several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0265] The features disclosed in several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0266] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A method for processing sound source data, characterized in that: The method comprises: Determining a sound image position of at least one first sound source data in a sound field space; determining second sound source data according to the sound image position of the first sound source data and the first sound source data; The second sound source data of the at least one first sound source data is mixed and then played.
2. The method according to claim 1, characterized in that Before determining the sound image position of the first sound source data in the sound field space, the method further includes: Acquire at least one third sound source data created; The third sound source data is processed according to the resampling parameters and the volume processing parameters to obtain the first sound source data.
3. The method according to claim 2, characterized in that When the sound scene of the third sound source data is a non-voice call scene, the third sound source data is processed according to the resampling parameters and the volume processing parameters to obtain the first sound source data.
4. The method according to claim 3, characterized in that The step of processing the third sound source data according to the resampling parameter and the volume processing parameter to obtain the first sound source data includes: resampling and volume processing the third sound source data according to the resampling parameter and the volume processing parameter to obtain fourth sound source data; The fourth sound source data is superimposed on an empty buffer to obtain the first sound source data.
5. The method according to claim 4, characterized in that Different third sound source data use the same resampling parameters.
6. The method according to claim 1, characterized in that Before determining the sound image position of the first sound source data in the sound field space, the method further includes: Acquire at least one third sound source data created; When the sound scene of the third sound source data is a voice call scene, the third sound source data is resampled and the data type is converted, or the third sound source data is converted to obtain the first sound source data.
7. The method according to claim 6, characterized in that Resampling and converting the data type of the third sound source data to obtain the first sound source data includes: When the first sampling rate of the third sound source data is less than or greater than the second sampling rate, and the first data type of the third sound source data is not the second data type, the third sound source data is resampled and the data type is converted according to the second sampling rate and the second data type to obtain the first sound source data.
8. The method according to claim 6, characterized in that Converting the data type of the third sound source data to obtain the first sound source data includes: When the first sampling rate of the third sound source data is equal to the second sampling rate, and the first data type of the third sound source data is not the second data type, the third sound source data is converted into a data type according to the second data type to obtain the first sound source data.
9. The method according to claim 7, characterized in that: The mixing and playing of the second sound source data of the at least one first sound source data comprises: According to the first data type and the first sampling rate, performing inverse data type conversion and inverse resampling on the second sound source data to obtain fifth sound source data; The fifth sound source data is played.
10. The method according to claim 8, characterized in that The mixing and playing of the second sound source data of the at least one first sound source data comprises: According to the first data type, performing an inverse data type conversion on the second sound source data to obtain fifth sound source data; Play the fifth sound source data.
11. The method according to claim 9 or 10, characterized in that: The playing of the fifth sound source data comprises: The fifth sound source data is played through a playback thread adapted to the voice call scenario.
12. The method according to any one of claims 2 to 5, characterized in that: The at least one third sound source data includes sound source data that does not need to be spatialized and is identified based on a corresponding destination cache address.
13. The method according to claim 12, characterized in that The sound source data that does not need to be spatialized includes non-multi-channel sound source data.
14. The method according to any one of claims 1 to 13, characterized in that: Determining the sound image position of the first sound source data in the sound field space includes: The sound image position of the first sound source data is determined according to the sound scene to which the first sound source data belongs.
15. The method according to claim 14, characterized in that The step of determining the sound image position of the first sound source data according to the sound scene to which the first sound source data belongs includes: According to the sound scene to which the first sound source data belongs, the sound and image position of the first sound source data is determined from a first mapping table; wherein the first mapping table records the sound and image position corresponding to at least one sound scene.
16. The method according to claim 15, characterized in that In the first mapping table, different sound scenes correspond to different sound image positions.
17. The method according to claim 15, characterized in that The sound image position corresponding to at least one sound scene in the first mapping table is predefined and / or configured according to the sound image position of the corresponding sound scene used by the second sound source data previously determined.
18. The method according to claim 15, characterized in that The method further comprises: receiving configuration information, the configuration information including specifying a sound image position of at least one sound scene; According to the configuration information, the sound and image positions of the corresponding sound scenes in the first mapping table are configured.
19. The method according to claim 15, characterized in that The method further comprises: receiving an update signal, the update signal being used to instruct to update the sound image position of at least one specific sound scene, the update signal comprising identification information of the at least one specific sound scene and a new sound image position; According to the update signal, the sound image position of the at least one specific sound scene in the first mapping table is updated.
20. The method according to any one of claims 1 to 13, characterized in that: Determining the sound image position of the first sound source data in the sound field space includes: Receiving a control signal based on a UI interface input, the control signal including an application identifier and a corresponding sound and image position; According to the control signal, the sound image position of the first sound source data of the application corresponding to the application identifier is determined.
21. The method according to any one of claims 14 to 19, characterized in that: Before determining the sound image position of the first sound source data according to the sound scene to which the first sound source data belongs, the method further includes: Obtaining an application identifier and a stream type of the first audio source data; A sound scene to which the first sound source data belongs is determined according to an application identifier and a stream type of the first sound source data.
22. The method according to claim 21, characterized in that The determining, according to the application identifier and the stream type of the first sound source data, the sound scene to which the first sound source data belongs includes: Determining a corresponding default sound scene according to an application identifier of the first sound source data; When the stream type of the first sound source data does not belong to the first stream type, the sound scene to which the first sound source data belongs is the default sound scene.
23. The method according to claim 22, characterized in that When the stream type of the first sound source data does not belong to the first stream type and the sampling rate of the first sound source data is not equal to the third sampling rate, the sound scene to which the first sound source data belongs is the default sound scene; wherein the third sampling rate is the sampling rate of the sound source data of human voice.
24. The method according to claim 23, characterized in that The method further comprises: When the stream type of the first sound source data does not belong to the first stream type and the sampling rate of the first sound source data is equal to the third sampling rate, the sound scene to which the first sound source data belongs is a voice call; wherein the third sampling rate is the sampling rate of the sound source data of human voice.
25. The method according to claim 22, characterized in that The method further comprises: When the stream type of the first sound source data belongs to the first stream type, the sound scene to which the first sound source data belongs is the stream type of the first sound source data.
26. The method according to any one of claims 22 to 25, characterized in that: The first stream type includes at least one of the following: a ring tone, a notification tone, a voice call, and an alarm.
27. The method according to claim 21, characterized in that The method further comprises: After determining the sound scene to which the first sound source data belongs, first attribute information of the first sound source data is generated; wherein the first attribute information is used to record a scene identifier of the sound scene to which the first sound source data belongs.
28. The method according to claim 27, characterized in that The determining the sound image position of the first sound source data from a first mapping table according to the sound scene to which the first sound source data belongs includes: According to the first attribute information of the first sound source data, the sound image position corresponding to the scene identifier is searched in a first mapping table.
29. The method according to any one of claims 14 to 28, characterized in that: In the case that there are plural pieces of the at least one first sound source data, the sound image position of the first sound source data is determined according to the sound scene to which the first sound source data belongs.
30. The method according to any one of claims 1 to 13, characterized in that: The determining of the sound image position of at least one first sound source data in the sound field space includes: In the case where the at least one first sound source data is one, the sound image position of the first sound source data is a center position.
31. The method according to claim 29, characterized in that When the sound and image state of the first sound source data is in a non-transition stage and there are multiple first sound source data, the sound and image position of the first sound source data is determined according to the sound scene to which the first sound source data belongs; wherein the non-transition stage means that the sound and image position of the first sound source data in different time periods of the same audio track remains unchanged.
32. The method according to any one of claims 1 to 13, characterized in that Determining the sound image position of the first sound source data in the sound field space includes: When the sound and image state of the first sound source data is in a transition stage, the sound and image position of the first sound source data is determined according to the sound and image position and the sound and image position update step of the first sound source data played in the previous time period; wherein the transition stage refers to the change of the sound and image position of the first sound source data in different time periods of the same audio track.
33. The method according to claim 32, characterized in that When the sound image state of the first sound source data is in a transitional stage, the sound image position of the first sound source data played in the initial time period is equal to the first space coordinate.
34. The method according to claim 32 or 33, characterized in that At the end of the transition phase of the first sound source data, the sound and image position of the first sound source data played in the current time period is equal to the sound and image position corresponding to the sound scene to which the first sound source data belongs; wherein the first spatial coordinate is the same as or different from the sound and image position corresponding to the sound scene to which the first sound source data belongs.
35. The method according to any one of claims 31 to 34, characterized in that The method further includes: generating second attribute information of the first sound source data, wherein the second attribute information is used to record whether the sound and image state of the first sound source data is a transition stage or a non-transition stage.
36. The method according to claim 35, characterized in that The method further comprises: When the transition phase of the first sound source data ends, the second attribute information of the first sound source data is updated to indicate that the sound image state is in a non-transition phase.
37. The method according to any one of claims 32 to 36, characterized in that: The transition stage means that when the first sound source data is played at different time periods, the sound and image position gradually transitions from the first spatial coordinate belonging to the scene position to the center position according to the sound and image position update step, and / or gradually transitions from the center position to the sound and image position corresponding to the sound scene to which the first sound source data belongs according to the sound and image position update step.
38. The method according to any one of claims 14, 21-28, characterized in that: According to the sound scene of the at least one first sound source data, the first sound source data occupying the center position is determined.
39. The method according to claim 38, characterized in that For multiple first sound source data of the same priority, the first sound source data that occupies the center position is determined in order of creation time from near to far from the current moment.
40. The method according to any one of claims 1 to 39, characterized in that The step of determining the second sound source data according to the sound image position of the first sound source data and the first sound source data includes: The first sound source data is rendered according to the sound and image position of the first sound source data to obtain second sound source data.
41. The method according to claim 40, characterized in that The method further includes: after determining the sound image position of the first sound source data in the sound field space, generating third attribute information, the third attribute information being used to record the sound image position of the first sound source data; The rendering of the first sound source data according to the sound and image position of the first sound source data to obtain the second sound source data includes: rendering the first sound source data according to the third attribute information to obtain the second sound source data.
42. The method of claim 30, 37, 38 or 39, wherein: Based on the number of channels of the first sound source data belonging to the center position, the sound image position of the first sound source data belonging to the center position is determined.
43. The method according to any one of claims 2-5, 12-13, characterized in that: The mixing and playing of the second sound source data of the at least one first sound source data comprises: Performing superposition processing on the second sound source data of the at least one first sound source data to obtain multi-channel data to be played; The multi-channel data to be played is played through a playing thread.
44. The method according to any one of claims 1 to 43, characterized in that The at least one first audio source data belongs to the same playback thread type.
45. The method according to claim 44, characterized in that Before mixing and playing the second sound source data of the at least one first sound source data, the method further includes: Determining a playback thread to which the at least one first audio source data belongs; According to the playback thread to which the first sound source data belongs, fourth attribute information is generated, and the fourth attribute information is used to record the playback thread to which the first sound source data belongs.
46. The method according to any one of claims 1 to 43, characterized in that The at least one first sound source data satisfies a validity condition; wherein the validity condition includes: The application identifier of the application to which the at least one first audio source data belongs is in the application whitelist; and / or, The sampling rate of the at least one first audio source data is at least one supported sampling rate.
47. The method according to claim 46, characterized in that Before determining the sound image position of at least one first sound source data in the sound field space, the method further includes: Based on determining that the at least one first sound source data satisfies the validity condition, fifth attribute information is generated, and the fifth attribute information is equal to a first value; wherein the first value is used to indicate that the at least one first sound source data is played using the sound source processing method.
48. A sound source data processing device, characterized in that: include: A first determination module, configured to determine a sound image position of at least one first sound source data in a sound field space; A second determination module is configured to determine second sound source data according to the sound image position of the first sound source data and the first sound source data; The mixing and playing module is configured to mix and play the second sound source data of the at least one first sound source data.
49. An electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 47 is implemented.
50. An audio system comprising the electronic device and the listening device of claim 49; wherein: The listening device is used to listen to the mixed sound source data played by the electronic device.
51. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 47 is implemented.