Sound source data processing method and device, equipment and storage medium

By assigning different audio and video locations to the sound source data of music scenes and other scenes, the problem of content that cannot be heard clearly due to sound interference when multiple sounds are played concurrently is solved, and the user's listening experience is improved.

CN119946542APending Publication Date: 2025-05-06GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311452221.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-02
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In scenes where multiple types of sounds are played concurrently, electronic devices may cause problems that the content of a certain sound cannot be heard clearly.

Method used

The two audio source data are distinguished in the sense of hearing by assigning different audio and image locations to the first sound source data belonging to the music scene and outputting the sound source data based on these locations.

Benefits of technology

It effectively solves the problem that the content cannot be heard clearly due to the interference of sounds when music type sound source data and other types of sound source data are played concurrently, and improves the user's listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946542A_ABST
    Figure CN119946542A_ABST
Patent Text Reader

Abstract

The invention provides a sound source data processing method and device, equipment and a storage medium. The method comprises the following steps: acquiring first sound source data and second sound source data; when the scene type of any one of the first sound source data and the second sound source data is a music scene, determining a first sound image position of the first sound source data in a sound field space; determining a second sound image position of the second sound source data in the sound field space; wherein the first sound image position is different from the second sound image position; and outputting the first sound source data and the second sound source data according to the first sound image position and the second sound image position.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to audio processing technology, and is related to but not limited to sound source data processing methods and devices, equipment, and storage media. Background Art

[0002] With the development of the digital intelligent era and electronic devices, there are more and more types of electronic devices, and the functions of each electronic device are becoming more and more abundant. Users can use different functions of an electronic device at the same time. For example, taking a mobile phone as an example, in one possible scenario, the user wants to listen to music while listening to a book; in another possible scenario, the user wants to listen to music while playing games.

[0003] However, in a scenario where an electronic device plays multiple types of sounds concurrently, there may be a problem in which a user cannot clearly hear the content of a certain sound (such as the content of a book). Summary of the invention

[0004] In view of this, the sound source data processing method, device, equipment, and storage medium provided in the present application can solve the problem that the content of a certain sound cannot be heard clearly due to mutual interference of sounds when music-type sound source data and other types of sound source data are played concurrently.

[0005] According to one aspect of an embodiment of the present application, a sound source data processing method is provided, including: acquiring first sound source data and second sound source data; determining a first sound image position of the first sound source data in a sound field space when the scene type of any one of the first sound source data and the second sound source data is a music scene; determining a second sound image position of the second sound source data in the sound field space; wherein the first sound image position is different from the second sound image position; and outputting the first sound source data and the second sound source data according to the first sound image position and the second sound image position.

[0006] According to one aspect of an embodiment of the present application, a sound source data processing device is provided, including: an acquisition module, configured to acquire first sound source data and second sound source data; a determination module, configured to determine a first sound image position of the first sound source data in a sound field space when the scene type of any one of the first sound source data and the second sound source data is a music scene; and determine a second sound image position of the second sound source data in the sound field space; wherein the first sound image position is different from the second sound image position; and an output module, configured to output the first sound source data and the second sound source data according to the first sound image position and the second sound image position.

[0007] According to one aspect of an embodiment of the present application, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, the method described in the embodiment of the present application is implemented.

[0008] According to one aspect of an embodiment of the present application, an audio system is provided, comprising the electronic device and a listening device described in the embodiment of the present application; wherein the listening device is used to listen to the sound source data played by the electronic device.

[0009] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method provided in the embodiment of the present application is implemented.

[0010] In an embodiment of the present application, a sound source data processing method is provided. In the method, for the case where first sound source data (i.e., music-type sound source data) and other sound source data (i.e., second sound source data) belonging to a music scene are played concurrently, different sound and image positions are first assigned to the first sound source data and the second sound source data, and based on the assigned sound and image positions, the first sound source data and the second sound source data are output; in this way, the output first sound source data and the second sound source data have different sound and image positions in terms of auditory sense, thereby solving the problem of inability to hear the content of a certain sound due to mutual interference of sounds when music-type sound source data and other types of sound source data are played concurrently.

[0011] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and are used together with the specification to illustrate the technical solution of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0013] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0014] Figure 1 Schematic diagram of the implementation process of the sound source data processing method provided in the embodiment of the present application Figure 1 ;

[0015] Figure 2 A schematic diagram of a specific implementation flow of step 104 provided in an embodiment of the present application;

[0016] Figure 3 A schematic diagram of a specific implementation flow of step 1043 provided in an embodiment of the present application;

[0017] Figure 4 Schematic diagram of the implementation process of the sound source data processing method provided in the embodiment of the present application Figure 2 ;

[0018] Figure 5 A schematic diagram of the position distribution of the sound of books and music relative to the human head provided in an embodiment of the present application;

[0019] Figure 6 Schematic diagram of the implementation process of the sound source data processing method provided in the embodiment of the present application Figure 3 ;

[0020] Figure 7 A schematic diagram of the structure of a sound source data processing device provided in an embodiment of the present application;

[0021] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0022] Fig. 9 A schematic diagram of the structure of an audio system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the specific technical solution of the present application will be further described in detail below in conjunction with the drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.

[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0025] In the following description, reference is made to “some embodiments”, “this embodiment”, “embodiments of the present application” and examples, etc., which describe a subset of all possible embodiments, but it can be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0026] The descriptions of “first, second, third” etc. that appear in the embodiments of the present application are only for illustration and distinction of the description objects. There is no distinction of order, nor do they indicate any special limitation on the number of devices in the embodiments of the present application, and cannot constitute any limitation on the embodiments of the present application.

[0027] The embodiment of the present application provides a method for processing audio source data, which is applied to an electronic device. During the implementation, the electronic device can be various types of devices with audio processing and output capabilities, for example, the electronic device can include a mobile phone, a tablet computer, a television, a projector, etc. The function implemented by the method can be implemented by calling a program code by a processor in the electronic device. Of course, the program code can be stored in a computer storage medium. It can be seen that the electronic device at least includes a processor and a storage medium.

[0028] Figure 1 Schematic diagram of the implementation process of the sound source data processing method provided in the embodiment of the present application Figure 1 ,like Figure 1 As shown, the method may include the following steps 101 to 104:

[0029] Step 101, obtaining first sound source data and second sound source data;

[0030] Step 102, when the scene type of any one of the first sound source data and the second sound source data is a music scene, determining a first sound image position of the first sound source data in a sound field space; and

[0031] Step 103, determining a second sound image position of the second sound source data in the sound field space; wherein the first sound image position is different from the second sound image position;

[0032] Step 104: output the first sound source data and the second sound source data according to the first sound image position and the second sound image position.

[0033] In an embodiment of the present application, a sound source data processing method is provided. In the method, for the case where first sound source data (i.e., music-type sound source data) and other sound source data (i.e., second sound source data) belonging to a music scene are played concurrently, different sound and image positions are first assigned to the first sound source data and the second sound source data, and based on the assigned sound and image positions, the first sound source data and the second sound source data are output; in this way, the output first sound source data and the second sound source data have different sound and image positions in terms of auditory sense, thereby solving the problem of inability to hear the content of a certain sound due to mutual interference of sounds when music-type sound source data and other types of sound source data are played concurrently.

[0034] The following describes further optional implementations and related terms of each of the above steps.

[0035] In step 101, first sound source data and second sound source data are obtained.

[0036] In the embodiment of the present application, the first sound source data and the second sound source data can be created by the same application or by different applications. The first sound source data and the second sound source data can be understood as track data (Track), which includes audio data of at least one channel.

[0037] In step 102, when the scene type of any one of the first sound source data and the second sound source data is a music scene, a first sound image position of the first sound source data in the sound field space is determined; and

[0038] In step 103, a second sound image position of the second sound source data in the sound field space is determined; wherein the first sound image position is different from the second sound image position.

[0039] It is understandable that after acquiring the first sound source data and the second sound source data, the electronic device needs to at least determine whether the scene type of any of the sound source data is a music scene. If the scene type of any of the sound source data is a music scene, different sound and image positions are set for the acquired sound source data.

[0040] Exemplarily, in some embodiments, the electronic device may set different sound and image positions for the first sound source data and the second sound source data when the scene type of the first sound source data is a music scene and the scene type of the second sound source data is one of the scenes in audiobooks, videos, games, voice calls, navigation, or learning applications.

[0041] In the embodiments of the present application, there is no limitation on how to determine the scene type to which the sound source data belongs. For example, in some embodiments, the electronic device can determine the scene type of the sound source data according to the metadata or audio content of the sound source data; wherein the sound source data is the first sound source data or the second sound source data.

[0042] It is understandable that when a certain application creates sound source data, it will also specify corresponding metadata, such as application package name, sampling rate and / or stream type, etc. Therefore, in some embodiments, the electronic device can determine whether the scene type of the sound source data is a music scene according to whether the application identifier (such as application package name) of the created sound source data is in the music application whitelist; if its application identifier is in the music application whitelist, the scene type of the sound source data is determined to be a music scene.

[0043] For the embodiments of determining the scene type of the sound source data based on the metadata of the sound source data, further, in some embodiments, the electronic device can determine that the scene type of the sound source data is a music scene based on determining that the application identifier (such as the application package name) of the sound source data is in the music application whitelist.

[0044] In a possible implementation, the music application whitelist records the application package names of one or more music applications; if the electronic device finds the application package name that records the sound source data in the music application whitelist, it determines that the scene type of the sound source data is a music scene.

[0045] The application identifiers in the music application whitelist may include default application identifiers and / or user-set application identifiers, etc. These application identifiers may include application identifiers of music applications installed by the electronic device and / or application identifiers of music applications not installed by the electronic device.

[0046] For the embodiment of determining the scene type based on the audio content of the sound source data, further, in some embodiments, the electronic device can input the sound source data into a trained neural network model, thereby obtaining the scene type of the sound source data based on the model. The training data set of the neural network model includes sound source data of multiple different scene types and their corresponding scene type labels.

[0047] In some embodiments, taking the scene type of the first sound source data as a music scene as an example, the first sound image position is a scene position, and the second sound image position is a center position; wherein the center position is different from the scene position.

[0048] In a possible implementation, for the center position, that is, the second sound image position, the electronic device may determine the second sound image position according to the number of channels of the channel data included in the second sound source data.

[0049] In step 104, the first sound source data and the second sound source data are output according to the first sound image position and the second sound image position.

[0050] The purpose of the electronic device executing step 104 is to make the first sound source data and the second sound source data finally played have different sound image positions in terms of auditory perception, so that the user can clearly hear the contents of the different sound source data, thereby enhancing the user's listening experience.

[0051] In the embodiments of the present application, considering that the volume of different sound source data is also a factor that causes the user to be unable to hear different sounds clearly, therefore, in some embodiments, the electronic device needs to perform amplitude processing on the first sound source data and / or the second sound source data before outputting the first sound source data and the second sound source data, so that after the amplitude processing, the amplitude of the sound source data belonging to the music scene is lower than the amplitude of the other sound source data; the electronic device renders and outputs the sound source data and the other sound source data belonging to the music scene after the amplitude processing according to the first sound image position and the second sound image position; in this way, the first sound source data and the second sound source data that are finally output have not only different sound image positions, but also different volumes, and the volume of the music is lower, thereby solving the problem that when multiple sounds are played simultaneously, the user cannot hear the content of the sound because the volume of the sound at the center position is low, thereby further enhancing the user's listening experience.

[0052] In the embodiment of the present application, whether the electronic device needs to perform amplitude processing on at least one of the first sound source data and the second sound source data before outputting them can be unconditional or conditional. Among them, for the unconditional embodiment, it can be understood that no matter what the magnitude relationship between the first amplitude of the first sound source data and the second amplitude of the second sound source data is, at least one of them needs to be amplitude processed before output.

[0053] For conditional embodiments, illustratively, in some embodiments, such as Figure 2 As shown, the electronic device can implement step 104 through the following steps 1041 to 1044:

[0054] Step 1041, determine whether the first amplitude of the first sound source data is greater than or equal to the second amplitude of the second sound source data; if so, execute step 1042; otherwise, execute step 1044.

[0055] Alternatively, in some other embodiments, the electronic device may execute step 1042 if the result of subtracting the second amplitude from the first amplitude is greater than or equal to an amplitude threshold; otherwise, execute step 1044; wherein the amplitude threshold is greater than 0.

[0056] Step 1042, performing amplitude reduction processing on the first sound source data to obtain third sound source data; wherein the amplitude of the third sound source data is smaller than the second amplitude of the second sound source data.

[0057] Step 1043: Render and output the third sound source data and the second sound source data according to the first sound image position and the second sound image position.

[0058] Figure 3The specific implementation flow diagram of step 1043 provided in the embodiment of the present application is as follows: Figure 3 As shown, the electronic device can implement step 1043 through the following steps 301 to 303:

[0059] Step 301, rendering the third sound source data according to the first sound image position to obtain rendered third sound source data;

[0060] For step 301, in a possible implementation, the electronic device may input the first sound image position and the third sound source data into a spatial audio rendering algorithm, so that the spatial audio rendering algorithm performs spatial processing on the third sound source data according to the first sound image position to render and output the sound source data with the first sound image position. For example, the spatial audio rendering algorithm selects a corresponding head-related transfer function (HRTF) according to the input sound image position, and renders the corresponding audio signal through the HRTF function. Among them, the parameters of the HRTF function corresponding to different sound image positions are different.

[0061] Step 302: Render the second sound source data according to the second sound image position to obtain rendered second sound source data.

[0062] As for step 302, its implementation method is the same as that of step 301, except that the input sound image position and sound source data are different, so it will not be described here in detail.

[0063] Step 303: Mix the rendered third sound source data with the rendered second sound source data and output the mixed sound.

[0064] The so-called mixing includes linear superposition or nonlinear superposition of multiple sound source data. Taking the case where the rendered third sound source data includes left channel data and right channel data, and the rendered second sound source data includes left channel data and right channel data as an example, in a possible implementation, the electronic device can implement step 303 as follows: superimpose the left channel data of the rendered third sound source data with the left channel data of the rendered second sound source data (such as linear superposition or nonlinear superposition), and superimpose the right channel data of the rendered third sound source data with the right channel data of the rendered second sound source data (such as linear superposition or nonlinear superposition), thereby obtaining the mixed dual-channel data.

[0065] Step 1044: Render and output the first sound source data and the second sound source data according to the first sound image position and the second sound image position.

[0066] The electronic device implements step 1044 in the same manner as the above step 1043, except that the objects to be rendered are different, and in step 1044, the objects to be rendered are the first sound source data and the second sound source data. Therefore, further implementations of step 1044 are not described here, and can be understood by referring to steps 301 to 303.

[0067] Of course, for the above Figure 2 The embodiment shown may also have other alternatives. When the first amplitude of the first sound source data is greater than or equal to the second amplitude of the second sound source data, or the result of subtracting the second amplitude from the first amplitude is greater than or equal to the amplitude threshold, the electronic device may not perform step 1042, but may perform amplitude enhancement processing on the second sound source data to obtain fourth sound source data; wherein the amplitude of the fourth sound source data is greater than the first amplitude of the first sound source data; and then, the electronic device renders and outputs the first sound source data and the fourth sound source data according to the first sound image position and the second sound image position.

[0068] Alternatively, in other embodiments, when the first amplitude of the first sound source data is greater than or equal to the second amplitude of the second sound source data, or the result of subtracting the second amplitude from the first amplitude is greater than the amplitude threshold, the electronic device may not execute step 1042, but instead perform amplitude weakening processing on the first sound source data to obtain fifth sound source data; and perform amplitude enhancement processing on the second sound source data to obtain sixth sound source data; wherein the amplitude of the fifth sound source data is less than the amplitude of the sixth sound source data; and then, the electronic device renders and outputs the fifth sound source data and the sixth sound source data according to the first sound image position and the second sound image position.

[0069] Regardless of which sound source data is ultimately rendered and output based on the sound and image position, the method of achieving the rendering output is the same. Suppose we call the rendered and output sound source data the seventh sound source data and the eighth sound source data; wherein the seventh sound source data may be the first sound source data, and the eighth sound source data may be the second sound source data; or, the seventh sound source data may be the third sound source data, and the eighth sound source data may be the second sound source data; or, the seventh sound source data may be the first sound source data, and the eighth sound source data may be the fourth sound source data; or, the seventh sound source data may be the fifth sound source data, and the eighth sound source data may be the sixth sound source data.

[0070] Based on this, in some embodiments, according to the first sound image position and the second sound image position, the seventh sound source data and the eighth sound source data are rendered and output, including: according to the first sound image position, the seventh sound source data is rendered to obtain the rendered seventh sound source data; according to the second sound image position, the eighth sound source data is rendered to obtain the rendered eighth sound source data; the rendered seventh sound source data and the rendered eighth sound source data are mixed and output.

[0071] The present application provides a method for processing sound source data. Figure 4 Schematic diagram of the implementation process of the sound source data processing method provided in the embodiment of the present application Figure 2 ,like Figure 4 As shown, the method includes the following steps 401 to 408:

[0072] Step 401, obtaining first sound source data and second sound source data; executing steps 402 and 404;

[0073] Step 402, determine whether the scene types of the first sound source data and the second sound source data are respectively a music scene and a book listening scene; if yes, execute step 403; otherwise, end.

[0074] Step 403, setting a first sound image position of the first sound source data in the sound field space, and setting a second sound image position of the second sound source data in the sound field space; wherein the first sound image position is different from the second sound image position.

[0075] For example, Figure 5 A schematic diagram of the position distribution of the sound of books and music provided in an embodiment of the present application relative to the human head (i.e., the head wearing headphones), as shown in FIG. Figure 5 As shown, the book listening sound is placed in the middle position in front of the human head, and the music sound is placed behind the human head, so that it has the effect of background sound.

[0076] Step 404, determining a first amplitude of the first sound source data and a second amplitude of the second sound source data;

[0077] Step 405, determining whether the first amplitude is greater than or equal to the second amplitude; if yes, executing step 406; otherwise, executing step 408;

[0078] Alternatively, in some other embodiments, when the result of subtracting the second amplitude from the first amplitude is greater than or equal to the amplitude threshold, step 406 is executed; otherwise, step 408 is executed.

[0079] Step 406, performing amplitude weakening processing on the first sound source data to obtain third sound source data; wherein the amplitude of the third sound source data is smaller than the second amplitude of the second sound source data;

[0080] Step 407: Render and output the third sound source data and the second sound source data according to the first sound image position and the second sound image position.

[0081] Step 408: Render and output the first sound source data and the second sound source data according to the first sound image position and the second sound image position.

[0082] It is understandable that some users have the habit of listening to music while reading books, because the audio content in some audiobook applications is relatively monotonous, and playing the user's favorite music at this time will make the sound richer. However, various systems currently do not specifically optimize for the concurrent use of audiobook applications and music applications, but directly superimpose the audio source data of the two applications linearly and play them out.

[0083] The sound from different audio sources is directly superimposed linearly and then played back. The sound after such processing usually has serious problems in terms of sound confusion, inability to highlight the key points, or even unclear audio content. For example, the superimposed music sound may cover the sound of the audiobook, making the sound of the audiobook application inaudible. Or the sound of the music and the audiobook are superimposed and interfere with each other, which may even cause user disgust and bring a bad listening experience.

[0084] Based on this, an exemplary application of an embodiment of the present application in a practical application scenario will be described below.

[0085] In the embodiment of the present application, the sound object rendering capability of the spatial audio rendering algorithm is used to distribute the sounds of the book and music in different locations. Figure 5 The schematic diagram of the position distribution of the sound of the book and music relative to the human head provided in the embodiment of the present application is as follows: Figure 5 As shown, the audio of the book is placed in the middle of the front of the head, and the music sound is placed behind the head, giving it a background sound effect. The two sounds processed in this way will coexist harmoniously, forming a better listening experience. When performing the above processing, it is also necessary to pay attention to the relative volume between the audio source data of the audio book application and the audio source data of the music application. Note that the music volume should always be kept lower than the audio book volume by a certain number of decibels so that the sound of the audio book application can always be highlighted.

[0086] Figure 6 Schematic diagram of the implementation process of the sound source data processing method provided in the embodiment of the present application Figure 3 ,like Figure 6 As shown, the method includes the following steps 601 to 607:

[0087] Step 601, determine whether the current sound source data includes sound source data belonging to the music scene and sound source data belonging to the book listening scene; if yes, execute steps 602 and 603; otherwise, end;

[0088] Step 602, setting corresponding sound and image positions for the sound source data created by the music application;

[0089] Step 603, setting the corresponding sound and image position for the sound source data created by the audiobook application;

[0090] Step 604, detecting the sound source amplitude (listening_dB) of the listening application and the sound source amplitude (music_dB) of the music application;

[0091] Step 605, determine whether music_dB-listening_dB is greater than or equal to the amplitude threshold; if yes, execute step 606; otherwise, execute step 607;

[0092] Step 606, reducing the sound source amplitude of the music application; then proceeding to step 607;

[0093] Step 607: Use a 3D audio rendering algorithm to process the audio source data of the audiobook application and the audio source data of the music application to obtain dual-channel data.

[0094] In the embodiment of the present application, the problem of mutual interference between the audiobook application and the music application when they are running simultaneously is solved.

[0095] In the sound field space, place the audiobook application in the front middle position and the music application in the back position, so that the sound directions of the two applications are separated from each other and do not interfere with each other.

[0096] In some embodiments, if the user plays multiple audiobook applications or multiple music applications, the multiple audiobook applications are distributed on the front left and right sides, and the music applications are distributed on the back left and right sides according to the scene judgment.

[0097] It should be noted that although the steps of the method in the present application are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all the steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps, etc.; or, steps in different embodiments may be combined into a new technical solution.

[0098] Based on the foregoing embodiments, an embodiment of the present application provides a sound source data processing device, which includes the modules included and the units included in the modules, which can be implemented by a processor; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be an AI acceleration engine (such as NPU, etc.), GPU, central processing unit (CPU), microprocessor (MPU), digital signal processor (DSP) or field programmable gate array (FPGA), etc.

[0099] Figure 7 A schematic diagram of the structure of the sound source data processing device provided in the embodiment of the present application, such as Figure 7 As shown, the sound source data processing device 70 includes:

[0100] An acquisition module 701 is configured to acquire first sound source data and second sound source data;

[0101] The determination module 702 is configured to determine a first sound image position of the first sound source data in the sound field space when the scene type of any one of the first sound source data and the second sound source data is a music scene; and determine a second sound image position of the second sound source data in the sound field space; wherein the first sound image position is different from the second sound image position;

[0102] The output module 703 is configured to output the first sound source data and the second sound source data according to the first sound image position and the second sound image position.

[0103] In some embodiments, the determination module 702 is further configured to: determine the scene type of the sound source data according to the metadata or audio content of the sound source data; wherein the sound source data is the first sound source data or the second sound source data.

[0104] In some embodiments, the metadata includes an application identifier of an application that creates corresponding sound source data; the sound source data is the first sound source data or the second sound source data; based on the metadata of the sound source data, the scene type of the sound source data is determined, including: based on determining that the application identifier of the sound source data is in a music application whitelist, determining that the scene type of the sound source data is a music scene.

[0105] In some embodiments, the scene type of the first sound source data is a music scene, the first sound image position is a scene position in the sound field space, and the second sound image position is a center position in the sound field space.

[0106] In some embodiments, determining the second sound image position of the second sound source data in the sound field space includes: determining the second sound image position according to the number of channels of the channel data included in the second sound source data.

[0107] In some embodiments, the scene type of the first sound source data is a music scene; outputting the first sound source data and the second sound source data according to the first sound image position and the second sound image position includes: performing amplitude processing on the first sound source data and / or the second sound source data so that after the amplitude processing, the amplitude of the sound source data belonging to the music scene is smaller than the amplitude of another sound source data; rendering and outputting the sound source data belonging to the music scene and the other sound source data after the amplitude processing according to the first sound image position and the second sound image position.

[0108] In some embodiments, outputting the first sound source data and the second sound source data according to the first sound image position and the second sound image position includes: when a first amplitude of the first sound source data is greater than or equal to a second amplitude of the second sound source data, or a result of subtracting the first amplitude from the second amplitude is greater than or equal to an amplitude threshold, performing amplitude weakening processing on the first sound source data to obtain third sound source data; wherein the amplitude threshold is greater than 0, and the amplitude of the third sound source data is less than the second amplitude of the second sound source data; and rendering and outputting the third sound source data and the second sound source data according to the first sound image position and the second sound image position.

[0109] In some embodiments, outputting the first sound source data and the second sound source data according to the first sound image position and the second sound image position includes: when a first amplitude of the first sound source data is greater than or equal to a second amplitude of the second sound source data, or a result of subtracting the first amplitude from the second amplitude is greater than or equal to an amplitude threshold, performing amplitude enhancement processing on the second sound source data to obtain fourth sound source data; wherein the amplitude threshold is greater than 0, and the amplitude of the fourth sound source data is greater than the first amplitude of the first sound source data; and rendering and outputting the first sound source data and the fourth sound source data according to the first sound image position and the second sound image position.

[0110] In some embodiments, the outputting of the first sound source data and the second sound source data according to the first sound image position and the second sound image position includes: when the first amplitude of the first sound source data is greater than or equal to the second amplitude of the second sound source data, or the result of subtracting the second amplitude from the first amplitude is greater than or equal to an amplitude threshold, performing amplitude weakening processing on the first sound source data to obtain fifth sound source data; and performing amplitude enhancement processing on the second sound source data to obtain sixth sound source data; wherein the amplitude threshold is greater than 0, and the amplitude of the fifth sound source data is less than the amplitude of the sixth sound source data; and rendering and outputting the fifth sound source data and the sixth sound source data according to the first sound image position and the second sound image position.

[0111] In some embodiments, according to the first sound image position and the second sound image position, the seventh sound source data and the eighth sound source data are rendered and output, including: according to the first sound image position, the seventh sound source data is rendered to obtain the rendered seventh sound source data; according to the second sound image position, the eighth sound source data is rendered to obtain the rendered eighth sound source data; the rendered seventh sound source data and the rendered eighth sound source data are mixed and output; wherein the seventh sound source data is the first sound source data, and the eighth sound source data is the second sound source data; or, the seventh sound source data is the third sound source data, and the eighth sound source data is the second sound source data; or, the seventh sound source data is the first sound source data, and the eighth sound source data is the fourth sound source data; or, the seventh sound source data is the fifth sound source data, and the eighth sound source data is the sixth sound source data.

[0112] In some embodiments, the scene type of the first sound source data is a music scene, and the scene type of the second sound source data is one of the following scenes: listening to books, videos, games, voice calls, navigation, and learning applications.

[0113] The description of the above device embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of the present application, please refer to the description of the method embodiment of the present application for understanding.

[0114] It should be noted that the division of modules in the embodiments of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, each functional unit in each embodiment of the present application may be integrated into a processing unit, or may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit. It may also be implemented in the form of a combination of software and hardware.

[0115] It should be noted that in the embodiment of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application can be essentially or partly embodied in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium, including several instructions to enable an electronic device to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.

[0116] An embodiment of the present application provides an electronic device, Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application, such as Figure 8 As shown, the electronic device 80 includes a memory 801 and a processor 802. The memory 801 stores a computer program that can be run on the processor 802. When the processor 802 executes the program, the steps in the method provided in the above embodiment are implemented.

[0117] It should be noted that the memory 801 is configured to store instructions and applications executable by the processor 802, and can also cache data to be processed or already processed by the processor 802 and various modules in the electronic device 80 (for example, image data, audio data, voice communication data, and video communication data), which can be implemented through flash memory (FLASH) or random access memory (Random Access Memory, RAM).

[0118] The present application embodiment provides an audio system, Fig. 9 A schematic diagram of the structure of the audio system provided in the embodiment of the present application is shown in FIG. Fig. 9 As shown, the audio system 90 includes an electronic device 80 and a listening device 901 ; wherein the listening device 901 is used to listen to the audio source data played by the electronic device 80 .

[0119] An embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the method provided in the above embodiment are implemented.

[0120] An embodiment of the present application provides a computer program product including instructions, which, when executed on a computer, enables the computer to execute the steps of the method provided in the above method embodiment.

[0121] It should be noted here that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium, storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0122] It should be understood that "one embodiment" or "an embodiment" or "some embodiments" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in one embodiment" or "in some embodiments" appearing throughout the specification may not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application. The above-mentioned sequence numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments. The above description of each embodiment tends to emphasize the differences between the various embodiments, and the same or similar aspects can be referenced to each other. For the sake of brevity, this article will not repeat them.

[0123] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there may be three relationships. For example, object A and / or object B can represent three situations: object A exists alone, object A and object B exist at the same time, and object B exists alone.

[0124] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0125] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as: multiple modules or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be electrical, mechanical or other forms.

[0126] The modules described above as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules; they may be located in one place or distributed on multiple network units; some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment.

[0127] In addition, all functional modules in the embodiments of the present application may be integrated into one processing unit, or each module may be a separate unit, or two or more modules may be integrated into one unit; the above-mentioned integrated modules may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0128] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, etc., various media that can store program codes.

[0129] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can essentially or in other words, the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling an electronic device to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0130] The methods disclosed in several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0131] The features disclosed in several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0132] The features disclosed in several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0133] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A method for processing sound source data, characterized in that: The method comprises: Acquire first sound source data and second sound source data; When the scene type of any one of the first sound source data and the second sound source data is a music scene, determining a first sound image position of the first sound source data in a sound field space; and Determine a second sound image position of the second sound source data in the sound field space; wherein the first sound image position is different from the second sound image position; The first sound source data and the second sound source data are outputted according to the first sound image position and the second sound image position.

2. The method according to claim 1, characterized in that: The method further comprises: The scene type of the sound source data is determined according to metadata or audio content of the sound source data; wherein the sound source data is the first sound source data or the second sound source data.

3. The method according to claim 2, characterized in that The metadata includes an application identifier of an application that creates the corresponding sound source data; Determining the scene type of the sound source data according to metadata of the sound source data includes: Based on determining that the application identifier of the sound source data is in the music application whitelist, determining that the scene type of the sound source data is a music scene.

4. The method according to any one of claims 1 to 3, characterized in that: The scene type of the first sound source data is a music scene, the first sound image position is a scene position in the sound field space, and the second sound image position is a center position in the sound field space.

5. The method according to claim 4, characterized in that The determining of the second sound image position of the second sound source data in the sound field space comprises: The second sound image position is determined according to the number of channels of the channel data included in the second sound source data.

6. The method according to any one of claims 1 to 5, characterized in that: The scene type of the first sound source data is a music scene; The step of outputting the first sound source data and the second sound source data according to the first sound image position and the second sound image position comprises: Performing amplitude processing on the first sound source data and / or the second sound source data, so that after the amplitude processing, the amplitude of the sound source data belonging to the music scene is smaller than the amplitude of another sound source data; According to the first sound image position and the second sound image position, the sound source data belonging to the music scene and the another sound source data after the amplitude processing are rendered and output.

7. The method according to claim 6, characterized in that The step of outputting the first sound source data and the second sound source data according to the first sound image position and the second sound image position comprises: When the first amplitude of the first sound source data is greater than or equal to the second amplitude of the second sound source data, or the result of subtracting the second amplitude from the first amplitude is greater than or equal to an amplitude threshold, performing amplitude weakening processing on the first sound source data to obtain third sound source data; wherein the amplitude threshold is greater than 0, and the amplitude of the third sound source data is less than the second amplitude of the second sound source data; and The third sound source data and the second sound source data are rendered and output according to the first sound image position and the second sound image position.

8. The method according to claim 6, characterized in that The step of outputting the first sound source data and the second sound source data according to the first sound image position and the second sound image position comprises: When the first amplitude of the first sound source data is greater than or equal to the second amplitude of the second sound source data, or the result of subtracting the second amplitude from the first amplitude is greater than or equal to an amplitude threshold, performing amplitude enhancement processing on the second sound source data to obtain fourth sound source data; wherein the amplitude threshold is greater than 0, and the amplitude of the fourth sound source data is greater than the first amplitude of the first sound source data; and The first sound source data and the fourth sound source data are rendered and output according to the first sound image position and the second sound image position.

9. The method according to claim 6, characterized in that The step of outputting the first sound source data and the second sound source data according to the first sound image position and the second sound image position comprises: When the first amplitude of the first sound source data is greater than or equal to the second amplitude of the second sound source data, or the result of subtracting the second amplitude from the first amplitude is greater than or equal to an amplitude threshold, performing amplitude weakening processing on the first sound source data to obtain fifth sound source data; and Performing amplitude enhancement processing on the second sound source data to obtain sixth sound source data; wherein the amplitude threshold is greater than 0, and the amplitude of the fifth sound source data is less than the amplitude of the sixth sound source data; The fifth sound source data and the sixth sound source data are rendered and output according to the first sound image position and the second sound image position.

10. The method according to claim 6, characterized in that Rendering and outputting seventh sound source data and eighth sound source data according to the first sound image position and the second sound image position includes: Rendering the seventh sound source data according to the first sound image position to obtain rendered seventh sound source data; Rendering the eighth sound source data according to the second sound image position to obtain rendered eighth sound source data; The rendered seventh sound source data and the rendered eighth sound source data are mixed and then output; wherein, The seventh sound source data is the first sound source data, and the eighth sound source data is the second sound source data; or, The seventh sound source data is the third sound source data, and the eighth sound source data is the second sound source data; or, The seventh sound source data is the first sound source data, and the eighth sound source data is the fourth sound source data; or, The seventh sound source data is the fifth sound source data, and the eighth sound source data is the sixth sound source data.

11. The method according to any one of claims 1 to 10, characterized in that: The scene type of the first sound source data is a music scene, and the scene type of the second sound source data is one of the following scenes: listening to books, videos, games, voice calls, navigation, and learning applications.

12. A sound source data processing device, characterized in that: The device comprises: An acquisition module configured to acquire first sound source data and second sound source data; a determination module configured to determine a first sound image position of the first sound source data in a sound field space when the scene type of any one of the first sound source data and the second sound source data is a music scene; and determine a second sound image position of the second sound source data in the sound field space; wherein the first sound image position is different from the second sound image position; The output module is configured to output the first sound source data and the second sound source data according to the first sound image position and the second sound image position.

13. An electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 11 is implemented.

14. An audio system comprising the electronic device and the listening device according to claim 13; wherein: The listening device is used to listen to the sound source data played by the electronic device.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.