Audio data processing method and electronic device
By receiving and utilizing channel layout information, determining the speakers corresponding to the audio data, the problem of channel data rendering errors in the playback device is solved, and better playback effect and listener experience is achieved.
Patent Information
- Application Number
- PCT/CN2024/122337
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-02
- Filing Date
- 2024-09-29
- Publication Date
- 2025-08-07
AI Technical Summary
In the prior art, the playback device easily renders audio data from different channels onto the wrong speakers when rendering the audio data, resulting in poor playback effect and affecting the listener's listening experience.
By receiving channel layout information, determine the speakers corresponding to the audio data, ensuring that each channel data is played through the matching speakers, or mixing when the number of channel data does not match the number of speakers to obtain the correct channel layout.
Optimize the playback effect, improve the listener's listening experience, and avoid the problem that the channel data cannot be played normally or the playback effect does not meet expectations.
Smart Images

Figure CN2024122337_07082025_PF_FP_ABST
Abstract
Description
Method and electronic device for processing audio data
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on February 2, 2024, with application number 202410163384.2 and application name “Method and electronic device for processing audio data”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of terminal technology, and in particular to a method for processing audio data, an electronic device, and a computer-readable storage medium. Background Art
[0003] Audio Vivid is a 3D sound technology independently developed in my country. It was jointly developed by the Ultra High Definition World Association (UWA) and the Audio Video Coding Standard (AVS) Working Group. 3D sound creates a sense of space, direction, and presence, making it feel like sound is coming from all directions, creating a truly realistic and immersive experience.
[0004] The current Audio Vivid only contains information about the number of channels. When a playback device (such as a player) renders and plays audio data based on the channel number information, audio data of different channels may be rendered to the wrong speakers. For example, both 5.1.4 and 7.1.2 channels have 10 channels. If the audio data in the 7.1.2 channel format is rendered to the corresponding speakers as audio data in the 5.1.4 channel format for playback, the audio data originally played on the left and right surround speakers may be played from the overhead speakers, thereby affecting the playback effect and reducing the audience's listening experience.
[0005] Summary of the Invention
[0006] The present application provides a method for processing audio data, an electronic device, and a computer-readable storage medium, which can optimize playback effects and enhance the audience's listening experience.
[0007] To achieve the above objectives, this application adopts the following technical solutions:
[0008] In a first aspect, a method for processing audio data is provided, the method comprising: a first device receiving first audio data and channel layout information, the first audio data comprising M channel data, the channel layout information being used to indicate speakers corresponding to the M channel data, where M is a positive integer greater than 1; the first device determining, from at least two speakers, the speakers corresponding to the M channel data based on the channel layout information.
[0009] The above method can be executed by a first device, which can be a playback device or a module in the playback device (such as a processor, chip, or chip system, etc.). The above method can also be implemented by a logical node, logical module or software that can realize all or part of the functions of the first device. In the prior art, the speakers of the playback device playing the channel data are random. In this application, the first device determines the speakers corresponding to the M channel data based on the received channel layout information, so that each channel data can be played through its respective matching speaker (for example, the front center channel data is played through the front center speaker), thereby optimizing the playback effect.
[0010] In one possible implementation, the first device determines the speakers corresponding to the M channel data from at least two speakers based on the channel layout information, including: when the number of speakers corresponding to the M channel data matches the number of at least two speakers, the first device determines the speakers corresponding to the M channel data from at least two speakers based on the channel layout information.
[0011] When the number of speakers corresponding to the M channel data matches the number of at least two speakers, the first device can directly send the M channel data to the corresponding speakers for playback without performing mixing processing, thereby achieving a good playback effect.
[0012] In one possible implementation, the first device determines the speakers corresponding to the M channel data from at least two speakers based on the channel layout information, including: when the channel layout information matches the speaker layout information, the first device determines the speakers corresponding to the M channel data from at least two speakers based on the channel layout information, and the speaker layout information is the position distribution information of each speaker in the at least two speakers.
[0013] When the number of speakers corresponding to the M channel data matches the number of at least two speakers, and when the channel layout information matches the speaker layout information, the first device can directly render the M channel data to the corresponding speakers according to the speaker layout for playback to obtain better playback effects.
[0014] In a possible implementation, the metadata includes extended information, and the extended information includes channel layout information.
[0015] In some scenarios, the channel layout information is carried through extended information in the metadata without the need to redesign a new data storage model, which makes the storage simple and convenient.
[0016] In one possible implementation, the first device determines the speakers corresponding to M channel data from at least two speakers based on the channel layout information, including: when the number of speakers corresponding to the M channel data does not match the number of at least two speakers, the first device mixes the M channel data according to the channel layout information, the speaker layout information and the preset mapping relationship to obtain N channel data and the channel layout of the N channel data, the speaker layout information is the position distribution information of each speaker in the at least two speakers, the preset mapping relationship is used to map the number of channels indicated by the channel layout information to the number of channels indicated by the speaker layout information, N matches the number of at least two speakers, the channel layout information of the N channel data is consistent with the speaker layout information, and N is a positive integer greater than 1; the first device determines the speakers corresponding to the N channel data from at least two speakers based on the channel layout of the N channel data.
[0017] In some cases, when the number of speakers corresponding to the M channel data does not match the number of at least two speakers, the first device first mixes the M channel data according to the channel layout information, the speaker layout information and the preset mapping relationship to obtain N channel data that matches the number of at least two speakers; and then accurately renders the N channel data to each speaker of the at least two speakers according to the channel layout of the N channel data, instead of randomly sending the M channel data to the at least two speakers. This not only avoids the problem of channel data not being able to be played normally when the number of channel data does not match the number of speakers, but also obtains better playback effects.
[0018] In a possible implementation, the metadata includes extended information, where the extended information includes channel layout information and a preset mapping relationship.
[0019] The channel layout information and preset mapping relationship are carried by the extended information in the metadata without the need to redesign a new data storage model, which makes the implementation simple and convenient for storage.
[0020] In a possible implementation, the extended information includes first indication information, second indication information, first information, and second information, wherein the first indication information is used to indicate that the first information includes channel layout information, and the second indication information is used to indicate that the second information includes a preset mapping relationship.
[0021] By parsing the extended information, the first device can quickly and accurately locate the information location of the channel layout information and the preset mapping relationship according to the first indication information and the second indication information (for example, the channel layout information is in the first information, and the preset mapping relationship is in the second information), which is conducive to improving the processing efficiency of the device.
[0022] In one possible implementation, the preset mapping relationship includes: an upmix mapping relationship and a downmix mapping relationship, the upmix mapping relationship is used to map the number of channels indicated by the channel layout information to the number of channels indicated by the speaker layout information through an upmix rule, and the downmix mapping relationship is used to map the number of channels indicated by the channel layout information to the number of channels indicated by the speaker layout information through a downmix rule, the second information includes third indication information, fourth indication information, third information and fourth information, the third indication information is used to indicate that the third information includes an upmix mapping relationship, and the fourth indication information is used to indicate that the fourth information includes a downmix mapping relationship.
[0023] By parsing the second information, the first device can quickly and accurately locate the information locations of the upmix mapping relationship and the downmix mapping relationship according to the third indication information and the fourth indication information (for example, the upmix mapping relationship is in the third information, and the downmix mapping relationship is in the fourth information), which is conducive to further improving the processing efficiency of the device.
[0024] In a possible implementation, after the first device determines the speakers corresponding to the N channel data, the method further includes: the first device configuring a correspondence between the N channel data and at least two speakers according to channel layout information of the N channel data.
[0025] In the case where the number of channel data does not match the number of speakers, the first device, after determining the speakers corresponding to the N channel data, configures the correspondence between the N channel data and each speaker based on the channel layout information of the N channel data, to avoid incorrectly using the channel layout information corresponding to the original M channel data to configure the correspondence between the N channel data and each speaker, resulting in the playback effect not meeting expectations or being unable to be played.
[0026] In a second aspect, a method for processing audio data is provided, the method including: a second device generates channel layout information, the channel layout information is used to indicate the speakers corresponding to M channel data in the first audio data, M is a positive integer greater than 1; the second device sends the first audio data and the channel layout information to the first device, and the first device is used to play the first audio data according to the channel layout information.
[0027] The above method can be executed by the second device (for example, a recording device), or by a module (such as a processor, a chip, or a chip system, etc.) applied to the second device, or by a logic node, a logic module or software that can realize all or part of the functions of the second device. In the existing audio data production scheme, when producing audio data, only the number of channels is recorded without the channel layout information, which may cause the channel data to not correspond to the speaker when the playback end plays the channel data; in the present application, when the second device produces audio data, it not only records the number of channels (or the total number of channels), but also generates channel layout information corresponding to different channel data, so that the playback end can send different channel data to the corresponding speaker for playback according to the channel layout information, thereby optimizing the playback effect and improving the audience's listening experience.
[0028] In a possible implementation, the metadata includes extended information, and the extended information includes channel layout information.
[0029] The channel layout information is carried by adding extended information to the metadata without the need to redesign a new data storage model, which makes the implementation simple and convenient for storage.
[0030] In a third aspect, a method for determining a channel layout is provided, the method comprising: acquiring second audio data, the second audio data comprising at least two channel data; and determining the channel layout of the second audio data through a detection algorithm.
[0031] The above method can be executed by the first device, which can be a playback device, or by a module (such as a processor, chip, or chip system, etc.) applied to the playback device, or by a logic node, logic module or software that can realize all or part of the functions of the first device. Since the produced audio data (for example, the second audio data) may have a situation where the channel layout information is lost or the channel layout is non-standard; therefore, in this application, the first device performs a channel layout detection on the second audio data through a detection algorithm, and determines the channel layout of the second audio data, and sends at least two channel data in the second audio data to the corresponding speakers for playback according to the channel layout, thereby solving the problem that the audio data cannot be played normally when the channel layout information is lost or the audio data is not produced in the channel order indicated by the standard channel layout.
[0032] In a possible implementation, determining the channel layout of the second audio data by using a detection algorithm includes: determining the channel layout of the second audio data by using a detection algorithm in response to the second audio data.
[0033] In some embodiments, regardless of whether the second audio data includes channel layout information or whether the channel layout indicated by the included channel layout information is a standard channel layout, after receiving the second audio data, the first device will re-determine the channel layout of the second audio data through a detection algorithm to avoid the situation where the erroneous channel layout information carried by the second audio data affects the playback effect.
[0034] In one possible implementation, determining the channel layout of the second audio data through a detection algorithm includes: when it is determined through the detection algorithm that the first channel data meets a first condition, determining that the first channel data is bass channel data, wherein the first condition includes: an energy proportion of the low-frequency signal in the first channel data is greater than an energy proportion of the low-frequency signal of other channel data in the second audio data except the first channel data, and an energy proportion of the high-frequency signal in the first channel data is less than an energy proportion of the high-frequency signal of other channel data in the second audio data except the first channel data, and the first channel data is any one of the at least two channel data.
[0035] By detecting each channel data in the second audio data through a detection algorithm, it is determined that the first channel data that meets the first condition is bass channel data. The electronic device can render the bass channel data to the bass speaker for playback, thereby avoiding the situation where the bass channel data is sent to speakers in other positions (for example, overhead speakers) and cannot be played or the playback effect does not meet expectations.
[0036] In one possible implementation, determining the channel layout of the second audio data through a detection algorithm further includes: when the second channel data is determined by the detection algorithm to meet a second condition, determining that the second channel data is sky channel data, wherein the second condition includes: the duration of the audio signal in the second channel data is less than a first preset value, and the energy of the audio signal in the second channel data is less than an energy threshold, and the second channel data is any one of at least two channel data.
[0037] By detecting each channel data in the second audio data through a detection algorithm, it is determined that the second channel data that meets the second condition is the sky channel data. The electronic device can render the sky channel data to the sky speaker (for example, the left overhead speaker) for playback to avoid the situation where the sky channel data is sent to speakers in other locations (for example, a subwoofer), resulting in inability to play or the playback effect not meeting expectations.
[0038] In one possible implementation, determining the channel layout of the second audio data through a detection algorithm further includes: when it is determined through the detection algorithm that the third channel data does not satisfy the third condition and does not satisfy the fourth condition, determining that the third channel data belongs to surround channel data, and the third channel data is one of at least two channel data, the third condition includes: the energy proportion of the low-frequency signal in the third channel data is greater than the energy proportion of the low-frequency signal of the other channel data in the second audio data except the third channel data, and the energy proportion of the high-frequency signal in the third channel data is less than the energy proportion of the high-frequency signal of the other channel data in the second audio data except the third channel data, and the fourth condition includes: the duration of the audio signal in the third channel data is less than a first preset value, and the energy of the audio signal in the third channel data is less than an energy threshold.
[0039] The detection algorithm is used to detect the channel data in the second audio data. If the electronic device determines that the third channel data does not meet the first and second conditions, it can be determined that the third channel data is surround channel data; the electronic device can render the surround channel data to the surround speakers (for example, the left front speaker) for playback to avoid the situation where the surround channel data is sent to speakers in other positions (for example, a subwoofer), resulting in a situation where the playback cannot be performed or the playback effect does not meet expectations.
[0040] In one possible implementation, the surround channel data includes front channel data and rear channel data, and determining that the third channel data belongs to the surround channel data includes: when the third channel data meets a fifth condition, determining that the third channel data belongs to the front channel data, or when the third channel data does not meet the fifth condition, determining that the third channel data belongs to the rear channel data, wherein the fifth condition includes: an energy value of the audio signal in the third channel data is greater than a second preset value.
[0041] The fifth condition is used to further determine from the surround channel data whether the third channel data belongs to the front channel data or the rear channel data, so as to accurately send the third channel data to the corresponding front speaker (for example, FL speaker or FR speaker) or rear speaker (for example, BL speaker or BR speaker) to ensure a good playback effect.
[0042] In one possible implementation, the front channel data includes fourth channel data, the rear channel data includes fifth channel data, the fourth channel data and the fifth channel data are two different channel data among the at least two channel data, and the method further includes: when the absolute value of the correlation coefficient between the fourth channel data and the fifth channel data is greater than or equal to a coefficient threshold, the fourth channel data and the fifth channel data are left channel data, or the fourth channel data and the fifth channel data are right channel data.
[0043] The first device determines whether the fourth channel data and the fifth channel data are same-side channel data by calculating the correlation coefficient between the fourth channel data and the fifth channel data, so as to further determine the speakers corresponding to each channel data in the front channel data and the rear channel data, thereby avoiding the situation where the channel data corresponds to the wrong speaker and affects the playback effect.
[0044] In a fourth aspect, an embodiment of the present application provides an audio processing system, which includes a first device, a second device and at least two speakers, wherein the first device is used to execute the method described in the first aspect and various possible implementations of the first aspect, and the second device is used to execute the method described in the second aspect and various possible implementations of the second aspect.
[0045] In a fifth aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory, the memory being used to store computer programs, and the processor being used to call and run the computer programs from the memory, so that the electronic device executes the method described in the first aspect and various possible implementations of the first aspect.
[0046] In the sixth aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory, the memory being used to store computer programs, and the processor being used to call and run the computer programs from the memory, so that the electronic device executes the method described in the second aspect and various possible implementations of the second aspect.
[0047] In the seventh aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory, the memory being used to store computer programs, and the processor being used to call and run the computer programs from the memory, so that the electronic device executes the method described in the third aspect and various possible implementations of the third aspect.
[0048] In an eighth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the processor executes the method described in the first aspect and various possible implementations of the first aspect.
[0049] In the ninth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the processor executes the method described in the second aspect and various possible implementations of the second aspect.
[0050] In the tenth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the processor executes the method described in the third aspect and various possible implementations of the third aspect.
[0051] In the eleventh aspect, an embodiment of the present application provides a computer program product, which includes: computer program code, which, when executed by an electronic device, enables the electronic device to execute the method described in the first aspect and various possible implementations of the first aspect.
[0052] In the twelfth aspect, an embodiment of the present application provides a computer program product, which includes: computer program code, which, when executed by an electronic device, enables the electronic device to execute the method described in the second aspect and various possible implementations of the second aspect.
[0053] In the thirteenth aspect, an embodiment of the present application provides a computer program product, which includes: computer program code, which, when executed by an electronic device, enables the electronic device to execute the method described in the third aspect and various possible implementations of the third aspect.
[0054] In the fourteenth aspect, an embodiment of the present application provides a chip system, which includes a processing circuit and a storage medium, in which computer program instructions are stored; when the computer program instructions are executed by the processing circuit, the method described in the first aspect and various possible implementations of the first aspect is implemented.
[0055] In the fifteenth aspect, an embodiment of the present application provides a chip system, which includes a processing circuit and a storage medium, in which computer program instructions are stored; when the computer program instructions are executed by the processing circuit, the method described in the second aspect and various possible implementations of the second aspect is implemented.
[0056] In the sixteenth aspect, an embodiment of the present application provides a chip system, which includes a processing circuit and a storage medium, in which computer program instructions are stored; when the computer program instructions are executed by the processing circuit, the method described in the third aspect and various possible implementations of the third aspect is implemented.
[0057] Optionally, the processing circuit in the above chip system can be replaced by a processor, and the storage medium can be replaced by a memory. Optionally, the chip system can also include a communication interface, which is used to realize communication between the chip system and external devices.
[0058] It should be noted that the beneficial effects of the technical solutions of the fourth to sixteenth aspects of this application can refer to the beneficial effects of the technical solutions of the first, second or third aspects mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] FIG1 is a schematic structural diagram of an electronic device 100 provided in an embodiment of the present application;
[0060] FIG2 is a schematic diagram of a software structure of an electronic device 100 provided in an embodiment of the present application;
[0061] FIG3 is a schematic diagram of a system architecture 300 applicable to the present application provided in an embodiment of the present application;
[0062] FIG4 is a schematic diagram of another system architecture applicable to the present application provided in an embodiment of the present application;
[0063] FIG5 is a flow chart of a method 500 for processing audio data provided in an embodiment of the present application;
[0064] FIG6 is a flow chart of another method 600 for processing audio data provided in an embodiment of the present application;
[0065] FIG7 is a schematic diagram illustrating a speaker layout in a home scenario provided by an embodiment of the present application;
[0066] FIG8 is a flow chart of a method 800 for determining a channel layout according to an embodiment of the present application;
[0067] FIG9 is a schematic diagram of a waveform display interface of an audio software provided in an embodiment of the present application;
[0068] FIG10 is a schematic diagram of a speaker layout in a cinema scenario provided by an embodiment of the present application;
[0069] FIG11 is a schematic diagram of a detection process for determining a channel layout according to an embodiment of the present application;
[0070] FIG12 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0071] The technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.
[0072] FIG1 shows a schematic structural diagram of an electronic device 100 .
[0073] The electronic device 100 may include at least one of a recording device, an audio processing device, an audio workstation, a digital audio workstation (or a digital audio workstation (DAW)), a projection device, a mobile phone, a foldable electronic device, a tablet computer, a display screen, a laptop computer, an extended reality (XR), and an in-vehicle playback system. The specific type of the electronic device 100 is not particularly limited in this embodiment of the application.
[0074] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) connector 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor 180, a button 190, an indicator 191, a subscriber identification module (SIM) card interface 192, etc.
[0075] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than those in FIG. 1 , or may combine or separate certain components, or may have different component arrangements. The components in FIG. 1 may be implemented in hardware, software, or a combination of software and hardware.
[0076] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.
[0077] The processor 110 can generate an operation control signal according to the instruction operation code and the timing signal to complete the control of instruction fetching and execution.
[0078] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 may be a cache memory. This memory can store instructions or data that have been used or are frequently used by processor 110. When processor 110 needs to use the instruction or data, it can directly access it from this memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.
[0079] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface. The processor 110 may be connected to modules such as a touch sensor, an audio module, a wireless communication module, a display, and a camera through at least one of the above interfaces.
[0080] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.
[0081] The USB connector 130 is an interface that complies with USB standard specifications and can be used to connect the electronic device 100 and peripheral devices. The charging management module 140 is used to receive charging input from a charger. The charger can be a wireless charger or a wired charger. The power management module 141 is used to connect the battery 142, the charging management module 140 and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, the internal memory 121, the display screen 194 and the wireless communication module 160. In other embodiments, the power management module 141 and the charging management module 140 can also be set in the same device.
[0082] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.
[0083] The mobile communication module 150 can provide solutions for wireless communications, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. In some embodiments, at least some functional modules of the mobile communication module 150 may be provided in the same device as at least some modules of the processor 110.
[0084] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs the sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.). In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 110 and be set in the same device as the mobile communication module 150 or other functional modules.
[0085] The wireless communication module 160 can provide wireless communication solutions for the electronic device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), near field communication (NFC), etc. In some embodiments, the antenna 1 of the electronic device 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the electronic device 100 can communicate with the network and other terminal devices through wireless communication technology. The wireless communication technology may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), etc.
[0086] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, audio, video, and other files can be stored on the external memory card or transferred from the electronic device to the external memory card.
[0087] The internal memory 121 can be used to store computer executable program code, which includes instructions. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application required for at least one function (for example, a sound playback function, an image playback function, etc.), etc. The data storage area can store data created during the use of the electronic device 100 (for example, audio data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional methods or data processing of the electronic device 100 by running instructions stored in the internal memory 121, and / or instructions stored in a memory provided in the processor.
[0088] The electronic device 100 can implement audio functions such as playing audio and recording audio through the audio module 170 , the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.
[0089] The sensor 180 may include a sound sensor, a gyro sensor, an acceleration sensor, a touch sensor, an ambient light sensor, etc., which is used to convert various signals from the outside world (such as sound signals) into electrical signals or other required forms of information output.
[0090] The buttons 190 may include a power button, a volume button (e.g., a user can adjust the volume via the volume button), etc. The buttons 190 may be mechanical buttons or touch buttons. The electronic device 100 may receive key inputs and generate key signal inputs related to user settings and function control of the electronic device 100.
[0091] The indicator 191 may be an indicator light, which may be used to indicate charging status, power level changes, notifications, and the like.
[0092] Optionally, the terminal device may further include a SIM card interface 192 for connecting a SIM card. The SIM card can be connected to and separated from the electronic device 100 by inserting it into the SIM card interface 192 or removing it from the SIM card interface 192. The electronic device 100 may support one or more SIM card interfaces 192. The SIM card interface 192 may support Nano SIM cards, Micro SIM cards, SIM cards, and the like. Multiple cards may be inserted into the same SIM card interface 192 at the same time. The types of the multiple cards may be the same or different. The SIM card interface 192 may also be compatible with different types of SIM cards. The SIM card interface 192 may also be compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to implement functions such as calls and data communications. In some embodiments, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card may be embedded in the electronic device 100 and cannot be separated from the electronic device 100.
[0093] The electronic device 100 may also adopt other architectures. The embodiment of the present invention may also take the Harmony system as an example to exemplify the software architecture of the electronic device 100. It should be understood that the solution provided in this application may also be applied to other types of operating systems such as the Android operating system, the Apple operating system, and the Windows operating system.
[0094] FIG2 is a schematic diagram of the software structure of the electronic device 100 according to an embodiment of the present application.
[0095] In some implementations, the Harmony system includes four layers, from bottom to top: the kernel layer, the system basic service layer, the framework layer, and the application layer.
[0096] The Harmony system uses a multi-kernel design, optionally including the Linux kernel, the Hongmeng microkernel, and the lightweight IoT operating system kernel (Lite OS). This design allows devices with different device capabilities to select the appropriate system kernel. The kernel layer also includes a kernel abstraction layer, which provides basic kernel capabilities to other Harmony layers, such as process management, thread management, memory management, file system management, network management, and peripheral management.
[0097] The system basic service layer is the core capability set of the Harmony system, which supports the Harmony system to provide services to application services through the framework layer in the scenario of multi-device deployment. This layer optionally includes the following parts:
[0098] The system's basic capability subsystems provide fundamental capabilities for running, scheduling, and migrating distributed applications across multiple devices in the Harmony system. These subsystems comprise a distributed soft bus, distributed data and file management, distributed task scheduling, the Ark runtime, and distributed security and privacy protection. The Ark runtime provides a C / C++ / JavaScript multi-language runtime and basic system class libraries. It also provides a runtime for Java programs statically compiled using the Ark compiler (i.e., applications or frameworks developed in Java).
[0099] Basic software service subsystems: These provide common, general-purpose software services for the Harmony system. These subsystems include graphics and imaging, distributed media, distributed AI, multimodal input, mobile sensing development platform (MSDP) and device virtualization (DV), event notification, phone services, and design for X (DFX). These subsystems can be tailored to the specific functionalities of each device, tailored to the deployment environment.
[0100] The enhanced software service subsystem set (see the enhanced software section outlined in the dashed line in Figure 2) provides the Harmony system with differentiated, device-specific, capability-enhancing software services. This subsystem consists of subsystems such as tablet software, smart screen software, vehicle software, and Internet of Things (IoT) software. The enhanced software service subsystem set can be tailored to the subsystem granularity based on the deployment environment of different device form factors, and each subsystem can be tailored to the functional granularity.
[0101] Harmony driver foundation (HDF) and hardware abstraction layer (HAL): They are the foundation of the open hardware ecosystem of the Harmony system, providing hardware capability abstraction to the hardware upward and providing a development framework and operating environment for various peripheral drivers downward.
[0102] Hardware Service Subsystem Set: This provides common, adaptive hardware services for the Harmony system and consists of hardware service subsystems such as general sensors, location, power, USB, and biometrics. The hardware service subsystem set can be tailored to the deployment environment of different device form factors, and each subsystem can be tailored to the functional granularity.
[0103] The proprietary hardware service subsystem (see the proprietary hardware section outlined by the dashed line in Figure 2) provides the Harmony system with differentiated hardware services for different devices. This subsystem optionally includes tablet proprietary hardware services, vehicle hardware services, wearable hardware services, and IoT hardware services. The proprietary hardware service subsystem can be tailored to the subsystem granularity, and each subsystem can be tailored to the functional granularity.
[0104] The framework layer provides Harmony system applications with a user program framework and meta-capability framework in multiple languages, such as Java / C / C++ / JavaScript, as well as a multi-language framework application programming interface (API) that is open to various software and hardware services.
[0105] The application layer includes system applications and third-party applications (or extended applications), including recording applications, mixing applications, audio applications, cameras, galleries, calendars, calls, navigation, Wi-Fi, music, and videos. Applications in the Harmony system are built based on atomic capabilities (AA) and feature capabilities (FA).
[0106] The following uses a digital audio working device (e.g., a recording device) and a playback device (e.g., a projector, etc.) having the structures shown in Figures 1 and 2 as an example to introduce the application scenarios and system architecture applicable to this application.
[0107] FIG3 illustrates a schematic diagram of a system architecture 300 applicable to the present application. As shown in FIG3(a), the system architecture 300 includes at least one DAW 301, at least one playback device 302, and at least two speakers 303. The DAW 301 generates and stores audio multimedia data (e.g., audio data, audio data from a video, etc.). The playback device 302 receives and plays the audio multimedia data. For example, the playback device 302 sends an audio data acquisition request to the DAW 301 via a network. The DAW 301 responds to the audio data acquisition request by sending the audio data to the playback device 302 via the network. Upon receiving the audio data, the playback device 302 decodes the audio data to obtain multi-channel data, and then renders the multi-channel data to the corresponding speakers 303 for playback. The speakers 303 receive and play the multi-channel data sent by the playback device 302.
[0108] In one example, as shown in (b) of FIG3 , the digital audio working device 301 includes physical units such as a recording module (e.g., a microphone), an audio processing module, an encoding module, a digital effects processing module, a processor, a memory, and a hard disk. The recording module can be used to record and play sound. For example, the recording module can play data of multiple audio tracks simultaneously. During playback, the user can hear the sound and see the audio waveform displayed by the audio software running on the recording module. The audio processing module is used for audio editing to achieve audio movement, segmentation, copying, conversion, and output; the digital effects processing module is used to tune, equalize, level, reverberate, delay, and noise reduction of the audio; and the encoding module is used to encode and package the audio according to a specified format to generate audio multimedia data (e.g., audio data, etc.). The playback device 302 includes a detection algorithm module, a decoding module, and a rendering module, wherein the detection algorithm module is used to detect and verify the channel layout information of the audio multimedia data. For example, when the audio data received by the playback device 302 does not have channel layout information, the playback device 302 will trigger the detection algorithm module to determine the channel layout of the audio data. Please refer to the detailed description of the relevant embodiments below, which will not be repeated here; the decoding module is used to parse the audio data (for example, multi-channel audio data or multi-channel digital audio data) in the audio multimedia data; the rendering module is used to control the audio data (for example, multi-channel digital audio data) to be rendered and played on the corresponding speaker. The speaker 303 includes a digital-to-analog converter (DAC) module, and the speaker 303 converts the received digital audio data (for example, multi-channel digital audio data) into analog audio data (for example, multi-channel analog-to-digital audio data) through the DAC module before playing it.
[0109] It should be noted that the above system architecture 300 can be applied to audio playback scenarios and video playback scenarios; in different application scenarios, the audio multimedia data played can be audio data or audio data in a video. In this application, audio multimedia data can be described as an audio multimedia file instead, and audio data can be described as an audio file (for example, an audio file in MP3 or MP4 format). Considering the convenience of description, this application uses audio multimedia data and audio data for relevant descriptions.
[0110] FIG4 shows a schematic diagram of a process from generating to playing audio data. The process includes:
[0111] (1) Signal Acquisition. A digital audio working device (e.g., digital audio working device 301 in FIG3 ) acquires sound object signals, multi-channel signals, and higher-order ambisonics (HOA) signals through an audio acquisition module (e.g., a recording module).
[0112] (2) Preprocessing: The digital audio working equipment preprocesses the collected object signal, multi-channel signal and HOA signal (for example, filtering, etc.) to obtain preprocessed audio data.
[0113] (3) Transcoding and packaging: The digital audio working equipment transcodes and packages the pre-processed audio data to obtain packaged audio data.
[0114] (4) Encoding and generating audio multimedia data: The digital audio processing device encodes the packaged audio data and generates audio multimedia data (eg, first audio data or second audio data).
[0115] (5) Decoding. The playback device (e.g., mobile phone, tablet, large screen) obtains audio multimedia data (e.g., first audio data) from the audio workstation via the network and decodes the obtained audio multimedia data or stream.
[0116] (6) Obtaining multi-channel data. The playback device decodes and parses the corresponding multi-channel data, such as the number of channels, channel layout, and / or mixing rules.
[0117] (7) Rendering and playback. The playback device determines the speakers corresponding to the multi-channel data based on the channel layout information and sends the multi-channel data to the rendering module of the playback device; the rendering module renders the multi-channel data to the corresponding speakers for playback.
[0118] The above describes the system architecture and other information applicable to this application. Below, taking the electronic device 100 having the structure shown in Figures 1 and 2 as an example, in conjunction with the accompanying drawings and application scenarios, the method for processing audio data provided by the embodiment of this application is specifically described. It should be understood that the electronic device applicable to this method is not limited to the hardware and software structures shown in Figures 1 and 2. In actual applications, the hardware and software system architecture shown in Figures 1 and 2 can be modified according to the specific application scenario, and this application does not limit this.
[0119] As explained in the background technology, current audio data (e.g., songs) in the Audio Vivid format typically only contain information about the number of channels. When playback devices render and play audio data based on this information, they can render audio data from different channels to the wrong speakers, impacting playback quality and reducing the listener's listening experience. This application proposes a method for processing audio data that optimizes playback quality and enhances the listener's listening experience.
[0120] Before introducing the method for processing audio data provided in this application, a brief description is first given of the execution subject involved in the method (for example, the first device or the second device).
[0121] As shown in Figure 5, a flow chart of a method 500 for processing audio data is shown; the execution subject involved in this method 500 can be a first device, or a chip, chip system, or processor applied to the first device, or a logic module or software that can realize all or part of the functions of the first device.
[0122] Exemplarily, the first device may be the electronic device 100 having the structure shown in FIG. 1 and FIG. 2 or the playback device 302 in FIG. 3 ; and the electronic device 100 or the playback device 302 may be a mobile phone, a tablet computer, or a car playback system.
[0123] Figure 6 shows a flow chart of another method 600 for processing audio data; the execution subject involved in this method 600 can be a second device, or a chip, chip system, or processor applied to the second device, or a logic module or software that can realize all or part of the functions of the second device.
[0124] Exemplarily, the second device may be the electronic device 100 having the structure shown in FIG. 1 and FIG. 2 or the digital audio working device 301 in FIG. 3 ; and the electronic device 100 or the digital audio working device 301 may be an audio processing device or an audio acquisition device, etc.
[0125] The following embodiments are described by taking the execution subject of method 500 as the first device and the execution subject of method 600 as the second device as an example.
[0126] Before introducing method 500, method 600 is first described in detail. Method 600 includes S601 to S602:
[0127] S601: The second device generates channel layout information, where the channel layout information is used to indicate speakers corresponding to M channel data in the first audio data, where M is a positive integer greater than 1.
[0128] The second device is used to generate and save audio multimedia data, which may be audio data or audio data in a video, and this application does not limit this.
[0129] It should be noted that audio data includes mono audio data and multi-channel audio data; in some cases, mono audio data can be described as mono digital audio data; multi-channel audio data can be described as multi-channel digital audio data, which is not limited in this application.
[0130] The above-mentioned second device and the device at the playback end (for example, the first device) can transmit data to each other through the network. For example, the first device at the playback end sends an audio multimedia data acquisition request to the second device, and the second device responds to the acquisition request and sends audio multimedia data to the first device.
[0131] For another example, the first device at the playback end sends a first audio data acquisition request to the second device, and the second device responds to the first audio data acquisition request and sends the first audio data to the first device.
[0132] The first audio data may refer to multi-channel audio data generated by the second device after mixing or performing other processing on the original audio data. The first audio data may refer to pure audio data or audio data in a video. For example, the first audio data may be audio data of a song or audio data of a movie.
[0133] The M channel data mentioned above may refer to audio data of M channels, where M may be 2, 3, or 5, for example.
[0134] Since different channel data correspond to different sound propagation directions, in order to obtain a good playback effect, different channel data needs to be rendered to the speakers in the corresponding directions for playback. The channel layout information can be used to indicate the speakers corresponding to the M channel data. In other words, the channel layout information is used to indicate the channel layout of the M channel data; wherein the channel layout includes but is not limited to stereo channels (i.e., 2 channels), surround channels (i.e., 4 channels), 4.1 channels, 5.0 channels, 5.1 channels, 5.1.4 channels, 6.0 channels, 7.1 channels, 7.1.2 channels, and 9.1.4 channels.
[0135] It should be noted that the channel layout may also be described alternatively as an audio channel layout or a sound bed layout or a multi-channel sound bed layout or a channel configuration.
[0136] For example, M=2 channel data includes left channel data and right channel data, and the channel layout information is stereo channel information, wherein the stereo channel information indicates that the speaker corresponding to the left channel data is the left speaker, and the speaker corresponding to the right channel data is the right speaker.
[0137] For another example, M=6 channel data includes left front channel data, right front channel data, center channel data, bass channel data, left channel data and right channel data, and the channel layout information is 5.1 channel information, wherein the 5.1 channel information indicates that the speaker corresponding to the left front channel data is the left front speaker, the speaker corresponding to the right front channel data is the right front speaker, the speaker corresponding to the center channel data is the center speaker, the speaker corresponding to the bass channel data is a subwoofer (or woofer), the speaker corresponding to the left channel data is the left speaker and the speaker corresponding to the right channel data is the right speaker.
[0138] It should be noted that, in some scenarios, the bass channel data may also be replaced by any one of the descriptions of heavy bass channel data, subwoofer data, or super bass channel data.
[0139] When producing the first audio data, the second device can carry the channel number information (for example, M=10) and channel layout information corresponding to the first audio data; when the first device sends a first audio data acquisition request to the second device, the second device responds to the request and sends the first audio data carrying the channel number information and channel layout information to the first device; the first device parses the first audio data to obtain the channel number information and channel layout information corresponding to the first audio data.
[0140] It should be noted that when the second device generates (or produces) the first audio data, it can store data information related to the first audio data through the audio model defined by the standard, where the standard can be the ITU-R BS.2076 series standard, such as ITU-R BS.2076-1 or ITU-R BS.2076-2.
[0141] For example, taking the number of channels as 10 and the channel layout as 7.1.2 channels, when the second device generates the first audio data, it can store the first audio data and related channel number information and channel layout information in the data format of the audio model shown in Table 1.
[0142] Table 1
[0143] The meaning of each field in Table 1 is as follows:
[0144] Frame: The unit for reading memory. The number of bytes corresponding to a frame is "number of channels * bit width". For example, 7.1.2 audio has 10 channels, and the corresponding bit width for 16-bit audio is 2 bytes. 1 frame = 10 * 2 bytes, where "*" is the multiplication operator.
[0145] Frame header (magic): The size is 1frame and is used to mark the frame header. For example, the first two bytes of magic are 0x3d and 0xaf respectively. Except for the first two bytes, the subsequent bytes of magic are filled with 0xff.
[0146] Size 1 (size1): The size is 1 frame. The first two bytes of size1 are filled or read in CPU byte order and are used to indicate the number of bytes occupied by the metadata following the frame. Except for the first two bytes, the subsequent bytes of size1 are filled with 0xff.
[0147] Metadata: used to store audio-related data (such as the number of channels); the metadata consists of two parts: one part is basic information, also known as the basic metadata part (audio format extended), and the other part is extended information, also known as the extended metadata part (VRext). The basic metadata part is compatible with the audio model defined in the ITU-R BS.2076 standard. For example, the basic metadata part can reference the metadata part defined in the ITU-R BS.2076-2 standard); the extended metadata part is a data part extended by adding new fields on the basis of the basic metadata; the extended metadata part can be used to store (or carry) channel layout information and / or preset mapping relationships.
[0148] The preset mapping relationship includes an upmix mapping relationship and a downmix mapping relationship, wherein the upmix mapping relationship is used to map the number of channels (or channel layout) indicated by the channel layout information to the number of channels (or speaker layout) indicated by the speaker layout information through an upmix rule; the downmix mapping relationship is used to map the number of channels (or channel layout) indicated by the channel layout information to the number of channels (speaker layout) indicated by the speaker layout information through a downmix rule; wherein the upmix rule and the downmix rule can be collectively referred to as mixing rules.
[0149] It should be noted that the mixing rules can also be described as mixing algorithms or mixing methods instead. The mixing rules include upmixing rules (or upmixing algorithms) and downmixing rules (or downmixing algorithms). The upmixing rules include but are not limited to a summation method, a weighted average method, an adaptive mixing weighting method (or an attenuation factor method or a normalization algorithm or an improved normalization algorithm), and an automatic alignment algorithm; the downmixing rules include but are not limited to a summation method, a weighted average method, a weighted clamping method (i.e., taking the maximum value instead when overflow occurs), an adaptive mixing weighting method (or an attenuation factor method or a normalization algorithm or an improved normalization algorithm), and an automatic alignment algorithm.
[0150] In some embodiments, the channel layout information (or the preset mapping relationship) may be carried by the extended metadata portion in the metadata. In other words, the extended information of the metadata includes the channel layout information (or the preset mapping relationship).
[0151] For example, if the number of channels is 10 and the channel layout is 7.1.2 channels, the extended metadata portion may be used to carry information about the number of channels (eg, 10) and the channel layout (eg, 7.1.2 channels) (or a preset mapping relationship).
[0152] In some scenarios, the channel layout information is carried through extended information in the metadata without the need to redesign a new data storage model, which makes the storage simple and convenient.
[0153] In other embodiments, the extended information (ie, the extended metadata portion) includes channel layout information and preset mapping relationships.
[0154] For example, the number of channels is 10, the channel layout is 5.1.4 channels, and the preset mapping relationship is X; the extended information may carry the number of channels 10, 7.1.2 channels, and the preset mapping relationship.
[0155] For example, the second device can generate the first audio data according to the data format shown in Table 1, wherein the extended information (i.e., the extended metadata part) is used to carry the number of channels, channel layout information and preset mapping relationship of the first audio data, and the basic information (i.e., the basic metadata part) can be used to carry the first audio data and other information (for example, the spherical harmonic coding order used by the renderer on the playback end).
[0156] It can be seen that in some scenarios, the channel layout information and preset mapping relationship can be carried by the extended information in the metadata without the need to redesign a new data storage model, which makes the implementation simple and convenient for storage.
[0157] In some embodiments, the extended information (extended metadata portion) includes indication information 1 and indication information 2; indication information 1 is used to indicate whether the metadata includes the extended metadata portion. For example, indication information 1 is indicated by 1 bit. When indication information 1 is 0, it indicates that the metadata includes basic metadata but not extended metadata. When indication information 1 is 1, it indicates that the metadata includes basic metadata and extended metadata. Indication information 2 is used to indicate whether the extended information includes channel layout information. For example, indication information 2 is indicated by 1 bit. When indication information 2 is 0, it indicates that the extended information does not include channel layout information. When indication information 2 is 1, it indicates that the extended information includes channel layout information.
[0158] For example, when indication information 2 indicates that the extended information includes channel layout information, some fields of the extended information may be used to carry the channel layout information. For example, the channel layout information that may be carried by some fields includes, but is not limited to, the following channel layouts:
[0159] Mono (CHANNEL_OUT_MONO), Stereo (CHANNEL_OUT_STEREO), 3-channel (CHANNEL_OUT_THREE), 4-channel (CHANNEL_OUT_QUAD), 5-channel (CHANNEL_OUT_FIVE), Surround channel (CHANNEL_OUT_SURROUND), 5.1 channel (CHANNEL_OUT_5POINT1), 5.1.4 channel (CHANNEL_OUT_5POINT1POINT4), 7-channel (CHANNEL_OUT_SEVEN), 7.1 channel (CHANNEL_OUT_7POINT1), 7.1 surround channel (CHANNEL_OUT_7POINT1_SURROUND), 9-channel (CHANNEL_OUT_NINE), 7.1.2-channel (CHANNEL_OUT_7POINT1POINT2), 11-channel (CHANNEL_OUT_ELEVEN), 7.1.4-channel (CHANNEL_OUT_7POINT1POINT4), 13-channel (CHANNEL_OUT_13POINT_360RA), 14-channel (CHANNEL_OUT_FOURTEEN), 15-channel (CHANNEL_OUT_FIFTEEN) and 16-channel (CHANNEL_OUT_SIXTEEN).
[0160] In other embodiments, the extended information (i.e., the extended metadata part) includes first indication information, second indication information, first information, and second information, wherein the first indication information is used to indicate that the first information includes channel layout information, and the second indication information is used to indicate that the second information includes a preset mapping relationship. The first indication information may refer to the address information of the first information; the second indication information may refer to the address information of the second information. For example, the first indication information is indicated by 8 bits (for example, 00000001), and the first information includes the channel layout information. The second indication information is also indicated by 8 bits (for example, 00000011), and the second information includes the preset mapping relationship. The extended information of the metadata indicates the channel layout information and the preset mapping relationship through the first indication information and the second indication information, respectively, so that the first device can parse the extended information and quickly and accurately locate the information location of the channel layout information and the preset mapping relationship according to the first indication information and the second indication information (for example, the channel layout information is in the first information, and the preset mapping relationship is in the second information), which is conducive to improving the processing efficiency of the device.
[0161] In some further embodiments, the preset mapping relationship includes: an upmix mapping relationship and a downmix mapping relationship, the second information includes third indication information, fourth indication information, third information and fourth information, the third indication information is used to indicate that the third information includes an upmix mapping relationship, and the fourth indication information is used to indicate that the fourth information includes a downmix mapping relationship; the third indication information may refer to address information of the third information; and the fourth indication information may refer to address information of the fourth information.
[0162] For example, the third indication information is indicated by 8 bits (for example, 00000101), and the third information includes an upmix mapping relationship. The fourth indication information is also indicated by 8 bits (for example, 00011011), and the fourth information includes a downmix mapping relationship. The extended information of the metadata indicates the upmix mapping relationship and the downmix mapping relationship through the third indication information and the fourth indication information, respectively, so that the first device can parse the extended information and quickly and accurately locate the information locations of the upmix mapping relationship and the downmix mapping relationship according to the third indication information and the fourth indication information (for example, the upmix mapping relationship is in the third information, and the downmix mapping relationship is in the fourth information), which is conducive to further improving the processing efficiency of the device.
[0163] For example, when the third indication information indicates that the third information includes an upmix mapping relationship, the extended metadata portion is used to carry upmixing rules; for example, some fields of the extended metadata portion can be used to carry upmixing rules, and the upmixing rules that can be carried by these fields include, but are not limited to, at least one of a summation method, a weighted average method, an adaptive mixing weighting method (or an attenuation factor method, a normalization algorithm, or an improved normalization algorithm), and an automatic alignment algorithm; when the third indication information indicates that the fourth information includes a downmix mapping relationship, the extended metadata portion can also be used to carry downmixing rules; for example, another part of the fields of the extended metadata portion is used to carry downmixing rules, and the downmixing rules that can be carried by these parts of the fields include, but are not limited to, at least one of a summation method, a weighted average method, a weighted clamping method (i.e., taking the maximum value as a replacement in case of overflow), an adaptive mixing weighting method (or an attenuation factor method, a normalization algorithm, or an improved normalization algorithm), and an automatic alignment algorithm.
[0164] Alignment: used to pad the metadata data length to an integer multiple of the frame length.
[0165] For example, if the size of metadata is 19824 bytes and 1 frame is 20 bytes, 16 bytes need to be padded with 0xff, so that the total length of metadata and align is 19840 bytes.
[0166] Size 2 (size2): The size is 1 frame. The first two bytes of size2 are filled or read in CPU byte order and are used to indicate the number of frames of PCM data following the frame. Subsequent bytes are filled with 0xff.
[0167] Pulse code modulation (PCM) part: The size is 1024 frames. This PCM part can be used to carry audio data (for example, M channel data in the first audio data). It should be noted that for PCM channel data with different channel layouts, the segmentation in the PCM channel data is different.
[0168] For example, as shown in Table 2, the distribution of PCM channel data for different channel layouts is as follows:
[0169] Table 2
[0170] In Table 2, abbreviations such as FL and FR represent channel data at different positions, where FL (left front) represents left front, FR (right front) represents right front, LFE (low frequency effect, LFE) represents low frequency effect (or bass), FC (front center) represents front center, BL (back left) represents left back, BR (back right) represents right back, SL (surround left) represents left surround, SR (surround right) represents right surround, LTF (left top front) represents left top front, RTF (right top front) represents right top front, LTR (left top rear) represents left rear of the head, RTR (right top rear) represents right rear of the head, LTM (left top middle) represents left top (or left side of the head), and RTM (right top middle) represents right top (or right side of the head).
[0171] It should be noted that, in some scenarios, the left back can also be represented by LB (left back), the right back can also be represented by RB (right back), the left surround can also be represented by LS (left surround), and the right surround can also be represented by RS (right surround).
[0172] S602: The second device sends first audio data and channel layout information to the first device, and the first device is configured to play the first audio data according to the channel layout information.
[0173] It should be noted that the first audio data and the channel layout information can be carried by the same information or by different information.
[0174] In some embodiments, the first audio data and channel layout information can be carried using the data format shown in Table 1. For example, the PCM part carries the M channel data in the first audio data, and the extended metadata part of the metadata is used to carry the channel layout information and / or mixing rules of the first audio data.
[0175] In other embodiments, the second device may also generate the first audio data using the audio model defined by the ITU-R BS.2076 standard, and the channel layout information and / or mixing rules of the first audio data may be carried in a separate information format, which is not limited in this application.
[0176] The second device sends the first audio data and channel layout information to the first device in the form of a data packet via the network; after receiving the data packet, the first device parses the first audio data and the channel layout information, and determines the speakers corresponding to the M channel data in the first audio data based on the channel layout information; finally, the first device outputs the first audio data through the speakers corresponding to the M channel data.
[0177] To sum up, the existing audio data production scheme only records the number of channels but not the channel layout information when producing audio data, which may lead to the problem that the channel data and the speakers may not correspond when the playback end plays the channel data; in this application, when the second device produces audio data, it not only records the number of channels (or the total number of channels), but also generates channel layout information corresponding to different channel data, so that the playback end can send different channel data to the corresponding speakers for playback according to the channel layout information, thereby optimizing the playback effect and improving the audience's listening experience.
[0178] As shown in FIG5 , it is a flowchart of a method 500 for processing audio data provided in an embodiment of the present application. The method 500 includes steps S501 to S502 , and these steps are described in detail below.
[0179] S501: A first device receives first audio data and channel layout information, where the first audio data includes M channel data. The channel layout information is used to indicate speakers corresponding to the M channel data, where M is a positive integer greater than 1.
[0180] The first device and the speaker can communicate with each other via a wired or wireless method, which is not limited in this application. For example, the first device and the speaker communicate with each other via Bluetooth.
[0181] The first device is a playback device, and the first device is used to receive and play audio multimedia data; for example, the first device is used to receive and play first audio data.
[0182] It should be noted that, for the introduction of the first audio data and the channel layout information, reference may be made to the relevant description in method 600 , which will not be repeated here.
[0183] The first device receives a data packet sent by the second device through the network, where the data packet includes first audio data, channel layout information, number of channels and other information; the first device parses the data packet to obtain M channel data, channel layout information and number of channels information.
[0184] For example, the second device generates a first audio data file carrying information such as channel layout information and the number of channels according to the data format of the audio model shown in Table 1, and encodes the first audio data file to obtain a data packet; according to the above description, the basic information and extended information of the metadata in the audio model are used to store information such as the number of channels, channel layout information and / or preset mapping relationships, and PCM is used to carry M channel data. After the first device obtains the data packet from the second device, it decodes the data packet and parses M channel data from the PCM part according to the data format of the audio model shown in Table 1, and parses information such as the number of channels, channel layout information and / or preset mapping relationships from the basic information and extended information of the metadata.
[0185] S502: The first device determines, from at least two speakers according to the channel layout information, speakers corresponding to the M channel data.
[0186] Among them, the speakers are also called horns. The layout of the speakers is different in different application scenarios. The speakers can be divided into left front speakers, right front speakers, center speakers, left rear speakers, right rear speakers, subwoofers, etc. according to their placement positions.
[0187] The above-mentioned layout of at least two speakers can correspond to multiple channel layouts. For example, the layout of two speakers can correspond to a channel layout of stereo channels; for another example, the layout of four speakers can correspond to a channel layout of 3.1 channels or 4.0 channels; for another example, the layout of 10 speakers can correspond to a channel layout of 7.1.2 channels or 5.1.4 channels.
[0188] It should be noted that the layout of the speakers can generally be arranged according to the channel layout to form speaker layouts corresponding to different channel layouts.
[0189] The first device may determine, from the at least two speakers according to the channel layout information, a speaker corresponding to each channel data in the M channel data.
[0190] For example, in some scenarios such as television programs, movies, or concerts, the first audio data generated by the second device generally includes 6-channel data or 8-channel data. For example, the first audio data includes: M=6 channel data, where the 6 channel data are left front channel data, center channel data, right front channel data, left rear channel data, right rear channel data, and subwoofer channel data; the channel layout information is 5.1 channel information, and the 5.1 channel information indicates that the speaker corresponding to the left front channel data is the left front speaker, the speaker corresponding to the center channel data is the center speaker, the speaker corresponding to the right front channel data is the right front speaker, the speaker corresponding to the left rear channel data is the left rear speaker, the speaker corresponding to the right rear channel data is the right rear speaker, and the speaker corresponding to the subwoofer channel data is the subwoofer. The first device can render each channel data to the corresponding speaker for playback according to the channel layout information.
[0191] To sum up, in the prior art, the speakers used by the playback device to play channel data are random. In this application, the first device determines the speakers corresponding to M channel data based on the received channel layout information, so that each channel data can be played through its own matching speaker (for example, the front center channel data is played through the front center speaker), thereby optimizing the playback effect.
[0192] In one possible implementation, the first device determines the speakers corresponding to the M channel data from at least two speakers based on the channel layout information, including: when the number of speakers corresponding to the M channel data matches the number of at least two speakers, the first device determines the speakers corresponding to the M channel data from at least two speakers based on the channel layout information.
[0193] It should be noted that matching can be understood as being equal or identical; usually, the number of speakers corresponding to the M channel data is the same as the number of the M channel data, that is, the number of speakers corresponding to the M channel data is M.
[0194] When the first device detects that the first audio data includes M channel data, it determines that the M channel data need to correspond to M speakers; at this time, the first device will compare whether M matches the number of at least two speakers (that is, whether they are equal); when the number M of speakers corresponding to the M channel data is equal to the number of at least two speakers, it means that the first device does not need to perform mixing processing on the M channel data, but directly determines the speakers corresponding to the M channel data from at least two speakers (or determines M speakers) based on the channel layout information.
[0195] For example, the first audio data includes M=3 channel data; under normal playback conditions, M=3 channel data needs to correspond to M=3 speakers; taking the channel layout of 2.1 channels as an example, the 2.1 channels include a left channel, a right channel and a bass channel, wherein the left channel corresponds to the left speaker, the right channel corresponds to the right speaker, and the bass channel corresponds to the bass speaker; since the channel layout corresponding to the first audio data is 2.1 channels, the three channel data included in the first audio data are left channel data, right channel data and bass channel data respectively; the first device can determine the speakers corresponding to the three channel data according to the 2.1 channels, that is, the left channel data corresponds to the left speaker, the right channel data corresponds to the right speaker, and the bass channel data corresponds to the bass speaker.
[0196] When the number of speakers corresponding to the M channel data matches the number of at least two speakers, the first device can directly send the M channel data to the corresponding speakers for playback without performing mixing processing, thereby achieving a good playback effect.
[0197] In other embodiments, the first device determines the speakers corresponding to the M channel data from at least two speakers based on the channel layout information, including: when the channel layout information matches the speaker layout information, the first device determines the speakers corresponding to the M channel data from at least two speakers based on the channel layout information, and the speaker layout information is the position distribution information of each speaker in the at least two speakers.
[0198] Among them, the speaker layout information is used to indicate different speaker layouts formed by at least two speakers; and the speaker layout can be deployed according to the channel layout. For example, if the channel layout is 2.1 channels, the speaker layout can be deployed according to the 2.1 channels to obtain a 2.1 speaker layout, wherein the 2.1 speaker layout includes a left speaker, a right speaker and a subwoofer, wherein the left speaker is used to play left channel data, the right speaker is used to play right channel data, and the subwoofer is used to play subwoofer channel data.
[0199] In some embodiments, the first device can send a speaker layout acquisition request to the speaker control device, and the speaker control device responds to the request and feeds back the speaker layout information of the current speaker to the first device; wherein, the speaker control device can be a separate hardware device, or a functional module in a certain speaker, or a chip in a certain speaker, and this application does not limit this.
[0200] It should be noted that the first device may obtain the speaker layout information before determining the number of speakers corresponding to the M channel data and the number of at least two speakers, or after determining the number of speakers corresponding to the M channel data and the number of at least two speakers. This application does not limit this.
[0201] The matching of the above-mentioned channel layout information and speaker layout information can be understood as the channel layout indicated by the channel layout information is consistent with the speaker layout indicated by the speaker layout information; for example, the channel layout indicated by the channel layout information is 5.1.4 channels, and the speaker layout indicated by the speaker layout information is a 5.1.4 speaker layout (that is, speakers deployed according to the 5.1.4 channel layout).
[0202] When the number of speakers corresponding to the M channel data matches the number of at least two speakers, the first device determines whether the channel layout information matches the speaker layout information; when the first device determines that the channel layout information matches the speaker layout information, it means that the current speaker layout is deployed according to the channel layout indicated by the channel layout information. At this time, the first device can determine the speakers corresponding to the M channel data from at least two speakers according to the channel layout information, and render each channel data to the corresponding speaker for playback.
[0203] In this embodiment, when the number of speakers corresponding to the M channel data matches the number of at least two speakers, when the channel layout information matches the speaker layout information, the first device can directly render the M channel data to the corresponding speakers according to the speaker layout for playback to obtain better playback effects.
[0204] In one possible implementation, the first device determines the speakers corresponding to M channel data from at least two speakers based on the channel layout information, including: when the number of speakers corresponding to the M channel data does not match the number of at least two speakers, the first device mixes the M channel data according to the channel layout information, the speaker layout information and the preset mapping relationship to obtain N channel data and the channel layout of the N channel data, the speaker layout information is the position distribution information of each speaker in the at least two speakers, the preset mapping relationship is used to map the number of channels indicated by the channel layout information to the number of channels indicated by the speaker layout information, N matches the number of at least two speakers, the channel layout information of the N channel data is consistent with the speaker layout information, and N is a positive integer greater than 1; the first device determines the speakers corresponding to the N channel data from at least two speakers based on the channel layout of the N channel data.
[0205] It should be noted that mismatch can be understood as unequal or different.
[0206] When the number of speakers corresponding to the M channel data does not match the number of the at least two speakers, it indicates that the number of speakers corresponding to the M channel data may be greater than the number of the at least two speakers or may be less than the number of the at least two speakers.
[0207] The preset mapping relationship can be referred to the relevant description above and will not be repeated here.
[0208] When the number of speakers corresponding to the M channel data does not match the number of at least two speakers, the first device can determine the mapping relationship to be used based on the size relationship between the number of speakers corresponding to the M channel data and the number of at least two speakers; for example, when the first device determines that the number of speakers corresponding to the M channel data is greater than the number of at least two speakers, the first device determines to use the downmix mapping relationship to mix the M channel data; when the first device determines that the number of speakers corresponding to the M channel data is less than the number of at least two speakers, the first device determines to use the upmix mapping relationship to mix the M channel data.
[0209] The number of speakers corresponding to the N channel data is consistent with (or equal to) the number of speakers indicated by the speaker layout information; the channel layout of the N channel data is the channel layout indicated by the speaker layout information.
[0210] After the first device determines a mapping relationship to be used based on the number of speakers corresponding to the M channel data and the number of the at least two speakers, the first device mixes the M channel data based on the channel layout information, the speaker layout information, and a preset mapping relationship (e.g., an upmix mapping relationship or a downmix mapping relationship) to obtain N channel data and a channel layout for the N channel data. The first device determines the speakers corresponding to the N channel data from the at least two speakers based on the channel layout of the N channel data, and renders the N channel data to the corresponding speakers for playback.
[0211] For example, the channel layout indicated by the channel layout information is 5.1 channels, and the speaker layout indicated by the speaker layout information is a 5.1.4 speaker layout. After the first device obtains the channel layout information and the speaker layout information, it determines that the mapping relationship to be used is the upmix mapping relationship based on the relationship between the number of speakers corresponding to the M=6 channel data and the number of at least two speakers 10; the first device uses the upmix mapping relationship to map the channel layout information of the M=6 channel data to the speaker layout information corresponding to the N=10 channel data (that is, the channel layout information of the N=10 channel data), that is, the first device maps the number of channels M=6 indicated by the channel layout information to the number of channels N=10 indicated by the speaker layout information through the upmix mapping relationship.
[0212] For another example, if the upmixing rule corresponding to the upmix mapping relationship is a weighted average method, the first device upmixes the 6-channel data with a 5.1-channel layout to the 10-channel data with a 7.1.2-channel layout through the upmix mapping relationship, as shown in Table 3. In Table 3, the first column on the left is the 5-channel data of the 5.1-channel layout, and the first horizontal row is the 10-channel data of the 7.1.2-channel layout. The first device uses the mapping coefficients (or proportional coefficients) in the table to upmix the 6-channel data of the 5.1-channel layout to the 10-channel data of the 7.1.2-channel layout. During the mapping process, the first device multiplies the channel data of the 5.1-channel layout by the corresponding coefficients in the table to obtain the channel data corresponding to the 7.1.2-channel layout.
[0213] For example, as shown in Table 3, the first device maps the FL channel data of the 5.1 channel to the FL channel data of the 7.1.2 channel, that is, multiplying the FL channel data of the 5.1 channel by the corresponding coefficient 1.0 to obtain the FL channel data of the 7.1.2 channel; maps the FR channel data of the 5.1 channel to the FR channel data of the 7.1.2 channel, that is, multiplying the FR channel data of the 5.1 channel by the corresponding coefficient 1.0 to obtain the FR channel data of the 7.1.2 channel; similarly, the FL channel data of the channel is mapped to the SL channel data of the 7.1.2 channel, that is, multiplying the FL channel data of the 5.1 channel by the corresponding coefficient 0.5 to obtain the SL channel data; similarly, the mapping coefficients in Table 3 can be used to calculate the other channel data of the 5.1 channel corresponding to the other channel data of the 7.1.2 channel.
[0214] Table 3
[0215] For another example, when the number of speakers corresponding to the M channel data is greater than the number of at least two speakers, the first device needs to downmix the M channel data through the downmix mapping relationship to obtain N channel data and the channel layout of N channel data; at this time, M is greater than N, and N matches the number of at least two speakers.
[0216] For example, assuming the downmixing rule corresponding to the downmix mapping relationship is a weighted average method, the first device downmixes 10 channel data with a 7.1.2 channel layout to 6 channel data with a 5.1 channel layout using the downmix mapping relationship, as shown in Table 4. The first column on the left in Table 4 is the 10 channel data for 7.1.2 channels, and the first horizontal row is the 6 channel data for 5.1 channels. The first device downmixes the 10 channel data for 7.1.2 channels to the 6 channel data for 5.1 channels using the mapping coefficients (or proportional coefficients) in the table. During mapping, the first device multiplies the channel data for 7.1.2 channels by the corresponding coefficients in the table to obtain the channel data corresponding to 5.1 channels.
[0217] For example, as shown in Table 4, the first device maps the FL channel data of the 7.1.2 channel to the FL channel data of the 5.1 channel, that is, multiplying the FL channel data of the 7.1.2 channel by the corresponding coefficient 1.0 to obtain the FL channel data of the 5.1 channel; and maps the FR channel data of the 7.1.2 channel to the FR channel data of the 5.1 channel, that is, multiplying the FR channel data of the 7.1.2 channel by the corresponding coefficient 1.0 to obtain the FR channel data of the 5.1 channel; similarly The 7.1.2-channel SL channel data is mapped to the 5.1-channel FC channel data by multiplying the 7.1.2-channel SL channel data by the corresponding coefficient 0.5. For the 5.1-channel FC channel data, the coefficient is 1.0*7.1.2-channel FC channel data + 0.5*7.1.2-channel SL channel data. Similarly, the mapping coefficients in Table 4 can be used to calculate the corresponding 7.1.2-channel other channel data to the 5.1-channel other channel data. It should be noted that "*" is a multiplication operator.
[0218] Table 4
[0219] In some embodiments, after the first device determines the speakers corresponding to the M channel data from at least two speakers based on the channel layout information, it can configure the speakers at different locations based on the channel layout information corresponding to the M channel data. In other embodiments, after the first device determines the speakers corresponding to the N channel data, it can configure the speakers at different locations based on the channel layout information corresponding to the N channel data, i.e., N channel data corresponds to N speakers at different locations, for example, configuring an FL speaker to play FL channel data, configuring an FR speaker to play FR channel data, and so on. If the number of channel data does not match the number of speakers, the first device, after determining the speakers corresponding to the N channel data, then configures the correspondence between the N channel data and the speakers based on the channel layout information of the N channel data, thereby avoiding the erroneous use of the channel layout information corresponding to the original M channel data to configure the correspondence between the N channel data and the speakers, resulting in unintended playback or even unavailable playback.
[0220] For example, taking the example of the first device configuring speakers at different positions through an XML file, when the number of channel data 6 does not match the number of speakers 10, the first device first matches the number of channel data with the number of speakers through the channel layout information, speaker layout information and preset mapping relationship, and then configures the speakers at 10 different positions according to the channel layout 7.1.2 corresponding to N=10 channel data, so that the rendering module of the first device plays the corresponding channel data according to the speakers configured in the XML file.
[0221] For example, the first device may configure speakers at different locations through an XML file; specifically, the first device may configure speakers at different locations through audio channel format names in the XML file:
[0222] <audioChannelFormat audioChannelFormatName="FL" / > Used to configure FL speakers;
[0223] <audioChannelFormat audioChannelFormatName="FR" / > Used to configure FR speakers;
[0224] <audioChannelFormat audioChannelFormatName="FC" / > Used to configure FC speakers;
[0225] <audioChannelFormat audioChannelFormatName="LEF" / > Used to configure LEF speakers;
[0226] <audioChannelFormat audioChannelFormatName="BL" / > Used to configure BL speakers;
[0227] <audioChannelFormat audioChannelFormatName="BR" / > Used to configure BR speakers;
[0228] <audioChannelFormat audioChannelFormatName="SL" / > Used to configure SL speakers;
[0229] <audioChannelFormat audioChannelFormatName="SR" / > Used to configure SR speakers;
[0230] <audioChannelFormat audioChannelFormatName="LTM" / > Used to configure LTM speakers;
[0231] <audioChannelFormat audioChannelFormatName="RTM" / > Used to configure the RTM speaker.
[0232] After the first device configures the speakers through the XML file, it sends the M channel data (or N channel data) and the XML file to the rendering module. The rendering module renders the M channel data (or N channel data) to the corresponding speakers according to the speakers configured in the XML file; after each speaker receives the corresponding channel data, it converts its own channel data into analog data before playing it.
[0233] It should be noted that when the number of speakers corresponding to the M channel data does not match the number of at least two speakers, the number of speakers corresponding to the M channel data can be matched with the number of at least two speakers through a preset mapping relationship; however, in some cases, if the first device cannot detect the preset mapping relationship, it will feedback an error message to the user instead of performing a hard match, thereby avoiding the situation where the channel data is sent to the speaker in the wrong position and affecting the playback effect.
[0234] In this embodiment, when the number of speakers corresponding to the M channel data does not match the number of at least two speakers, the first device first mixes the M channel data according to the channel layout information, the speaker layout information and the preset mapping relationship to obtain N channel data that matches the number of at least two speakers; and then accurately renders the N channel data to each speaker of the at least two speakers according to the channel layout of the N channel data, instead of randomly sending the M channel data to the at least two speakers. This not only avoids the problem of channel data not being able to be played normally when the number of channel data does not match the number of speakers, but also obtains a better playback effect.
[0235] The following describes a method for processing audio data in conjunction with an application scenario; as shown in FIG7 , a schematic diagram of a speaker layout in a home scenario is shown; after the second device performs mixing and other processing on the original audio data, it generates first audio data according to the 5.1.4 channel shown in FIG7 (a); wherein, the 5.1.4 channel includes 5 main speakers, 1 woofer, and 4 overhead speakers, namely, the left front speaker, the right front speaker, ..., the left rear overhead speaker, and the right rear overhead speaker; speakers in different positions are used to play different channel data; for example, the main speaker corresponds to the surround channel data (for example, the left front speaker, the right front speaker, etc.), The woofer corresponds to the bass channel data and the overhead speaker corresponds to the sky channel data; the second device generates first audio data carrying the number of channels, channel layout and / or mixing rules according to the data format of Table 1; the first device obtains a data packet including the first audio data from the second device, and parses the data packet according to the data format of Table 1 to obtain 10 channel data; the 10 channel data are respectively left front channel data, right front channel data, center channel data,..., right rear channel data; after the first device parses the 10 channel data and 5.1.4 channels from the data packet, it configures the speakers corresponding to the 10 channel data according to the 5.1.4 channels.
[0236] (b) in FIG7 shows a 5.1.4-channel speaker layout; four speakers for playing the sky channel data, such as the left front overhead speaker, the right front overhead speaker, the left rear overhead speaker, and the right rear overhead speaker, are installed in the upper space (e.g., the top of the room, etc.); five speakers for playing the surround channel data, such as the left front speaker and the right front speaker, are installed on a plane flush with the human ear; one speaker for playing the bass channel data, such as a subwoofer or a woofer or a subwoofer, is installed directly in front of or to the right of the listener, and the left and right speakers and the subwoofer are placed on the same line as far as possible to obtain a better playback effect. After the first device configures the speakers corresponding to the 10 channel data according to the 5.1.4 channel, it renders the 10 channel data to the speakers at the corresponding positions shown in FIG7 (b) for playback.
[0237] It should be noted that in some scenarios, the position descriptions in Figure 7 may also have different alternative descriptions. For example, "left front" can be replaced by "left front", "right front" can be replaced by "right front", "bass" can be replaced by "subwoofer", "left front overhead" can be replaced by "left front overhead", "right front overhead" can be replaced by "right front overhead", "left rear overhead" can be replaced by "left rear overhead", "right rear overhead" can be replaced by "right rear overhead", "left rear" can be replaced by "left rear" and "right rear" can be replaced by "right rear".
[0238] Methods 500 and 600 are introduced above. Now, a method 800 for determining channel layout proposed in this application is introduced below, as shown in FIG8 . Before introducing method 800 , a brief introduction to the technical background of method 800 is first given.
[0239] In some cases, due to differences in production tools, audio production habits of mixers, etc., channel layout information may be lost during the audio production process, or the audio file may not be produced in the channel order indicated by the standard channel layout; for audio files without channel layout, after the playback end obtains the audio file, there is no available channel layout information to determine the speakers corresponding to each channel data; for audio files that are not produced in the channel order indicated by the standard channel layout, after the playback end obtains the audio file, if the speakers corresponding to each channel data are determined in accordance with the channel order indicated by the standard channel layout, the channel data may not correspond to the speakers. For example, the left front channel data is mapped to the woofer, and the woofer plays the left front channel data, which produces a playback effect that is inconsistent with the expectation. Therefore, the present application proposes a method 800 for determining the channel layout, which can solve the problem that the audio file cannot be played normally when the channel layout information of the audio file is lost or the audio file is not produced in the channel order indicated by the standard channel layout.
[0240] Figure 8 shows a flow chart of a method 800 for determining a channel layout; the execution subject involved in this method 800 may be a first device, or a chip, chip system, or processor applied to the first device, or a logic module or software that can implement all or part of the functions of the first device.
[0241] Exemplarily, the first device may be the electronic device 100 having the structure shown in FIG. 1 and FIG. 2 or the playback device 302 in FIG. 3 ; and the electronic device 100 or the playback device 302 may be a mobile phone, a tablet computer, or a car playback system.
[0242] The following embodiment describes the method 800 by taking the first device as an example. The method 800 includes steps S801 to S802, which are described in detail below.
[0243] S801: A first device obtains second audio data, where the second audio data includes at least two channel data.
[0244] The first device may refer to a playback device, for example, and the second audio data may refer to audio data or audio data in a video. For example, the second audio data may be audio data of a song or audio data of a movie.
[0245] The first device can obtain the second audio data from an audio workstation (or audio work device) or from some audio database, which is not limited in this application. The audio workstation is used to collect raw audio data and perform a series of processing (e.g., mixing) on the raw audio data to obtain produced audio data, which includes one or more channels of data. The second audio data can be understood as produced audio data (or audio files). The audio database is used to store some produced audio data (or audio files), such as audio data from music or videos.
[0246] It should be noted that the second audio data may include channel quantity information but not channel layout information, or may include channel quantity information and channel layout information. The channel layout indicated by the channel layout information may be a standard channel layout, a non-standard channel layout, or even erroneous channel layout information. A standard channel layout can be understood as each channel data being arranged in the channel order indicated by the standard channel layout, and a non-standard channel layout can be understood as each channel data not being arranged in the channel order indicated by the standard channel layout. Incorrect channel layout information can be understood as channel layout information that does not conform to the channel layout used in the actual audio data production process. In short, both a non-standard channel layout and an erroneous channel layout are considered erroneous channel layout information for the playback end.
[0247] For example, the arrangement order of the channel data of the 10 channels in the standard 7.1.2 channel is shown in Table 5, while for non-standard 7.1.2 channels, the arrangement order of the channel data of some channels may be disordered, as shown in Table 6; compared with the arrangement order of the channel data corresponding to channel numbers 4 to 10 in Table 5, the arrangement order of the channel data corresponding to channel numbers 4 to 10 in Table 6 is disordered. It should be noted that in Table 5, the channel number indicates the arrangement order of different channel data in the 7.1.2 channel; the channel position indicates the channel data at different positions in the 7.1.2 channel, for example, left front indicates left front channel data, right front indicates right front channel data, bass indicates bass channel data, left overhead indicates left overhead channel data, and so on; the channel name indicates the abbreviation of different channel data in the 7.1.2 channel.
[0248] It should be noted that "FL" corresponds to the left front channel data, "FR" corresponds to the right front channel data, "FC" channel corresponds to the center channel data, "LEF" corresponds to the bass channel data, "BL" corresponds to the left rear channel data, "BR" corresponds to the right rear channel data, "SL" corresponds to the left surround channel data, "SR" corresponds to the right surround channel data, "LTM" corresponds to the left overhead channel data, and "RTM" corresponds to the right overhead channel data.
[0249] Table 5
[0250] Table 6
[0251] For example, in some cases, when the mixer produces the second audio data, he actually produces the audio data according to 5.1.4 channels, but when marking the channel layout information, he marks it as 7.1.2 channels. Although the number of channels is the same, the different channel layouts correspond to different channel data arrangement orders. This incorrect channel layout information will also cause the playback end to be unable to play the audio correctly.
[0252] S802: The first device determines the channel layout of the second audio data through a detection algorithm.
[0253] After the first device receives the second audio data, regardless of whether the second audio data includes channel layout information, the first device will determine the channel layout of the second audio data through a detection algorithm; in other words, when the second audio data does not include channel layout information, the first device determines the channel layout of the second audio data through a detection algorithm; when the second audio data includes channel layout information, the first device proofreads the channel layout information included in the second audio data through a detection algorithm to avoid the situation where the channel layout information carried by the second audio data is a non-standard channel layout and affects the playback effect.
[0254] Among them, the detection algorithm can also be called a channel layout detection algorithm or a layout detection algorithm, which is used to detect or verify the channel layout of audio data (for example, the second audio data); the detection algorithm can be a detection algorithm module in the first device, or it can be a detection chip, which is not limited in this application; the relevant description of the channel layout can refer to the relevant description in the above method 600, and will not be repeated here; and the detection algorithm can filter the second audio data (for example, use a low-pass filter to filter the various channel data in the second audio data), calculate (for example, calculate the energy, energy proportion, duration of the audio signal, etc. of each channel data in the second audio data), compare and calculate (for example, compare the energy proportion of each channel data, etc.) and calculate the correlation coefficient (for example, calculate the correlation coefficient of part of the channel data) and other processing.
[0255] For example, when the second audio data includes channel quantity information but does not include channel layout information, the first device may detect (or determine) the channel layout of the second audio data through a detection algorithm.
[0256] For another example, when the second audio data includes channel quantity information and channel layout information, regardless of whether the channel layout indicated by the channel layout information is a standard channel layout or an incorrect channel layout, the first device can verify the channel layout included in the second audio data through a detection algorithm to ensure that the channel layout used by the first device when determining the corresponding speaker based on the channel layout information is the verified channel layout.
[0257] In some embodiments, the first device determines the channel layout of the second audio data by using a detection algorithm, including: the first device determines the channel layout of the second audio data by using a detection algorithm in response to the second audio data.
[0258] In order to avoid using non-standard channel layout information or erroneous channel layout information to determine the speakers corresponding to each channel data, after receiving the second audio data, the first device triggers the detection algorithm module in response to the second audio data to verify the channel layout information included in the second audio data; for example, the channel layout of the second audio data is 2.1 channels, the left channel data corresponds to the left speaker, the right channel data corresponds to the right speaker, and the bass data corresponds to the bass speaker; the first device filters the each channel data in the second audio data through the detection algorithm, and determines that the 2.1 channels included in the second audio data are the standard channel layout.
[0259] For another example, the second audio data includes the number of channels but does not include channel layout information. In this case, the first device, in response to the second audio data, identifies the speakers corresponding to the respective channel data in the second audio data through a detection algorithm, i.e., determines the channel layout of the second audio data through the detection algorithm. For example, the second audio data includes two channel data, and the first device identifies the speaker corresponding to the first channel data as the left speaker and the speaker corresponding to the second channel data as the right speaker through the detection algorithm. The first device determines the channel layout of the second audio data as 2.0 channels through the detection algorithm. The first device can send the two channel data to the corresponding speakers for playback based on the 2.0 channel layout. It should be noted that the specific processing process of the first device determining the channel layout of the second audio data through the detection algorithm can be referred to in the embodiments below and will not be described in detail here.
[0260] In this embodiment, regardless of whether the second audio data includes channel layout information or whether the channel layout indicated by the included channel layout information is a standard channel layout, after the first device receives the second audio data, it will re-determine the channel layout of the second audio data through a detection algorithm to avoid the situation where the erroneous channel layout information carried by the second audio data affects the playback effect.
[0261] To sum up, since the produced audio data (for example, the second audio data) may have lost channel layout information or non-standard channel layout; therefore, in this application, the first device performs channel layout detection on the second audio data through a detection algorithm, and determines the channel layout of the second audio data, and sends at least two channel data in the second audio data to the corresponding speakers for playback according to the channel layout, thereby solving the problem that the audio data cannot be played normally when the channel layout information is lost or the audio data is not produced in the channel order indicated by the standard channel layout.
[0262] In one possible implementation, the first device determines the channel layout of the second audio data through a detection algorithm, including: when the first device determines through the detection algorithm that the first channel data meets a first condition, determining that the first channel data is bass channel data, wherein the first condition includes: the energy proportion of the low-frequency signal in the first channel data is greater than the energy proportion of the low-frequency signal of other channel data in the second audio data except the first channel data, and the energy proportion of the high-frequency signal in the first channel data is less than the energy proportion of the high-frequency signal of other channel data in the second audio data except the first channel data, and the first channel data is any one of the at least two channel data.
[0263] Among them, the energy proportion of the low-frequency signal can refer to the ratio of the energy of the low-frequency signal of a single channel data to the total energy of the single channel data in the entire time period; for example, the total duration of the first channel data is T, the energy of the low-frequency signal of the first channel data is e1, and the total energy of the first channel data in the T time period is E, the energy proportion of the low-frequency signal of the first channel data is the ratio of e1 to E (i.e., e1 / E).
[0264] The energy proportion of the high-frequency signal can refer to the ratio of the energy of the high-frequency signal of a single channel data to the total energy of the single channel data in the entire time period; for example, the total duration of the first channel data is T, the energy of the high-frequency signal of the first channel data is e2, and the total energy of the first channel data in the T time period is E, the energy proportion of the high-frequency signal of the first channel data is the ratio of e2 to E (i.e., e2 / E).
[0265] The first device detects each channel data in at least two channel data through a detection algorithm; first, the first device calculates the energy of the low-frequency signal of each channel data (for example, the first channel data, the second channel data, etc.) and calculates the energy of the high-frequency signal of each channel data (for example, the first channel data, the second channel data, etc.) through the detection algorithm; second, the first device compares the channel data with the largest proportion of low-frequency signal energy in each channel data and compares the channel data with the smallest proportion of high-frequency signal energy in each channel data through the detection algorithm; finally, the first device determines the channel data with the largest proportion of low-frequency signal energy and the smallest proportion of high-frequency signal energy from the at least two channel data. For example, the first device determines that the channel data that meets the first condition from the at least two channel data is the first channel data, and the first channel data is bass channel data (or subwoofer data); and the bass channel data can be played through any one of a subwoofer, a subwoofer speaker, a subwoofer horn, or a subwoofer speaker.
[0266] For example, the second audio data includes three channel data, namely the first channel data, the second channel data and the third channel data; the first device calculates the energy proportion of the low-frequency signal of the first channel data as E1, the energy proportion of the low-frequency signal of the second channel data as E2 and the energy proportion of the low-frequency signal of the third channel data as E3 through the detection algorithm; the first device compares E1, E2 and E3 in pairs through the detection algorithm. For example, if E1 is greater than E2, E2 is less than E3 and E1 is greater than E3, then E1 is the maximum value of E1, E2 and E3, indicating that the energy proportion E1 of the low-frequency signal in the first channel data is greater than the energy proportion of the low-frequency signals of the other channel data except the first channel data in the second audio data (for example, E2 and E3); similarly, the first device calculates the energy proportion of the high-frequency signal of the first channel data as E4, the energy proportion of the high-frequency signal of the second channel data as E5 and the high-frequency signal of ... The energy proportion of the high-frequency signal in the first channel data is E5, and the energy proportion of the high-frequency signal in the third channel data is E6; the first device compares E4, E5 and E6 in pairs through the detection algorithm. For example, if E4 is less than E5, E5 is less than E6, and E4 is greater than E6, then E4 is the minimum value among E1, E2 and E3, indicating that the energy proportion E4 of the high-frequency signal in the first channel data is less than the energy proportion of the high-frequency signal of other channel data other than the first channel data in the second audio data (for example, E5 and E6); the first device determines that the first channel data meets the first condition based on the above comparison result, and therefore, the first channel data is bass channel data; the first device can render the bass channel data to the bass speaker for playback, thereby avoiding the situation where the bass channel data is sent to speakers at other positions (for example, sky position speakers) and cannot be played or the playback effect does not meet expectations.
[0267] In another possible implementation, the first device determines the channel layout of the second audio data through a detection algorithm, further including: when the first device determines through the detection algorithm that the second channel data meets a second condition, determining that the second channel data is sky channel data, wherein the second condition includes: the duration of the audio signal in the second channel data is less than a first preset value, and the energy of the audio signal in the second channel data is less than an energy threshold, and the second channel data is any one of at least two channel data.
[0268] The first preset value (or energy threshold) can be set based on experience or according to actual application scenarios, and this application does not limit this. For example, the first preset value can be 2 seconds or 3 seconds, and the energy threshold can be 5dB or 10dB.
[0269] The duration of the above-mentioned audio signal can be understood as the duration of the audio signal in a single channel data within the total duration T of the single channel data; for example, the total duration of the first channel data is T; within the T period, the duration of the audio signal in the first channel data is T1, where T1 is less than T.
[0270] The first device calculates the duration of the audio signal of each channel data in the second audio data and the energy of the audio signal through a detection algorithm; for example, if the duration of the audio signal of the second channel data is less than a first preset value, and the energy of the audio signal of the second channel data is less than an energy threshold, the first device can determine that the second channel data is sky channel data; and the sky channel data can be played through a sky speaker (for example, a left overhead speaker).
[0271] For example, the first preset value is T0, and the energy threshold is E0; the second audio data includes 10 channel data, namely, channel 01 data, channel 02 data, ..., channel 10 data; the first device calculates through a detection algorithm that the duration of the audio signal of channel 01 data is T1 and the energy of the audio signal is E1, the duration of the audio signal of channel 02 data is T2 and the energy of the audio signal is E2, ... and the duration of the audio signal of channel 10 data is T10 and the energy of the audio signal is E10; the first device compares T1, T2, ... and T10 with T0, and compares E1, E2, ... and E10 with E0 respectively through a detection algorithm, and if except that T1 is less than T0 and E1 is less than E0, T2 is less than T0 and E2 is less than E0, T5 is less than T0 and E5 is less than E0, 0 and T6 is less than T0 and E6 is less than E0, T3, T4, T7, T8, T9 and T10 are all greater than T0 and E3, E4, E7, E8, E9 and E10 are all greater than E0, indicating that the 01st channel data, the 02nd channel data, the 05th channel data and the 06th channel data are sky channels; the first device determines, based on the above comparison result, that the second channel data (for example, one of the 01st channel data, the 02nd channel data, the 05th channel data or the 06th channel data) meets the second condition, therefore, the second channel data is sky channel data; the first device can render the sky channel data to the sky speaker (for example, the left overhead speaker) for playback, thereby avoiding the situation where the sky channel data is sent to speakers at other positions (for example, the subwoofer) and cannot be played or the playback effect does not meet expectations.
[0272] It should be noted that the first device can first determine the bass channel data through the first condition and then determine the sky channel data through the second condition; or it can first determine the sky channel data through the second condition and then determine the bass channel data through the first condition; this application does not limit this.
[0273] In one possible implementation, the first device determines the channel layout of the second audio data through a detection algorithm, and also includes: when the first device determines through the detection algorithm that the third channel data does not meet the third condition and does not meet the fourth condition, determining that the third channel data belongs to surround channel data, and the third channel data is one of at least two channel data, the third condition includes: the energy proportion of the low-frequency signal in the third channel data is greater than the energy proportion of the low-frequency signal of other channel data in the second audio data except the third channel data, and the energy proportion of the high-frequency signal in the third channel data is less than the energy proportion of the high-frequency signal of other channel data in the second audio data except the third channel data, and the fourth condition includes: the duration of the audio signal in the third channel data is less than the first preset value, and the energy of the audio signal in the third channel data is less than the energy threshold.
[0274] Among them, the surround channel data occupies the energy of the main audio signal of the second audio data and can be played through the surround speakers; the surround channel data includes but is not limited to left front channel data, right front channel data, center channel data, left rear channel data and right rear channel data.
[0275] The surround speakers are usually placed at the same level as the human ears; the surround speakers include but are not limited to the left front speaker, right front speaker, center speaker, left rear speaker and right rear speaker; for example, the left front speaker, left side (or left center) speaker or left rear speaker are at the same level as the left ear, and the right front speaker, right side (or right center) speaker or right rear speaker are at the same level as the right ear.
[0276] It should be noted that the first device filters out bass channel data and sky channel data from at least two channel data through the first condition or the second condition; and if the remaining channel data in the second audio data, except for the bass channel data and the sky channel data, does not meet the third condition and does not meet the fourth condition, then the remaining channel data can be determined as surround channel data.
[0277] It should be noted that this application does not limit the order in which the first device determines the bass channel data, the sky channel data, and the surround channel data. For example, the first device may first determine the bass channel data, then determine the sky channel data, and finally determine the surround channel data; it may also first determine the sky channel data, then determine the bass channel data, and finally determine the surround channel data; it may also first determine the surround channel data, then determine the sky channel data, and finally determine the bass channel data.
[0278] For another example, the second audio data includes 10 channel data, namely channel 01 data, channel 02 data, channel 03 data, ..., channel 10 data; the channel layout is 5.1.4 channels. If the channel 07 data to the channel 10 data meet the second condition, the first device can determine that the channel 07 data to the channel 10 data are sky channels; if the channel 06 data meets the first condition, the first device can determine that the channel 06 data is bass channel data; and the channel 01 data to the channel 05 data neither meet the third condition nor the fourth condition. Therefore, the channel 01 data to the channel 05 data can be determined as surround channel data. The surround channel data needs to be sent to the surround speakers for playback to obtain high-quality playback effects. The first device can render the surround channel data to the surround speakers (for example, the left front speaker) for playback to avoid the situation where the surround channel data is sent to speakers in other positions (for example, the subwoofer) and the playback cannot be played or the playback effect does not meet expectations.
[0279] It should be noted that after the first device determines the surround channel data, it can further determine which channel data in the surround channel data is front channel data and which is rear channel data. The front channel data includes, but is not limited to, left front channel data, right front channel data, and center channel data; and the rear channel data includes, but is not limited to left rear channel data and right rear channel data.
[0280] In one possible implementation, the surround channel data includes front channel data and rear channel data, and the first device determines that the third channel data is surround channel data, including: when the third channel data meets a fifth condition, the first device determines that the third channel data belongs to the front channel data, or when the third channel data does not meet the third condition, the first device determines that the third channel data belongs to the rear channel data, wherein the fifth condition includes: an energy value of the audio signal in the third channel data is greater than a second preset value.
[0281] The second preset value can be set based on empirical values or actual application scenarios, and this application does not impose any restrictions on this. For example, since in actual scenarios, the sound perceived by the listener mainly comes from the front speakers, the energy of the front channel data played by the front speakers is usually very high. Therefore, the second preset value can be set higher, for example, the second preset value can be set to 120dB or 150dB, etc., so that the front channel data and the rear channel data can be distinguished by the second preset value.
[0282] For example, the first device calculates the energy value E of the audio signal in the third channel data through a detection algorithm; if the energy value E is greater than the second preset value, it means that the first channel data meets the fifth condition, and the first device can determine that the third channel data belongs to the front channel data; if the energy value E is less than the second preset value, it means that the third channel data does not meet the fifth condition, and the first device can determine that the third channel data belongs to the rear channel data.
[0283] In this embodiment, the first device further determines from the surround channel data whether the third channel data belongs to the front channel data or the rear channel data through the fifth condition, so as to accurately send the third channel data to the corresponding front speaker (for example, FL speaker or FR speaker) or rear speaker (for example, BL speaker or BR speaker) to ensure a good playback effect.
[0284] It should be noted that after the first device determines the front channel data and the rear channel data, it can further determine the channel data belonging to the same side in the front channel data and the rear channel data. For some complex channel layouts (such as 5.1.4 channels, 7.1.2 channels, or 7.1.4 channels), determining the channel data belonging to the same side from the front channel data and the rear channel data is of great significance.
[0285] In one possible implementation, the front channel data includes fourth channel data, the rear channel data includes fifth channel data, the fourth channel data and the fifth channel data are two different channel data among the at least two channel data, and the method 800 further includes: when the absolute value of the correlation coefficient between the fourth channel data and the fifth channel data is greater than or equal to a coefficient threshold, the fourth channel data and the fifth channel data are left channel data, or the fourth channel data and the fifth channel data are right channel data.
[0286] The coefficient threshold may be a positive real number less than 1, and may be set based on empirical values or actual application scenarios, which is not limited in this application. For example, the coefficient threshold may be 0.9 or 0.95.
[0287] It should be noted that when the third channel data belongs to one of the front channel data, the fourth channel data may be the third channel data; or when the third channel data belongs to one of the rear channel data, the fifth channel data may be the third channel data.
[0288] After the first device determines the front channel data and the rear channel data, it obtains a channel data from the front channel data, for example, the fourth channel data, and then obtains a channel data from the rear channel data, for example, the fifth channel data; the first device calculates the correlation coefficient r of the fourth channel data and the fifth channel data, and compares the correlation coefficient r with the coefficient threshold. If r is greater than or equal to the coefficient threshold, it means that the fourth channel data and the fifth channel data are highly correlated and belong to ipsilateral channel data, where the ipsilateral channel data includes left channel data or right channel data; for example, the fourth channel data is left front channel data, and the fifth channel data is left rear channel data; or, the fourth channel data is right front channel data, and the fifth channel data is right rear channel data.
[0289] If r is less than the coefficient threshold, it means that the fourth channel data and the fifth channel data have little correlation and are non-ipsilateral channel data. For example, the fourth channel data is the left front channel data or the center channel data, and the fifth channel data is the right rear channel data; for another example, the fourth channel data is the right front channel data or the center channel data, and the fifth channel data is the left rear channel data.
[0290] In this embodiment, the first device determines whether the fourth channel data and the fifth channel data are same-side channel data by calculating the correlation coefficient between the fourth channel data and the fifth channel data, so as to further determine the speakers corresponding to each channel data in the front channel data and the rear channel data, thereby avoiding the situation where the channel data corresponds to the wrong speaker and affects the playback effect.
[0291] Methods 500 to 800 and possible implementations are introduced in detail above. Methods 500 to 800 and possible implementations are further explained below in combination with different application scenarios.
[0292] As shown in Figure 9, a schematic diagram of a waveform display interface of an audio software is shown; the audio software can display audio data with different channel layouts; the user can view the audio waveform corresponding to each channel data in the second audio data on the audio software; when displaying the second audio data, the audio software usually detects whether the second audio data carries channel layout information. If it does, the audio software will display the audio waveform of each channel data in the second audio data according to the channel layout information; if it does not, the audio software will select a default layout based on the number of channels to display the audio waveform of each channel data in the second audio data. The user can use the audio software to view the audio waveform of each channel data before and after the first device processes the second audio data through the detection algorithm.
[0293] For example, as shown in (a) in Figure 9, the waveform display interface of the audio software displays second audio data with a channel layout of 7.1.2 channels; the channel order of the 10 channel data of the second audio data is arranged according to the channel order indicated by the standard 7.1.2 channels; wherein, the channel order indicated by the standard 7.1.2 channels can be referred to in Table 5; for another example, as shown in (b) in Figure 9, the waveform display interface of the audio software also displays second audio data with a channel layout of 7.1.2 channels; however, the channel order of the 10 channel data of the second audio data is not arranged according to the channel order indicated by the standard 7.1.2 channels; wherein, the channel order indicated by the non-standard 7.1.2 channels can be referred to in Table 6, which is not repeated here; when displaying the second audio data, the audio software displays it according to the channel order indicated by the non-standard 7.1.2 channels. Obviously, the channel order will be different from the channel order shown in (a) in Figure 9. For example, the fourth channel position shown in Figure 9 (a) is the LEF channel, while the fourth channel position shown in Figure 9 (b) is the Lw channel. It should be noted that the "L" channel on the waveform display interface corresponds to the FL channel data, the "R" channel corresponds to the FR channel data, the "C" channel corresponds to the FC channel data, the "LEF" channel corresponds to the LEF channel data, the "Ls" channel corresponds to the SL channel data, the "Rs" channel corresponds to the SR channel data, the "Lw" channel corresponds to the BL channel data, the "Rw" channel corresponds to the BR channel data, the "Tsl" channel corresponds to the LTM channel data, and the "Tsf" channel corresponds to the RTM channel data.
[0294] Through the audio software, it can be seen that the arrangement order of the channel data indicated by the non-standard channel layout has changed compared to the arrangement order of the channel data indicated by the standard channel layout. If the first device renders the audio data of the non-standard channel layout (for example, the second audio data) to the corresponding speaker according to the audio data of the standard channel layout, it will produce a playback effect that is inconsistent with expectations.
[0295] For example, in a scenario where the second audio data carries a channel layout, the user can use audio software to view the second audio data before or after processing by the first device using the detection algorithm, thereby determining whether the channel layout carried by the second audio data is a standard channel layout. Of course, in a scenario where the second audio data does not carry a channel layout, the user can also use audio software to view the second audio data after processing by the first device using the detection algorithm to determine which channel layout the second audio data was produced according to.
[0296] As shown in Figure 10, a schematic diagram of the speaker layout in a cinema scene is shown; (a) in Figure 10 shows a planar schematic diagram of a 7.1.2 speaker layout corresponding to 7.1.2 channels. The 7.1.2 speaker layout includes 7 main speakers, 1 subwoofer and 2 overhead speakers, namely the left front speaker, the right front speaker,..., the left overhead speaker and the right overhead speaker; and speakers in different positions are used to play different channel data; for example, the main speakers correspond to surround channel data, the subwoofer corresponds to bass channel data and the overhead speakers correspond to sky channel data. In actual applications, speakers are placed according to the 7.1.2 channel channel layout. For example, as shown in (b) in Figure 10, two speakers for playing the sky channel data, such as the left overhead speaker and the right overhead speaker, are installed in the upper space (for example, the top of the room, etc.); seven speakers for playing the surround channel data, such as the left front speaker and the right front speaker, are installed at a horizontal position flush with the human ear; one speaker for playing the bass channel data, such as a subwoofer or a woofer or a heavy woofer, is installed directly in front of or to the right of the listener, and the left and right speakers and the subwoofer are tried to be on the same line on the ground to obtain a better playback effect.
[0297] It should be noted that in some scenarios, the position descriptions in Figure 10 may also have different alternative descriptions. For example, "left front" can be replaced by "left front", "right front" can be replaced by "right front", "bass" can be replaced by "subwoofer", "left overhead" can be replaced by "left overhead", "right overhead" can be replaced by "right overhead", "left rear" can be replaced by "left rear" and "right rear" can be replaced by "right rear".
[0298] As shown in Figure 11, a schematic diagram of a detection process for determining channel layout is shown. Taking the case where a first device detects the channel layout of second audio data as an example, the second audio data includes 10 channel data but does not include channel layout information. The detection algorithm module of the first device includes: a low-pass filter, a calculation module, and a comparison module, wherein the low-pass filter is used to filter out high-frequency signals; the calculation module is used to calculate the energy and duration of the audio signal; and the comparison module is used to compare the relationship between energy ratio, correlation coefficient, and other values and a threshold value (e.g., a first preset value). The first device can first determine whether bass channel data exists in the 10 channel data based on a first condition. For example, the first device inputs the 10 channel data into the low-pass filter to obtain the 10 filtered channel data. The calculation module calculates the energy ratio of each of the 10 filtered channel data. The channel data with the largest energy ratio (e.g., the first channel data) can be determined as bass channel data or low-frequency effect (LEF) channel data. The bass channel data can be played through a woofer (or subwoofer or woofer).
[0299] The first device then determines the sky channel data in the other 9 channel data except the bass channel data through the second condition; the first device calculates the overall signal energy E of each of the other 9 channel data, and the duration T of the audio signal in each channel data, wherein the channel data whose energy E is less than the energy threshold and whose duration T is less than the first preset value can be determined as the sky channel data; the sky channel data includes 4 channel data, namely, left front (LTF) channel data, right front (RTF) channel data, left rear (LTR) channel data and right rear (RTR) channel data.
[0300] After the first device determines the bass channel data and the sky channel data, it can determine that the other five channel data are surround channel data according to the third condition and the fourth condition; the first device further determines which of the surround channel data are front channel data and which are rear channel data according to the fifth condition; the first device can compare the energy E of the other five channel data with the second preset value through the comparison module, wherein the channel data with energy E greater than the second preset value is the front channel data, and the channel data with energy E less than or equal to the second preset value is the rear channel data; wherein the front channel data includes left front (FL) channel data, right front (FR) channel data and center (FC) channel data; the rear channel data includes left back (BL) channel data and right back (BR) channel data; the first device arbitrarily obtains one channel data 1 (for example, left front channel data or right front channel data) from the front channel data channel data or center channel data) and arbitrarily obtain one channel data 2 (for example, left rear channel data or right rear channel data) from the rear channel data, and calculate the correlation coefficient of the two channel data through the calculation module. If the correlation coefficient is greater than or equal to the coefficient threshold, channel data 1 and channel data 2 are ipsilateral channel data; if the correlation coefficient is less than the coefficient threshold, channel data 1 and channel data 2 are non-side channel data; for example, if the correlation coefficient between channel data 1 and channel data 2 is greater than or equal to the coefficient threshold, it means that channel data 1 is left front channel data and channel data 2 is left rear channel data, or, channel data 1 is right front channel data and channel data 2 is right rear channel data; after the first device determines the left front channel data, left rear channel data, right front channel data and right rear channel data through the relationship between the correlation coefficient and the coefficient threshold, the remaining channel data is the center channel data. Finally, the first device determines that the 10 channels of audio data include 5 surround sound channels, 1 bass channel, and 4 overhead channels. Therefore, the channel layout of the second audio data is 5.1.4. The first device renders the 10 channels of audio data to the corresponding speakers according to the 5.1.4 layout for playback, achieving a good playback effect.
[0301] It should be noted that the method for processing audio data or determining the channel layout provided by this application is not limited to scenarios such as home and theaters, but can also be applied to scenarios such as vehicle-mounted systems, concerts, and TV programs. This application only introduces the method for processing audio data or determining the channel layout using home and theater scenarios as examples, and should not be understood as a limitation on the application scenarios of this application.
[0302] The above details examples of the method for processing audio data or the method for determining a channel layout provided by this application. It is understood that, to implement the aforementioned functions, the terminal device includes hardware structures and / or software modules corresponding to each function. Those skilled in the art will readily appreciate that, in conjunction with the units and algorithmic steps of the various examples described in the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is implemented in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application. This application may divide the method for processing audio data or the method for determining a channel layout into functional units based on the aforementioned method examples. For example, each function may be divided into separate functional units, or two or more functions may be integrated into a single unit. Such integrated units may be implemented in either hardware or software functional units. It should be noted that the division of units in this application is illustrative and merely represents a logical functional division; alternative divisions may be used in actual implementation.
[0303] FIG12 is a schematic diagram of the structure of an electronic device provided by the present application. The dotted lines in FIG12 indicate that the unit or module is optional. Electronic device 1200 can be used to implement the method described in the above method embodiment. Electronic device 1200 can be a server or a chip (system).
[0304] The electronic device 1200 includes one or more processors 1201, which can support the electronic device 1200 to implement the methods in the method embodiments corresponding to Figures 5, 6, and 8. The processor 1201 can be a general-purpose processor or a dedicated processor. For example, the processor 1201 can be a central processing unit (CPU). The CPU can be used to control the electronic device 1200, execute software programs, and process data of the software programs. The electronic device 1200 can also include a communication unit 1205 to implement signal input (reception) and output (transmission).
[0305] The electronic device 1200 may be a chip (system) including a memory and a processor, wherein the processor is configured to execute a computer program stored in the memory to implement the methods shown in the above embodiments.
[0306] The communication unit 1205 may be an input and / or output circuit of the chip (system), or the communication unit 1205 may be a communication interface of the chip (system), and the chip (system) may be a component of the electronic device 1200 .
[0307] For another example, the communication unit 1205 may be a transceiver of the electronic device 1200, or the communication unit 1205 may be a transceiver circuit of the electronic device 1200. The electronic device 1200 may include one or more memories 1202, on which a program 1204 is stored. The program 1204 can be executed by the processor 1201 to generate instructions 1203, so that the processor 1201 performs the method described in the above method embodiment according to the instructions 1203. Optionally, data may also be stored in the memory 1202. Optionally, the processor 1201 may also read data stored in the memory 1202. The data may be stored at the same storage address as the program 1204, or the data may be stored at a different storage address than the program 1204.
[0308] The processor 1201 and the memory 1202 may be provided separately or integrated together, for example, on a system-on-chip (SOC) of an electronic device. The specific manner in which the processor 1201 executes the method for processing audio data or the method for determining the channel layout may be found in the relevant description of the method embodiment.
[0309] It should be understood that each step of the above method embodiment can be completed by hardware logic circuits or software instructions in the processor 1201. The processor 1201 can be a CPU, a digital signal processor (DSP), a field programmable gate array (FPGA), or other programmable logic devices, such as discrete gates, transistor logic devices, or discrete hardware components.
[0310] The present application also provides a computer program product that, when executed by processor 1201, implements any method embodiment of the present application. The computer program product may be stored in memory 1202, for example, a program 1204. Program 1204 undergoes preprocessing, compilation, assembly, and linking to be converted into an executable object file that can be executed by processor 1201.
[0311] The present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a computer, implements any method embodiment of the present application. The computer program may be a high-level language program or an executable target program.
[0312] The computer-readable storage medium is, for example, memory 1202. Memory 1202 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus RAM (DRRAM).
[0313] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and equipment and the technical effects produced can refer to the corresponding processes and technical effects in the aforementioned method embodiments, and will not be repeated here.
[0314] In several embodiments provided in this application, the disclosed systems, devices, and methods can be implemented in other ways. For example, some features of the method embodiments described above can be ignored or not executed. The device embodiments described above are merely schematic, and the splitting of units is only a logical function splitting. There may be other splitting methods in actual implementation, and multiple units or components may be combined or integrated into another system. In addition, the coupling between the units or the coupling between the components may be direct coupling or indirect coupling, and the above coupling includes electrical, mechanical or other forms of connection.
[0315] The above embodiments are intended only to illustrate the technical solutions of the present application and are not intended to limit the same. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they may still modify the technical solutions described in the above embodiments or replace some of the technical features therein with equivalents, and that such modifications or replacements do not deviate from the essence of the corresponding technical solutions within the scope of the technical solutions of the embodiments of the present application and are therefore intended to be included within the scope of protection of the present application.
[0316] Finally, the above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application shall be covered by the scope of protection of the present application. Therefore, the scope of protection of the present application shall be based on the scope of protection of the claims.
Claims
1. A method for processing audio data, characterized in that: The method comprises: The first device receives first audio data and channel layout information, where the first audio data includes M channel data, and the channel layout information is used to indicate speakers corresponding to the M channel data, where M is a positive integer greater than 1; The first device determines, from at least two speakers according to the channel layout information, speakers corresponding to the M channel data.
2. The method according to claim 1, characterized in that The first device determines, from at least two speakers according to the channel layout information, speakers corresponding to the M channel data, including: When the number of speakers corresponding to the M channel data matches the number of the at least two speakers, the first device determines the speakers corresponding to the M channel data from the at least two speakers according to the channel layout information.
3. The method according to claim 2, characterized in that The first device determines, from the at least two speakers according to the channel layout information, a speaker corresponding to the M channel data, including: When the channel layout information matches the speaker layout information, the first device determines the speakers corresponding to the M channel data from the at least two speakers according to the channel layout information, where the speaker layout information is position distribution information of each speaker in the at least two speakers.
4. The method according to any one of claims 1 to 3, characterized in that The metadata includes extended information including the channel layout information.
5. The method according to claim 1, wherein The first device determines, from at least two speakers according to the channel layout information, speakers corresponding to the M channel data, including: When the number of speakers corresponding to the M channel data does not match the number of the at least two speakers, the first device performs mixing processing on the M channel data according to the channel layout information, the speaker layout information, and a preset mapping relationship to obtain N channel data and channel layout information of the N channel data, wherein the speaker layout information is position distribution information of each of the at least two speakers, and the preset mapping relationship is used to map the number of channels indicated by the channel layout information to the number of channels indicated by the speaker layout information, wherein N matches the number of the at least two speakers, the channel layout information of the N channel data is consistent with the speaker layout information, and N is a positive integer greater than 1; The first device determines, from the at least two speakers, a speaker corresponding to the N channel data according to the channel layout information of the N channel data.
6. The method according to claim 5, characterized in that The metadata includes extended information, and the extended information includes the channel layout information and the preset mapping relationship.
7. The method according to claim 6, characterized in that The extended information includes first indication information, second indication information, first information, and second information. The first indication information is used to indicate that the first information includes the channel layout information, and the second indication information is used to indicate that the second information includes the preset mapping relationship.
8. The method according to claim 7, characterized in that The preset mapping relationship includes: an upmix mapping relationship and a downmix mapping relationship, the upmix mapping relationship is used to map the number of channels indicated by the channel layout information to the number of channels indicated by the speaker layout information through an upmix rule, and the downmix mapping relationship is used to map the number of channels indicated by the channel layout information to the number of channels indicated by the speaker layout information through a downmix rule, the second information includes third indication information, fourth indication information, third information and fourth information, the third indication information is used to indicate that the third information includes the upmix mapping relationship, and the fourth indication information is used to indicate that the fourth information includes the downmix mapping relationship.
9. The method according to any one of claims 5 to 8, characterized in that After the first device determines the speakers corresponding to the N channel data, the method further includes: The first device configures a correspondence between the N channel data and the at least two speakers according to the channel layout information of the N channel data.
10. A method for processing audio data, characterized in that: The method comprises: The second device generates channel layout information, where the channel layout information is used to indicate speakers corresponding to M channel data in the first audio data, where M is a positive integer greater than 1; The second device sends the first audio data and the channel layout information to the first device, and the first device is configured to play the first audio data according to the channel layout information.
11. The method according to claim 10, characterized in that The metadata includes extended information including the channel layout information.
12. A method for determining channel layout, characterized in that: The method comprises: Acquire second audio data, where the second audio data includes at least two channel data; The channel layout of the second audio data is determined by a detection algorithm.
13. The method according to claim 12, characterized in that The determining the channel layout of the second audio data by using a detection algorithm includes: When it is determined through the detection algorithm that the first channel data satisfies a first condition, the first channel data is determined to be bass channel data, wherein the first condition includes: an energy proportion of low-frequency signals in the first channel data is greater than an energy proportion of low-frequency signals in other channel data other than the first channel data in the second audio data, and an energy proportion of high-frequency signals in the first channel data is less than an energy proportion of high-frequency signals in other channel data other than the first channel data in the second audio data, and the first channel data is any one of the at least two channel data.
14. The method according to claim 12 or 13, characterized in that The determining the channel layout of the second audio data by a detection algorithm further includes: When the detection algorithm determines that the second channel data satisfies a second condition, the second channel data is determined to be sky channel data, wherein the second condition includes: a duration of the audio signal in the second channel data is less than a first preset value, and energy of the audio signal in the second channel data is less than an energy threshold, and the second channel data is one of the at least two channel data.
15. The method according to any one of claims 12 to 14, characterized in that The determining the channel layout of the second audio data by a detection algorithm further includes: When it is determined through the detection algorithm that the third channel data does not satisfy the third condition and does not satisfy the fourth condition, the third channel data is determined to belong to surround channel data, and the third channel data is one of the at least two channel data. The third condition includes: an energy proportion of low-frequency signals in the third channel data is greater than an energy proportion of low-frequency signals of other channel data in the second audio data except the third channel data, and an energy proportion of high-frequency signals in the third channel data is less than an energy proportion of high-frequency signals of other channel data in the second audio data except the third channel data. The fourth condition includes: a duration of the audio signal in the third channel data is less than a first preset value, and the energy of the audio signal in the third channel data is less than an energy threshold.
16. The method according to claim 15, characterized in that The surround channel data includes front channel data and rear channel data, and determining that the third channel data belongs to the surround channel data includes: When the third channel data satisfies a fifth condition, the third channel data is determined to be front channel data, or when the third channel data does not satisfy the fifth condition, the third channel data is determined to be rear channel data, wherein the fifth condition includes: an energy value of an audio signal in the third channel data is greater than a second preset value.
17. The method according to claim 16, characterized in that The front channel data includes fourth channel data, the rear channel data includes fifth channel data, the fourth channel data and the fifth channel data are two different channel data among the at least two channel data, and the method further includes: When the absolute value of the correlation coefficient between the fourth channel data and the fifth channel data is greater than or equal to a coefficient threshold, the fourth channel data and the fifth channel data are left channel data, or the fourth channel data and the fifth channel data are right channel data.
18. An audio processing system, characterized in that: The audio processing system includes a first device, a second device, and at least two speakers. The first device is used to perform the method according to any one of claims 1 to 9, and the second device is used to perform the method according to claim 10 or 11.
19. An electronic device, characterized in that: The electronic device includes a processor and a memory, the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the electronic device executes the method described in any one of claims 1 to 9, or the electronic device executes the method described in any one of claims 5 or 6, or the electronic device executes the method described in any one of claims 12 to 17.
20. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the processor executes the method according to any one of claims 1 to 9, or the processor executes the method according to claim 5 or 6, or the processor executes the method according to any one of claims 12 to 17.
21. A chip system, characterized in that: The chip system includes a memory and a processor, and the processor is configured to execute a computer program stored in the memory to implement the method as claimed in any one of claims 1 to 9, or to implement the method as claimed in any one of claims 5 or 6, or to implement the method as claimed in any one of claims 12 to 17.
Citation Information
Patent Citations
Audio playing method and electronic equipment
CN110809226A
Audio playing method and device and electronic equipment
CN111857473A
Loudspeaker control method and device, electronic equipment and storage medium
CN114125655A
Display device, external device and audio output method
CN115884061A
Immersive audio fading
WO2023239639A1