Multichannel mixing method, device and medium

CN116962955BActive Publication Date: 2026-09-15HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210414876.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-15
Publication Date
2026-09-15
Estimated Expiration
2042-04-15

AI Technical Summary

Technical Problem

[0006]本申请实施例提供了一种多通道的混音方法、设备及介质,解决了目前声道下混方案中,下混后的音频数据破音,影响用户听觉体验的问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116962955B_ABST
    Figure CN116962955B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of audio processing, in particular to a multi-channel mixing method, device and medium. The method comprises the following steps: acquiring first multi-channel audio data, wherein the first multi-channel audio data comprises audio data of M to-be-mixed channels; determining that audio data with energy meeting a preset energy threshold exists in the first multi-channel audio data, and performing energy reduction processing on the audio data with energy greater than the preset energy threshold in the first multi-channel audio data; obtaining second multi-channel audio data according to the energy reduction processing result; and performing downmixing on the second multi-channel audio data to obtain mixed output data with N mixed channels, wherein M>N and N>=1. The multi-channel mixing method provided in the application can solve the problem of broken sound caused by excessively high energy of part of audio frames during channel downmixing, obtain a more ideal channel downmixing result, and improve the auditory experience of users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, specifically to a multi-channel mixing method, device, and medium. Background Technology

[0002] With the rapid development of modern technology, in various scenarios requiring audio playback, the mismatch between audio data and the number of channels on the audio output device often necessitates real-time multichannel mixing during audio output. This typically involves converting multichannel audio data into audio data with fewer channels, a process known as channel downmixing. For example, when playing AiMax content on a large screen, there may be 3.1, 5.1, or 7.1 multichannel audio data. However, when the large screen output device switches to a digital audio interface (Sony / Philips Digital Interface, SPDIF) / Audio Return Channel (ARC) / Bluetooth output, it may only output two channels. To preserve as much audio stream information as possible, it is necessary to downmix the multichannel data to generate two-channel data.

[0003] Currently, the following two schemes are generally used for multi-channel mixing to convert to two-channel downmixing:

[0004] 1) Use the first two channels of a multi-channel audio system as output, discarding the center channel, surround channels, and bass channel. This approach results in the loss of vocal data because some human voice audio data appears in the discarded channels. Additionally, using only two channels as output reduces the user's listening experience.

[0005] 2) Using the Dolby downmixing scheme, the data related to the left and right channels in the multi-channel audio data are weighted and summed to obtain the two-channel audio data output. However, for audio data that does not conform to Dolby specifications, such as audio data with high bass audio energy, distortion may occur when using the Dolby downmixing scheme for channel downmixing, resulting in a poor listening experience for the user. Summary of the Invention

[0006] This application provides a multi-channel mixing method, device, and medium, which solves the problem of audio data distortion after downmixing in current channel downmixing schemes, affecting the user's listening experience.

[0007] In a first aspect, embodiments of this application provide a multi-channel audio mixing method applied to an electronic device, comprising: acquiring first multi-channel audio data, the first multi-channel audio data including audio data of M channels to be mixed; determining that there are audio data in the first multi-channel audio data whose energy meets a preset energy threshold, and performing energy reduction processing on the audio data in the first multi-channel audio data whose energy is greater than the preset energy threshold; obtaining second multi-channel audio data based on the energy reduction processing result; and downmixing the second multi-channel audio data to obtain mixed output data with N mixing channels, wherein M>N and N≥1.

[0008] It is understandable that the first multi-channel audio data is the input data during channel downmixing, and the second multi-channel audio data is the output data after channel downmixing.

[0009] In some embodiments, the preset energy threshold is a pre-set minimum energy value that may cause distortion in the mixed audio. In some embodiments, the preset energy threshold is other pre-set energy values ​​that may cause distortion in the mixed audio and affect the user's listening experience. This application does not impose any limitations on this.

[0010] In some embodiments, the first multi-channel audio data can be 2.1 channel, 3.1 channel, 5.1 channel, 7.1 channel, or other multi-channel audio data, and the mixing output data can be mono audio data, two-channel audio data, or other multi-channel audio data not exceeding the first multi-channel audio data.

[0011] It is understood that the multi-channel mixing method in this application involves mixing multi-channel audio data (i.e., the first multi-channel audio data) into audio data with fewer channels, i.e., mixed output data. Before downmixing each channel's audio data, energy tracking is performed on the audio data of each channel to identify audio data exceeding a preset energy threshold, and energy suppression is applied to obtain energy-suppressed second multi-channel audio data. This second multi-channel audio data is then downmixed. The multi-channel mixing method in this application can fully adapt to and support downmixing of various multi-channel audio data, and can solve the distortion problem caused by excessively high energy in some audio frames during downmixing, resulting in a more ideal downmixing result and improving the user's listening experience.

[0012] In one possible implementation of the first aspect above, determining that there is audio data in the first multi-channel audio data with energy greater than a preset energy threshold includes: performing frame segmentation processing on the first multi-channel audio data to obtain multiple audio frames, and determining the frame energy of the multiple audio frames; determining that there are high-energy audio frames in the first multi-channel audio data with frame energy greater than the preset energy threshold.

[0013] It is understood that in some embodiments, audio frames whose frame energy does not exceed a preset energy threshold in the first multi-channel audio data are low-energy audio frames, and low-energy audio frames may not be subjected to energy reduction processing.

[0014] In one possible implementation of the first aspect above, the audio data in the first multi-channel audio data with energy greater than a preset energy threshold is subjected to energy reduction processing to obtain the second multi-channel audio data, including: determining the target gain of the high-energy audio frame, and determining the frame gain of the high-energy audio frame according to the target gain; and determining the target audio frame corresponding to the high-energy audio frame after energy reduction processing according to the frame gain of the high-energy audio frame.

[0015] It can be understood that the target gain is the energy suppression factor when performing energy reduction processing on high-energy audio frames, and the energy reduction of high-energy audio frames can be achieved by using this energy suppression factor.

[0016] In some embodiments, a target gain may also be applied to low-energy audio frames, where the target gain of the low-energy audio frame is 1, meaning that no energy reduction is applied to it.

[0017] In one possible implementation of the first aspect described above, the frame energy of the high-energy audio frame is determined by the following formula: The high-energy audio frame consists of L sampling points; β represents the frame energy smoothing coefficient; x i (n)(k) represents the audio data of the k-th sampling point in the n-th audio frame of the i-th audio channel to be mixed among M audio channels to be mixed; This represents the energy of the k-th sample point in the n-th audio frame of the i-th audio channel to be mixed out of M audio channels to be mixed; This represents the frame energy of the nth audio frame of the i-th channel among M channels to be mixed.

[0018] In some embodiments, each audio frame may include L = 512 sampling points, that is, the frame length of the audio frame is 512. In other embodiments, L may be other values, which are not limited in this application.

[0019] In one possible implementation of the first aspect above, the preset energy threshold includes a first threshold and / or a second threshold; the high-energy audio frame includes at least one of the following: among the multiple audio frames of the M mixing channels, the audio frame whose average frame energy is greater than the first threshold is a high-energy audio frame; among the audio frames of the same mixing channel, the audio frame whose maximum frame energy is greater than the second threshold is a high-energy audio frame among at least two consecutive audio frames of the corresponding audio frame.

[0020] It can be understood that the index of an audio frame is the sequence number of a certain audio frame in any one of the M channels to be mixed. For example, for the nth audio frame of the i-th channel to be mixed in the M channels to be mixed, its index is n.

[0021] In one possible implementation of the first aspect above, the maximum frame energy of each audio frame of the M channels to be mixed is determined based on the frame energy of the audio frame with the highest frame energy among the audio frames that correspond to the same mixing channel and have the same index as each audio frame.

[0022] In one possible implementation of the first aspect described above, the target gain of the high-energy audio frame is determined based on a preset energy threshold and the maximum frame energy of at least two audio frames consecutive to each high-energy audio frame.

[0023] In one possible implementation of the first aspect above, the frame gain is determined by the following formula: Where α represents the frame gain smoothing coefficient; This represents the target gain of the nth audio frame in the i-th channel of the M channels to be mixed. This represents the frame gain of the (n-1)th audio frame of the i-th audio channel to be mixed out of M audio channels to be mixed. This represents the frame gain of the nth audio frame in the i-th channel of the M channels to be mixed.

[0024] In some embodiments, low-energy audio frames in the first multi-channel audio data whose frame energy does not exceed a preset energy threshold can also have their frame gain calculated using the above formula, wherein the target gain of the low-energy audio frame is 1.

[0025] In one possible implementation of the first aspect above, determining the target audio frame corresponding to the high-energy audio frame after energy reduction processing based on the frame gain of the high-energy audio frame includes: determining the sampling point gain of each sampling point in the high-energy audio frame based on the frame gain of the high-energy audio frame; performing energy reduction processing on the audio data of each sampling point in the high-energy audio frame based on the sampling point gain to obtain the audio data of each sampling point in the target audio frame; and generating the target audio frame based on the audio data of each sampling point in the target audio frame.

[0026] In one possible implementation of the first aspect above, the gain at each sampling point is determined by the following formula:

[0027]

[0028] Where FrameLen represents the frame length of the target audio frame; FrameGain xi(n-1) FrameGain represents the frame gain of the (n-1)th audio frame of the i-th channel among M channels to be mixed. xi(n)Let GainBuff[i][k] represent the frame gain of the nth audio frame in the i-th channel of the M channels to be mixed. xi(n-1) GainBuff[i][k] represents the sampling gain of the k-th sample point in the (n-1)-th audio frame of the i-th channel to be mixed out of M channels; xi(n) This represents the sampling gain of the k-th sample point of the n-th audio frame in the i-th channel of the M-th audio channels to be mixed.

[0029] In some embodiments, low-energy audio frames in the first multi-channel audio data whose frame energy does not exceed a preset energy threshold can also have their sampling point gain calculated using the above formula, wherein the frame gain of the low-energy audio frame is calculated based on the target gain of the low-energy audio frame being 1.

[0030] In one possible implementation of the first aspect described above, the audio data of each sampling point in the target audio frame is determined by the audio data of each sampling point in the high-energy audio frame corresponding to the target audio frame and the corresponding sampling point gain.

[0031] In one possible implementation of the first aspect above, obtaining the second multi-channel audio data based on the energy reduction processing result includes: generating the second multi-channel audio data based on the target audio frame and low-energy audio frames in the first multi-channel audio data whose energy is not greater than a preset energy threshold.

[0032] In one possible implementation of the first aspect above, downmixing the second multi-channel audio data to obtain mixed output data with N second channels includes: weighted summation of target audio frames and low-energy audio frames corresponding to the same second channel in the second multi-channel audio data to obtain mixed output data.

[0033] It can be understood that the process of obtaining the mixed output data described above is to perform channel downmixing on the second multi-channel audio data using the Dolby downmixing method, where the weighting coefficients for weighted summation are preset parameters.

[0034] Secondly, embodiments of this application provide an electronic device, including: one or more processors; one or more memories; the one or more memories storing one or more programs, which, when executed by one or more processors, cause the electronic device to perform the aforementioned multi-channel mixing method.

[0035] Thirdly, embodiments of this application provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the aforementioned multi-channel mixing method.

[0036] Fourthly, embodiments of this application provide a computer program product, including a computer program / instructions, which, when executed by a processor, implement the aforementioned multi-channel mixing method. Attached Figure Description

[0037] Figure 1 The diagram shown is a scenario illustration of the multi-channel mixing method provided in an embodiment of this application.

[0038] Figure 2 The diagram shown is a flowchart illustrating the mixing method for converting six-channel downmixing into two-channel mixing according to an embodiment of this application.

[0039] Figure 3 The diagram shown is a flowchart of a multi-channel mixing method provided in an embodiment of this application.

[0040] Figure 4 The diagram shown is a flowchart of another multi-channel mixing method provided in an embodiment of this application.

[0041] Figure 5 The diagram shown is a flowchart of an energy suppression method provided in an embodiment of this application.

[0042] Figure 6 The figure shown is a schematic diagram of the bitstream waveform of multi-channel audio data provided in an embodiment of this application;

[0043] Figure 7 The diagram shows the bitstream waveform and energy spectrum of the mixed audio channel after channel downmixing.

[0044] Figure 8 The diagram shown is a schematic representation of the hardware structure of a mobile phone according to an embodiment of this application. Detailed Implementation

[0045] As mentioned earlier, audio data that does not conform to Dolby specifications will exhibit distortion issues when downmixed. Specifically, due to the excessively high energy of the bass channel audio data, the energy of the corresponding audio data will also be excessively high when weighted and summed to obtain the mixed audio data. Consequently, when this audio data is output, the user hears distorted audio, resulting in a poor listening experience.

[0046] To address the issue of audio distortion in the downmixed audio data, which negatively impacts the user's listening experience, as described in the aforementioned channel downmixing schemes, this application proposes a multi-channel mixing method. This method includes: an electronic device identifying audio data exceeding a preset energy threshold in each channel of the multi-channel audio data, performing energy suppression (i.e., energy reduction) on this data, and then calculating the suppressed audio data for each channel. Subsequently, the suppressed multi-channel audio data can be weighted and summed based on a preset channel downmixing algorithm to obtain the channel-downmixed audio data.

[0047] It is understood that in some embodiments, energy suppression can be achieved by calculating the frame gain of each audio frame in each channel and reducing the frame gain using a corresponding suppression factor to obtain suppressed audio frames, thereby obtaining energy-suppressed audio data.

[0048] It's understandable that the default channel downmixing algorithm is the correspondence between each channel before and after downmixing, as well as the weighting coefficients for the weighted summation. For example, in a six-channel downmixing scheme that is converted to two channels, where the six channels are the left channel, left surround channel, bass channel, center channel, right channel, and right surround channel, the default channel downmixing algorithm is: the left channel, left surround, bass, and center channel are downmixed as the left channel, and the right channel, right surround, bass, and center channel are downmixed as the right channel. The weighting coefficients for the weighted summation are the default weighting coefficients in the Dolby downmixing scheme.

[0049] It is understood that in some embodiments, multi-channel audio data can be divided into multiple audio frames, and energy tracking and energy suppression can then be performed based on these audio frames. Here, an audio frame is defined as multiple audio data segments divided through frame processing, with each segment constituting an audio frame. These audio frames overlap, and the segmentation is identical for each channel. In other embodiments, energy suppression can also be performed on audio data whose energy exceeds a preset energy threshold using other methods; this application does not impose limitations on this approach.

[0050] The multi-channel mixing method provided in this application embodiment, before performing weighted summation on the audio data of each channel, performs energy tracking on the audio frames of each channel to identify audio frames that exceed a preset energy threshold and performs energy suppression. This method can fully adapt to various audio data, support channel downmixing of various multi-channel audio data, solve the problem of distortion caused by excessive energy in some audio frames during channel downmixing, obtain a more ideal channel downmixing result, and improve the user's listening experience.

[0051] It is understood that the electronic devices in the embodiments of this application include, but are not limited to, mobile phones (including foldable screen phones), tablet computers, laptop computers, desktop computers, servers, wearable devices, head-mounted displays, mobile email devices, in-vehicle infotainment systems, portable game consoles, portable music players, e-reader devices, and televisions in which one or more processors are embedded or coupled. For ease of explanation, the following description uses a mobile phone as an example to illustrate this application.

[0052] The following is combined Figure 1 and Figure 2 Taking the output of audio data from mobile phone 100 to Bluetooth headset 200 as an example, the application scenarios in this application embodiment are introduced.

[0053] Figure 1 The diagram shows an application scenario of multi-channel mixing methods.

[0054] like Figure 1 As shown, the scenario includes a mobile phone 100 and a Bluetooth headset 200. The mobile phone 100 and the Bluetooth headset are wirelessly connected via Bluetooth.

[0055] When a user wears Bluetooth headset 200 and plays multi-channel audio data from mobile phone 100, mobile phone 100 needs to downmix the multi-channel audio data to convert it into two-channel audio data including left and right channels, and then send the audio data to Bluetooth headset 200 via Bluetooth. Bluetooth headset 200 will then play the audio data upon receiving it.

[0056] Figure 2 The diagram shows a flowchart of a mixing method that converts six channels into two channels.

[0057] Specifically, such as Figure 2 As shown, taking six-channel audio data as an example, the six channels include left channel A1, left surround channel A2, bass channel A3, center channel A4, right channel A5, and right surround channel A6. Before performing channel downmixing on the audio data, the mobile phone 100 can first perform frame processing on the six-channel audio data and determine the frame energy of each audio frame. When the frame energy of any channel's audio frame is greater than a preset energy threshold, the energy suppression factor of that audio frame is determined, energy suppression is performed on that audio frame, and the energy-suppressed six-channel audio data is calculated based on the suppressed audio frames. Then, the energy-suppressed six-channel audio data is downmixed using a preset channel downmixing algorithm to obtain two-channel audio data, including left channel a1 and right channel a2. The mobile phone 100 can then send the audio data of left channel a1 to the left earpiece of the Bluetooth headset 200 and the audio data of right channel a2 to the right earpiece of the Bluetooth headset 200.

[0058] It is understood that the multi-channel mixing method in this application can not only support the mixing of six channels into two channels, but also support the mixing of any M channels into N channels, where M>N.

[0059] The multi-channel mixing method in the embodiments of this application will be further described below with reference to the accompanying drawings.

[0060] Figure 3 The diagram shown is a flowchart of a multi-channel mixing method provided in an embodiment of this application.

[0061] like Figure 3 As shown, multi-channel mixing methods include:

[0062] 301: Get multi-channel audio data.

[0063] It is understandable that the different channels in multi-channel audio data (i.e., the first multi-channel audio data mentioned above) are different channel types, and each channel is a channel to be mixed, such as the left channel, right channel, left surround channel, right surround channel, etc. The audio data of different channels can be output through different speakers, or they can be output through the same channel through channel downmixing.

[0064] In some embodiments, the multi-channel audio data can be 3.1, 5.1, 7.1, or other multi-channel audio data. Specifically, a 3.1 channel includes a left channel, a bass channel, a center channel, and a right channel; a 5.1 channel includes a left channel, a left rear surround channel, a bass channel, a center channel, a right channel, and a right surround channel; and a 7.1 channel includes a left channel, a left rear surround channel, a bass channel, a center channel, a right channel, a right surround channel, a left rear surround channel, and a right rear surround channel. In some embodiments, the multi-channel audio data can also have more or fewer channels than those given in the examples above, and this application does not impose any limitations on this.

[0065] It is understood that in some embodiments, the acquired multi-channel audio data also includes a mask for the multi-channel audio data. The mask is used to overlay a layer on the multi-channel audio data to select or block some audio data. Based on the mask, the correspondence between the channels before and after mixing the multi-channel audio data can also be determined, that is, which channels' audio data before mixing are mixed to form the audio data of each channel after mixing.

[0066] In some embodiments, after acquiring multi-channel audio data, the multi-channel audio data can be initialized. Initialization may include determining a preset channel downmixing algorithm, specifically including: determining the correspondence between channels before and after mixing, and performing a weighted summation of the multi-channel audio data to obtain the weighting coefficients of the audio data for each channel after mixing. It can be understood that the audio data of at least one channel in the multi-channel audio data corresponding to the same mixed channel after channel downmixing (i.e., the output channel after channel downmixing) can be considered as a joint detection channel. Furthermore, when determining the energy of audio frames, the energy of the corresponding audio frames in the joint detection channel can be initially determined. For example, the average energy of the corresponding audio frames can be determined to see if it exceeds a set initial energy threshold. In some embodiments, initialization may also include initializing the parameters of the formulas in the preset channel downmixing algorithm or energy suppression algorithm.

[0067] 302: Divide the multi-channel audio data into frames to obtain multiple audio frames.

[0068] It is understandable that the characteristics of multichannel audio data and the parameters representing its essential features change over time; that is, audio data has time-varying characteristics. Therefore, the energy tracking and energy suppression of multichannel audio data discussed below can be based on short-time data, i.e., short-time analysis. Specifically, this can be achieved by dividing the multichannel audio data into multiple segments, each segment being an audio frame. Furthermore, the essential characteristics of the same audio frame remain unchanged or relatively stable.

[0069] Furthermore, in some embodiments, multi-channel audio data can be directly divided into frames, for example, each audio frame has a length of 10-30ms.

[0070] In other embodiments, multi-channel audio data can be sampled, continuous multi-channel audio data can be converted into discrete multi-channel audio data, and audio data from 512 consecutive sample points can be combined into an audio frame.

[0071] 303: If the frame energy of the m-th audio frame is greater than a preset energy threshold. The preset energy threshold is a pre-set value. If the frame energy of the m-th audio frame is greater than the preset energy threshold, it indicates that after channel downmixing, the audio portion corresponding to this frame may experience distortion, and energy suppression is required, i.e., step 304 is executed. If the frame energy of the m-th audio frame is less than or equal to the preset energy threshold, it indicates that after channel downmixing, the audio portion corresponding to this frame will not experience distortion, and energy suppression is not required, i.e., step 305 is executed.

[0072] It can be understood that the index of an audio frame is the sequence number of a certain audio frame in any one of the multiple channels of multi-channel audio data. For example, for the m-th audio frame of the i-th channel to be mixed in M ​​channels, its index is m.

[0073] In some embodiments, a preset energy threshold can be determined based on the minimum frame energy value that may cause distortion in the audio data after channel downmixing. For example, the preset energy threshold may be -6dB or -3dB. The specific threshold can be determined based on different electronic devices, audio output devices, and multi-channel audio data; this application does not impose any limitations on this.

[0074] In some embodiments, step 303 can determine whether the average frame energy of the m-th audio frame of each channel in each joint detection channel is greater than a set energy threshold by calculating the average frame energy and judging the average frame energy.

[0075] In some embodiments, step 303 may determine whether the frame energy of the m-th audio frame is greater than a set energy threshold simply by calculating the frame energy of the m-th audio frame for each channel and checking whether the frame energy of that audio frame is greater than a set energy threshold. In other embodiments, the frame energy of the m-th audio frame in each of the jointly detected channels may be taken as the frame energy of the m-th audio frame, and then it may be determined whether this frame energy is greater than a set energy threshold. Further, after determining the maximum frame energy of the m-th audio frame in each of the jointly detected channels as the frame energy of the m-th audio frame, it may be determined whether the maximum frame energy of at least one audio frame near the m-th audio frame is greater than a set energy threshold. For example, it may be determined whether the maximum frame energy of the m-th audio frame and the (m-1)-th audio frame is greater than a set energy threshold to determine whether the frame energy of the m-th audio frame is greater than a set energy threshold.

[0076] In some embodiments, setting the energy threshold includes an initial energy threshold (i.e., the first threshold mentioned above) and a fine-grained energy threshold (i.e., the second threshold mentioned above). Therefore, step 303 can perform an initial assessment of the frame energy of the m-th audio frame. Specifically, the average frame energy of the m-th audio frame in each joint detection channel is first calculated, and this average frame energy is then assessed to determine whether it exceeds the initial energy threshold. If it does, further fine-grained assessment is performed. Specifically, fine-grained assessment can be achieved by taking the maximum frame energy of the m-th audio frame in each joint detection channel as the frame energy of the m-th audio frame, and then determining whether the maximum frame energy among at least two consecutive audio frames near the m-th audio frame exceeds the fine-grained energy threshold. Step 303 will be further described below with reference to the formula.

[0077] In some embodiments, the frame energy of an audio frame can be obtained by performing a Fourier transform on the multichannel audio data to obtain the energy spectrum of the multichannel audio data, and the frame energy of the audio frame composed of multiple sampling points can be calculated based on the energy values ​​of multiple consecutive sampling points. The specific calculation method will be introduced in conjunction with the formula below.

[0078] 304: The m-th audio frame is suppressed using an energy suppression algorithm to obtain the m-th target audio frame.

[0079] It can be understood that the energy suppression algorithm is an algorithm that determines the target audio frame after energy suppression based on the frame energy of the m-th audio frame. In some embodiments, an energy suppression factor can be calculated based on the calculated frame energy of the m-th audio frame and a preset formula, and then energy suppression can be performed on the m-th audio frame based on the energy suppression factor to obtain the m-th target audio frame. In some embodiments, the energy suppression factor is a target gain, and thus energy suppression can be calculated based on the frame gain of the m-th audio frame according to the target gain, and the m-th target audio frame can be calculated based on the frame gain.

[0080] In some embodiments, when each audio frame includes multiple sampling points, when the target audio frame is calculated, the frame gain of the m-th audio frame can be calculated using the energy suppression factor, and the gain of each sampling point in the m-th audio frame (i.e., sampling point gain) can be calculated based on the frame gain of the m-th audio frame. Then, the audio data of each sampling point in the target audio frame can be determined based on the gain of each sampling point, and signal reconstruction can be performed to determine the audio data of the m-th target audio frame.

[0081] 305: Determine the m-th audio frame as the m-th target audio frame.

[0082] It is understood that in some embodiments, if the frame energy of the m-th audio frame does not exceed a preset energy threshold, then after channel downmixing, the audio data corresponding to the audio frame will not distort due to excessive energy, and thus will not affect the user's listening experience. Therefore, there is no need to suppress the energy of the audio frame, and the original audio data of the audio frame can be preserved.

[0083] 306: Based on a preset channel downmixing algorithm, channel downmixing is performed on each target audio frame of multiple channels to obtain mixed output data.

[0084] It can be understood that the preset channel downmixing rules include the corresponding channels before and after channel downmixing, as well as the channel downmixing calculation formula. That is, by using the corresponding channels before and after channel downmixing, the target audio frames of some channels in the multi-channel audio data of each mixed channel participating in the channel downmixing calculation can be determined. Then, based on the determined audio frames corresponding to the same mixed channel, they are substituted into the corresponding channel downmixing formula to achieve the mixed output data of that mixed channel. The mixed output data is the audio data of one channel output to the Bluetooth headset 200 mentioned earlier.

[0085] It is understood that in some embodiments, after obtaining the target audio frame in step 305 or 304, the Dolby downmixing algorithm can be used to perform channel downmixing to obtain the mixing output data of each mixing channel. Specifically, the target audio frames with corresponding indices in each jointly detected channel can be weighted and summed to obtain the mixed audio frames in the mixing output data corresponding to the target audio frame. Then, the mixed audio frames corresponding to the same mixing channel can be combined to obtain the output data of that mixing channel.

[0086] This application embodiment tracks the energy of speech data in multi-channel audio data using the aforementioned multi-channel mixing method, suppresses the energy of high-energy audio data, and then performs channel downmixing based on the energy-suppressed multi-channel audio data. The multi-channel mixing method in this application embodiment is applicable to both Dolby-compliant and non-Dolby-compliant multi-channel audio data; that is, it can be used for various multi-channel downmixing methods, achieving adaptive channel downmixing without distortion due to excessively high audio data energy. Furthermore, the multi-channel mixing method in this application embodiment only suppresses energy in a subset of high-energy audio data, without discarding any audio data. This solves the problem of downmixing distortion while preserving the audio data of each channel, improving the user's listening experience.

[0087] The following is combined Figure 4 Taking the mixing method of downmixing two channels into one channel as an example, the multi-channel mixing method in the embodiments of this application will be further introduced.

[0088] Figure 4 The diagram shown is a flowchart of another multi-channel mixing method according to an embodiment of this application.

[0089] like Figure 4 As shown, the multi-channel audio data includes the bitstream of audio data for channel 1 and the bitstream of audio data for channel 2. After acquiring the audio data for channel 1 and audio data for channel 2, the electronic device performs frame-based processing on the audio data of the two channels, and each channel can obtain 6 audio data frames.

[0090] From the bitstream of the audio data of each channel after frame segmentation, it can be seen that the second and fourth audio frames in channel 1 are not very smooth, and the second and third audio frames in channel 2 are also not very smooth. To avoid distortion in the mixed data after downmixing, energy suppression needs to be performed on the second and fourth audio frames in channel 1, and the second and third audio frames in channel 1, respectively, to obtain the bitstream of the target audio frame after suppression for each channel. Then, the Dolby downmixing algorithm can be used to perform a weighted summation of the corresponding target audio frames in the two downmixed channels to obtain the corresponding mixed audio frames in the mixed channel. The six mixed audio frames constitute the mixed output data.

[0091] The following is combined Figure 5 The present application will further describe an energy suppression method in one of its embodiments.

[0092] Figure 5 The diagram shown is a flowchart of an energy suppression method according to an embodiment of this application.

[0093] like Figure 5 As shown, the method includes:

[0094] 501: Multichannel data. The multichannel data obtained in step 501 is the same as the multichannel audio data obtained in step 301. Step 501 is similar to step 301, and will not be described in detail here.

[0095] 502: Generate an initial downmixing algorithm based on the channel layout to determine the mixing channels.

[0096] It can be understood that channel arrangement refers to the types and number of channels in multi-channel data. The initialization downmixing algorithm is a formula algorithm that generates the correspondence between channels before and after downmixing and the mixed channels based on the channel arrangement, and calculates the data of the mixed channels. Specifically, in some embodiments, generating the initialization downmixing algorithm may include determining the correspondence between M channels and N channels through a mask of multi-channel data. For example, when downmixing a six-channel system into two channels, the left channel, left surround, bass, and center channel are downmixed as the left channel, and the right channel, right surround, bass, and center channel are downmixed as the right channel. In some embodiments, generating the initialization downmixing algorithm also includes formulas for calculating the output data of each mixed channel and initializing the parameters in the formulas. Among them, the channels that generate the same mixed channel (i.e., the mixed channel) during channel downmixing can be represented as joint detection channels, that is, the data of each channel in the joint detection channel corresponding to the mixed channel are weighted and summed to obtain the data to be output by the mixed channel (i.e., the mixed channel).

[0097] 503: The VAD detection result is 1.

[0098] VAD detection, short for Voice Activity Detection, is a detection method where the frame energy of the audio data exceeds a set energy threshold.

[0099] In some embodiments, before performing VAD detection, the multi-channel data can be sampled and framed. Framed processing can involve dividing a portion of consecutive sampling points into an audio frame after sampling the multi-channel data, for example, taking 512 sampling points as one audio frame. Furthermore, in some embodiments, the VAD detection condition can be to detect whether the frame energy of the audio frame is greater than a preset VAD detection threshold (i.e., the first threshold mentioned above), where the VAD detection threshold can be, for example, a frame energy greater than -6dB. When the VAD detection decision result is true, that is, the VAD detection result is 1, it indicates that the frame energy of the audio frame exceeds the preset VAD detection threshold, and the audio frame may exhibit distortion after channel downmixing. It can be understood that VAD detection is the initial judgment mentioned above, and further refined judgment can be performed based on the initial judgment result.

[0100] In some embodiments, the average frame energy of each audio frame in a channel can be jointly detected to see if it exceeds a preset VAD detection threshold. If the average frame energy is too high, energy tracking needs to be performed on each audio frame in each channel, and energy suppression needs to be performed on audio frames that meet the energy suppression conditions.

[0101] It's understandable that energy tracking is only needed for the audio frame when the VAD detection result is 1. When the VAD detection result is false (i.e., 0), the audio will not distort due to excessive energy in the audio frame after downmixing. Then, VAD detection can be performed on the next audio frame.

[0102] 504: Calculate frame energy, track and jointly detect the audio channel and the maximum energy of the preceding and following frames.

[0103] It is understandable that if the VAD detection result is 1, then energy suppression needs to be performed on the corresponding audio frame.

[0104] Specifically, taking L sampling points as one speech frame, the frame energy of an audio frame can be calculated using the following formula:

[0105]

[0106] in, This represents the frame energy of the nth audio frame in the i-th channel. β is the smoothing coefficient used in calculating the frame energy, and x... i(n)(k) represents the input data of the k-th sampling point of the n-th audio frame in the i-th channel, which is a portion of the acquired multi-channel data. Here, i = 0, 1, 2, 3, 4, 5..., where the value of i is related to the number of channels, k can be an integer between 0 and L, and n represents the number of audio frames after framing the multi-channel data. In some embodiments, the smoothing coefficient β = 0.3.

[0107] It is understandable that in step 504 above, the frame energy is calculated. It can track the frame energy of data from each channel. When the frame energy of an audio frame is tracked and it is determined that the frame energy is greater than a set detection threshold, step 505 can be executed.

[0108] Specifically, the audio frame with the highest energy in the corresponding audio frame of the jointly detected channel can be calculated, which can be determined in the following way.

[0109]

[0110] in, This represents the maximum frame energy value of the nth audio frame in the data of each channel in the joint detection channel.

[0111] In some embodiments, after determining the maximum frame energy value of the current nth audio frame using the above formula (2), the maximum energy value between the preceding and following audio frames can be determined using the following formula:

[0112]

[0113] in, This represents the maximum energy value between the audio frames preceding and following the nth audio frame. This represents the maximum frame energy value of the (n-1)th audio frame in the data of each channel in the joint detection channel. This represents the maximum frame energy value of the (n+1)th audio frame in the data of each channel in the joint detection channel.

[0114] In some embodiments, It is also obtained by determining the maximum value of the maximum energy values ​​in the preceding and following audio frames of the nth and (n-1)th audio frames.

[0115] 505: Calculate the target gain and frame gain, and finally calculate the gain of each sampling point, multiply it by a fixed gain, and output the result.

[0116] It is understandable that the output results are the input data when performing channel downmixing.

[0117] In some embodiments, the calculation is obtained Then, we can first determine whether the energy of the current nth audio frame exceeds the set detection threshold (i.e., the second threshold mentioned above). If it does, it indicates that the audio frame has a risk of distortion after channel downmixing and needs to be suppressed. Then, we can calculate the target gain and frame gain of the audio frame to determine the energy suppression factor.

[0118] It is understandable that setting a detection threshold involves accurately judging the frame energy of each channel to determine whether audio frames with excessive energy need to be suppressed.

[0119] It is understandable that the target gain can be understood as an energy suppression factor, used to reduce the frame gain, thereby achieving the purpose of reducing the energy of the audio frame.

[0120] In some embodiments, the target gain can be calculated using the following formula:

[0121]

[0122] This is understandable, where Threshold represents the set detection threshold, Threshold = -3dB. This represents the target gain of the nth audio frame in the i-th channel.

[0123] As shown in formula (4), when the maximum frame energy value of the preceding and following frames is less than the set detection threshold Threshold, it indicates that the frame energy of the audio frame is appropriate and there will be no distortion due to excessive energy. When the maximum frame energy value of the preceding and following frames is greater than or equal to the set detection threshold Threshold, it indicates that the frame energy of the audio frame is too high, which may cause distortion due to excessive energy. It is necessary to calculate the target gain and suppress the frame gain of the audio frame.

[0124] In some embodiments, after calculating the target gain, the frame gain of the audio frame can be determined using the following formula:

[0125]

[0126] in, This represents the frame gain of the nth audio frame in the i-th channel. The frame gain represents the (n-1)th audio frame of the i-th channel, and α represents the smoothing coefficient when calculating the frame gain. In some embodiments, the smoothing coefficient α = 0.1.

[0127] In some embodiments, after calculating the frame gain according to the above formula (5), the sampling point gain of each sampling point in the audio frame can be calculated using the following formula:

[0128]

[0129] in, This represents the sampling gain of the k-th sample point in the n-th audio frame of the i-th channel. FrameLen represents the sampling gain of the k-th sample point in the (n-1)-th audio frame of the i-th channel, and FrameLen represents the frame length of the audio frame. For example, if an audio frame has 512 samples, then the frame length of that audio frame is 512.

[0130] As can be seen from the above formula (6), when calculating the sampling point gain, the sampling point gain of each sampling point is related to the frame gain of its corresponding speech frame and the sampling point gain of the corresponding index number of the previous speech frame. The index of each sampling point is the sequence number of each sampling point in the corresponding audio frame. For example, the index of the kth sampling point in an audio frame is k.

[0131] Furthermore, in some embodiments, the calculation formula for the target audio frame, obtained from the sampling point gain in Formula 6, is as follows:

[0132]

[0133] Where, α i (n)[i][k] represents the audio data of the k-th sample point of the n-th target audio frame of the i-th channel, x i (n)[i][k] represents the audio data of the k-th sample point of the n-th target audio frame of the i-th channel, i.e., x in Formula 1. i (n)(k).

[0134] It is understood that, in some embodiments, audio data of consecutive target audio frames can be obtained from the audio data of discrete sampling points in the target audio frame. The audio data of each target audio frame can then be used as input data for channel downmixing.

[0135] In some embodiments, after calculating the energy-suppressed audio frames, channel downmixing can be performed using the following formula:

[0136]

[0137] Among them, a j (n) represents the output result of the nth audio frame of the j-th mixing channel, w i,j (n) represents the mixing weight of the nth audio frame in the i-th channel of the j-th mixing channel, α i(n) represents the input data of the nth audio frame of the i-th channel during channel downmixing. When the channel downmixing is two-channel, j = 0, 1; when the channel downmixing is three-channel, j = 0, 1, 2, and so on. A = 0, 1, 2, 3, 4, 5... represents the multiple channels before mixing. For example, A = 0, 1, 2, 3, 4, 5 represents the left channel, left surround channel, bass channel, center channel, right channel, and right surround channel, respectively. J represents the number of channels before mixing corresponding to the j-th mixing channel.

[0138] To more clearly illustrate the positive effects of the multi-channel mixing method in the embodiments of this application, the following is combined with... Figure 6 and Figure 7 The multi-channel mixing method in the embodiments of this application is simulated. The simulation software is CLion. Furthermore, the multi-channel audio data is 5.1 channel audio data, with smoothing coefficients α = 0.1, β = 0.3, and a detection threshold Threshold = -3dB as the simulation conditions.

[0139] Figure 6 The diagram shown is a schematic diagram of the bitstream waveform of multi-channel audio data in an embodiment of this application.

[0140] Figure 7 The diagram shows the bitstream waveform and energy spectrum of the mixed audio channel after channel downmixing.

[0141] like Figure 6 As shown in the figure, the six bitstream waveforms, from top to bottom, represent the left channel, right channel, center channel, bass channel, left surround channel, and right surround channel. The horizontal axis represents time, and the vertical axis represents the audio data. During bitstream pass-through, each sample point is represented by 16 bits. Therefore, using a fixed-coefficient downmixing scheme, to ensure the mixed result is not distorted, the sum of the values ​​for each channel must be less than the range that 16 bits can represent; otherwise, the data will overflow and wrap around, causing data jumps and resulting in noise. Figure 6 If Dolby submixing is used for mixing, Figure 6 The data within the boxes exhibits some large energy peaks. To ensure distortion-free downmixing, the coefficients of the downmixing channels need to be very small. Furthermore, since volume and energy are positively correlated, the corresponding volume of the mixed channels will decrease, resulting in a lower overall audio volume after downmixing and negatively impacting the user's listening experience.

[0142] like Figure 7As shown, the first row of the bitstream waveform diagram is the bitstream waveform diagram of the audio data of the left channel mix obtained using the multi-channel mixing method of this application, and the second row of the bitstream waveform diagram is the bitstream waveform diagram of the audio data of the right channel mix obtained using the Dolby mixing method. In the first and second row of the bitstream waveform diagrams, the horizontal axis represents time, and the vertical axis represents the audio data. The third row corresponds to the audio data of the left channel mix in the first row and is the energy spectrum of the audio data of the left channel mix. The fourth row corresponds to the audio data of the right channel mix in the second row and is the energy spectrum of the audio data of the right channel mix. In the third and fourth row of energy spectra, the horizontal axis represents time, and the vertical axis represents energy.

[0143] Depend on Figure 7 It can be seen that the bitstream waveform of the audio data of the right channel after being processed by the multi-channel mixing method in this embodiment of the application is more stable in terms of envelope change, with very few large fluctuations, and its energy stripes are also relatively stable, without being excessively high and spreading throughout the entire frequency domain. In contrast, the bitstream waveform of the right channel mixed audio data that has not been processed by the multi-channel mixing method of this application shows that in its higher energy portion, for example... Figure 7 The audio data framed in the middle shows energy stripes across the entire frequency domain, indicating significant distortion and noticeable noise.

[0144] As can be seen, the multi-channel mixing method in this application embodiment performs energy tracking on the speech data of each channel, thereby suppressing the energy of audio data with excessively high energy. This reduces the risk of mixing distortion without losing audio data and improves the user's listening experience.

[0145] Figure 8 According to an embodiment of this application, a schematic diagram of the hardware structure of a mobile phone 100 is shown.

[0146] Mobile phone 100 is capable of executing the display method provided in the embodiments of this application. Figure 8 In this context, similar components share the same reference numerals. For example... Figure 8 As shown, the mobile phone 100 may include a processor 110, a power module 140, a memory 180, a camera 101, a mobile communication module 130, a wireless communication module 120, a sensor module 190, an audio module 150, an interface module 160, and a display screen 102, etc.

[0147] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the mobile phone 100. In other embodiments of this application, the mobile phone 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0148] Processor 110 may include one or more processing units, such as processing modules or circuits of a Central Processing Unit (CPU), Graphics Processing Unit (GPU), Image Signal Processor (ISP), Digital Signal Processor (DSP), Micro-programmed Control Unit (MCU), Artificial Intelligence (AI) processor, or Field Programmable Gate Array (FPGA). Different processing units may be independent devices or integrated within one or more processors. For example, in some embodiments of this application, processor 110 may be used to determine whether the energy of the m-th audio frame exceeds a set energy threshold and to calculate an energy suppression factor. In some embodiments, processor 110 may also be used to perform channel downmixing on the obtained target audio frame to obtain mixed output data.

[0149] The memory 180 can be used to store data, software programs, and modules. It can be volatile memory, such as random-access memory (RAM); or non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); or a combination of the above types of memory; or it can be a removable storage medium, such as a secure digital storage (SD) card. In some embodiments of the application, the memory 180 is used to store 100+ channel audio data of the mobile phone and a preset channel downmixing algorithm.

[0150] The power module 140 may include a power supply, a power management component, etc. The power supply may be a battery. The power management component is used to manage the charging of the power supply and the power supply to other modules. The charging management module is used to receive charging input from the charger; the power management module is used to connect to the power supply and the processor 110.

[0151] The mobile communication module 130 may include, but is not limited to, an antenna, a power amplifier, a filter, and a low-noise amplifier (LNA). The mobile communication module 130 can provide wireless communication solutions, including 2G / 3G / 4G / 5G, for use on the mobile phone 100. The mobile communication module 130 can receive electromagnetic waves via the antenna, filter and amplify the received electromagnetic waves, and then transmit them to a modem processor for demodulation. The mobile communication module 130 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via the antenna. In some embodiments, at least some functional modules of the mobile communication module 130 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 130 and at least some modules of the processor 110 may be housed in the same device.

[0152] The wireless communication module 120 may include an antenna, which enables the transmission and reception of electromagnetic waves. The wireless communication module 120 can provide solutions for wireless communication applications on the mobile phone 100, including Wireless Local Area Networks (WLANs) (such as Wireless Fidelity (Wi-Fi) networks), Bluetooth (BT), Global Navigation Satellite System (GNSS), Frequency Modulation (FM), Near Field Communication (NFC), and Infrared (IR) technologies. The mobile phone 100 can communicate with networks and other devices through wireless communication technologies.

[0153] In some embodiments, the mobile communication module 130 and the wireless communication module 120 of the mobile phone 100 may also be located in the same module.

[0154] Camera 101 is used to capture still images or videos. An object is projected onto a photosensitive element by an optical image generated through the lens. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to an ISP (Image Signal Processor) for conversion into a digital image signal. Mobile phone 100 can implement its shooting function through the ISP, camera 101, video codec, GPU (Graphics Processing Unit), display screen 102, and application processor. For example, in some embodiments of this application, camera 101 is used to capture facial images and QR code images for mobile phone 100 to perform facial recognition, QR code recognition, etc.

[0155] The display screen 102 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini LED, a Micro LED, a Micro OLED, a quantum dot light-emitting diode (QLED), etc. For example, the display screen 102 is used to display various UI interfaces of the mobile phone 100 in landscape / portrait modes, such as split-screen, parallel view, and single-app exclusive screen mode.

[0156] The sensor module 190 may include proximity sensors, pressure sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, accelerometers, distance sensors, fingerprint sensors, temperature sensors, touch sensors, ambient light sensors, bone conduction sensors, etc.

[0157] The audio module 150 can convert digital audio information into analog audio signal output, or convert analog audio input into digital audio signal. The audio module 150 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 150 can be located in the processor 110, or some functional modules of the audio module 150 can be located in the processor 110.

[0158] Interface module 160 includes an external memory interface, a Universal Serial Bus (USB) interface, and a Subscriber Identification Module (SIM) card interface. The external memory interface can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of mobile phone 100. The external memory card communicates with processor 110 through the external memory interface to perform data storage. The USB interface is used for communication between mobile phone 100 and other mobile phones. The SIM card interface is used to communicate with the SIM card installed in mobile phone 100, for example, to read or write phone numbers stored in the SIM card.

[0159] In some embodiments, the mobile phone 100 further includes buttons, a motor, and indicators. The buttons may include volume buttons, a power button, etc. The motor is used to generate a vibration effect in the mobile phone 100. The indicators may include a laser indicator, a radio frequency indicator, an LED indicator, etc.

[0160] The various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or a combination of these implementation methods. Embodiments of this application can be implemented as computer programs or program code executable on a programmable system, the programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0161] Program code can be applied to input instructions to execute the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, the processing system includes any system having a processor such as, for example, a Digital Signal Processor (DSP), a microcontroller, an Application Specific Integrated Circuit (ASIC), or a microprocessor.

[0162] The program code can be implemented using a high-level procedural language or an object-oriented programming language to communicate with the processing system. Assembly language or machine language can also be used to implement the program code when needed. In fact, the mechanisms described in this application are not limited to any particular programming language. In either case, the language can be a compiled language or an interpreted language. In some cases, the disclosed embodiments can be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments can also be implemented as instructions carried or stored thereon on one or more transient or non-transitory machine-readable (e.g., computer-readable) storage media, which can be read and executed by one or more processors. For example, the instructions can be distributed via a network or through other computer-readable media. Therefore, machine-readable media can include any mechanism for storing or transmitting information in a machine-readable (e.g., computer-readable) form, including but not limited to floppy disks, optical disks, CD-ROMs, magneto-optical disks, read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic cards or optical cards, flash memory, or tangible machine-readable storage for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in the form of electrical, optical, acoustic, or other forms of propagation signals. Therefore, machine-readable media includes any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a machine-readable (e.g., computer-readable) form.

[0163] In addition, the technical solution of this application also provides a computer-readable storage medium storing instructions, which, when executed on the electronic device 100, cause the electronic device 100 to perform the display method provided by the technical solution of this application.

[0164] In addition, the technical solution of this application also provides a computer program product, which includes instructions for implementing the display method provided by the technical solution of this application.

[0165] Furthermore, the technical solution of this application also provides a chip device, which includes: a communication interface for inputting and / or outputting information; and a processor for executing a computer-executable program, causing a device equipped with the chip device to execute the display method provided by the technical solution of this application.

[0166] In the accompanying drawings, some structural or methodological features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be necessary. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Furthermore, the inclusion of structural or methodological features in a particular figure does not imply that such features are required in all embodiments, and in some embodiments, these features may be omitted or may be combined with other features.

[0167] It should be noted that all units / modules mentioned in the device embodiments of this application are logical units / modules. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important factor; the combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed in this application. Furthermore, to highlight the innovative aspects of this application, the above-described device embodiments of this application have not introduced units / modules that are not closely related to solving the technical problems proposed in this application. This does not mean that the above-described device embodiments do not contain other units / modules.

[0168] It should be noted that in the examples and description of this patent, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0169] Although this application has been illustrated and described with reference to certain preferred embodiments thereof, those skilled in the art should understand that various changes in form and detail may be made thereto without departing from the spirit and scope of this application.

Claims

1. A multi-channel mixing method, applied to electronic devices, characterized in that, include: Acquire first multi-channel audio data, which includes audio data of M channels to be mixed; It is determined that there are audio data in the first multi-channel audio data with energy greater than a preset energy threshold, and the audio data in the first multi-channel audio data with energy greater than the preset energy threshold is subjected to energy reduction processing; Based on the energy reduction processing results, the second multi-channel audio data is obtained; The second multi-channel audio data is down-mixed to obtain mixed output data with N mixing channels, where M>N and N≥1; where, The step of determining that there is audio data in the first multi-channel audio data with energy greater than a preset energy threshold includes: The first multi-channel audio data is processed into frames to obtain multiple audio frames, and the frame energy of the multiple audio frames is determined. It was determined that there were high-energy audio frames in the first multi-channel audio data whose frame energy was greater than a preset energy threshold; The preset energy threshold includes a first threshold and / or a second threshold; The high-energy audio frame includes at least one of the following: Among the multiple audio frames of the M mixing channels, the audio frame with an average frame energy greater than the first threshold corresponding to at least one audio frame with the same index of the same mixing channel is the high-energy audio frame. Among the audio frames of the same audio channel to be mixed, the audio frame whose maximum frame energy is greater than the second threshold between at least two consecutive audio frames is the high-energy audio frame.

2. The multi-channel mixing method according to claim 1, characterized in that, The frame energy of the high-energy audio frame is determined by the following formula: ; The high-energy audio frame includes L sampling points; Represents the frame energy smoothing coefficient; This represents the audio data of the k-th sampling point in the n-th audio frame of the i-th channel to be mixed among the M channels to be mixed; This represents the energy of the k-th sampling point in the n-th audio frame of the i-th audio channel to be mixed among the M audio channels to be mixed; This represents the frame energy of the nth audio frame of the i-th audio channel among the M audio channels to be mixed.

3. The multi-channel mixing method according to claim 1, characterized in that, The maximum frame energy of each of the M audio frames to be mixed is determined based on the frame energy of the audio frame with the highest frame energy among the audio frames that correspond to the same mixing channel and have the same index as each audio frame.

4. The multi-channel mixing method according to claim 1, characterized in that, The energy reduction processing is performed on the audio data in the first multi-channel audio data whose energy is greater than the preset energy threshold, including: Determine the target gain of the high-energy audio frame, and determine the frame gain of the high-energy audio frame based on the target gain; Based on the frame gain of the high-energy audio frame, determine the target audio frame corresponding to the high-energy audio frame after energy reduction processing.

5. The multi-channel mixing method according to claim 4, characterized in that, The target gain of the high-energy audio frame is determined based on the preset energy threshold and the maximum frame energy of at least two consecutive audio frames with respect to each of the high-energy audio frames.

6. The multi-channel mixing method according to claim 5, characterized in that, The frame gain is determined by the following formula: ; in, Indicates the frame gain smoothing coefficient; This represents the target gain of the nth audio frame of the i-th channel among the M channels to be mixed; This represents the frame gain of the (n-1)th audio frame of the i-th audio channel to be mixed among the M audio channels to be mixed; This represents the frame gain of the nth audio frame of the i-th audio channel among the M audio channels to be mixed.

7. The multi-channel mixing method according to claim 4, characterized in that, The step of determining the target audio frame corresponding to the high-energy audio frame after energy reduction processing based on the frame gain of the high-energy audio frame includes: The sampling point gain of each sampling point in the high-energy audio frame is determined based on the frame gain of the high-energy audio frame. Based on the gain of each sampling point, the audio data of each sampling point in the high-energy audio frame is subjected to energy reduction processing to obtain the audio data of each sampling point in the target audio frame. The target audio frame is generated based on the audio data of each sampling point of the target audio frame.

8. The multi-channel mixing method according to claim 7, characterized in that, The gain of each sampling point is determined by the following formula: ; in, This indicates the frame length of the target audio frame; This represents the frame gain of the (n-1)th audio frame of the i-th audio channel to be mixed among the M audio channels to be mixed; This represents the frame gain of the nth audio frame of the i-th audio channel to be mixed among the M audio channels to be mixed; This represents the sampling gain of the k-th sampling point of the (n-1)-th audio frame of the i-th audio channel to be mixed among the M audio channels to be mixed; This represents the sampling gain of the k-th sampling point of the n-th audio frame of the i-th audio channel to be mixed among the M audio channels to be mixed.

9. The multi-channel mixing method according to claim 7, characterized in that, The audio data of each sampling point in the target audio frame is determined by the audio data of each sampling point in the high-energy audio frame corresponding to the target audio frame and the corresponding sampling point gain.

10. The multi-channel mixing method according to claim 4, characterized in that, The process of obtaining the second multi-channel audio data based on the energy reduction processing result includes: The second multi-channel audio data is generated based on the target audio frame and the low-energy audio frames in the first multi-channel audio data whose energy is not greater than a preset energy threshold.

11. The multi-channel mixing method according to claim 10, characterized in that, The downmixing of the second multi-channel audio data to obtain mixed output data with N mixing channels includes: The target audio frame and the low-energy audio frame corresponding to the same mixing channel in the second multi-channel audio data are weighted and summed to obtain the mixing output data.

12. An electronic device, characterized in that, include: Memory, used to store instructions executed by one or more processors of an electronic device, and A processor is one of the processors in an electronic device, used to control the execution of the multi-channel mixing method according to any one of claims 1 to 11.

13. A computer-readable storage medium, characterized in that, The storage medium stores instructions that, when executed on a computer, cause the computer to perform the multi-channel mixing method according to any one of claims 1 to 11.

14. A computer program product, characterized in that, The computer program product includes instructions for implementing the multichannel mixing method as described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Apparatus and method for generating a level parameter and apparatus and method for generating a multi-channel representation

    US20140236604A1