Real-time acquisition method and system for pulse code modulation data

By calculating the signal-to-noise ratio, spectrum flatness and channel coherence information of the audio channel, dynamic filtering and weighting mixing, the problem of unstable signal quality of multi-channel audio acquisition is solved, and the audio quality and user experience are improved.

CN120378799APending Publication Date: 2025-07-25TRONLONG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510490354.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing audio acquisition devices are difficult to ensure the stability of signal quality during the multi-channel pulse coding modulation data acquisition process, resulting in unbalanced volume, inaccurate frequency response, insufficient dynamic range, etc., affecting the overall audio quality and user listening experience.

Method used

By obtaining the pulse code modulation data of multiple audio channels in real time, calculating the signal-to-noise ratio, spectrum flatness and channel coherence information, normalizing and weighting to obtain the comprehensive weight, filtering the target mixing channel set, performing weighted mixing, and linearly adjusting the gain during channel switching to ensure smooth transitions of audio time domain signals and volume.

Benefits of technology

It improves the mixing effect and audio quality, improves the user's listening experience, ensures the continuity and stability of audio output, avoids the subjectivity of manual intervention, and improves the robustness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378799A_ABST
    Figure CN120378799A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a real-time acquisition method and system for pulse code modulation data. According to the technical scheme provided by the embodiment of the invention, the plurality of audio channels are screened from the plurality of audio channels based on the comprehensive weight, and weighted sound mixing is performed on the screened audio channels, so that dynamic screening of the pulse code modulation data is realized, the sound mixing effect and the audio quality are improved, and the hearing experience of a user is improved. In the process of outputting the mixed audio, the switched audio channel is added into the mixed audio, and the gain of the switched audio channel is gradually and linearly adjusted to 0 from a historical value, so that the smooth transition of the audio time domain signal and the volume during the switching of the audio channel is ensured, the audio time domain waveform is continuous, and the audio volume is kept stable; the audio output effect is ensured, and the hearing experience of the user is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of audio processing, and in particular, to a method and system for real-time acquisition of pulse code modulation data. Background Art

[0002] Currently, with the rapid development of digital technology, the acquisition and processing of audio data have become an important part of modern audio technology. In terms of audio acquisition, an audio acquisition device (such as a microphone) converts an analog sound wave signal into a digital signal that can be processed by a computer, namely PCM (Pulse Code Modulation data), through three key steps of signal sampling, quantization, and encoding. When outputting the pulse code modulation data acquisition, the method of collecting multiple audio channels is usually adopted to improve the audio quality and enhance the spatial perception ability. For example, in the ALSA (Advanced Linux Sound Architecture) framework, 8 channels are used to collect pulse code modulation data. By capturing audio information of different sound sources and different spatial positions through multiple channels, the audio acquisition in different scenarios such as meetings and performances can be realized. The multi-channel pulse code modulation data will be finely adjusted and balanced through audio parameters such as volume, frequency, dynamics, reverberation, and sound field, so as to ensure that the audio signal maintains high quality and consistency during the output process and meets the needs of users.

[0003] However, due to factors such as environmental noise interference, device performance limitations, and operation technical difficulties, it is often difficult for audio acquisition devices to ensure the stability of signal quality during real-time acquisition. Simply collecting the pulse code modulation data of all audio channels or fixedly selecting the pulse code modulation data of the corresponding audio channels for mixing may lead to problems such as uneven volume, inaccurate frequency response, and insufficient dynamic range, thus affecting the overall audio quality. This results in a poor audio playback effect and affects the user's listening experience. Summary of the Invention

[0004] The embodiments of the present application provide a method and system for real-time acquisition of pulse code modulation data, which can dynamically screen and switch audio channels to collect pulse code modulation data, improve the overall audio quality, and solve the technical problem of poor overall audio signal quality during the acquisition process of multi-channel pulse code modulation data.

[0005] In a first aspect, the embodiments of the present application provide a method for real-time acquisition of pulse code modulation data, including:

[0006] Obtain the pulse code modulation data of multiple audio channels in real time, calculate the signal-to-noise ratio, spectral flatness, and channel coherence information of each audio channel based on the pulse code modulation data, perform normalization processing on the signal-to-noise ratio, spectral flatness, and channel coherence information, and perform weighted summation on the normalized signal-to-noise ratio, spectral flatness, and channel coherence information to obtain the comprehensive weight of each audio channel;

[0007] Based on the comprehensive weight, screen several audio channels from multiple audio channels as the target mixing channel set for the current acquisition period. Based on the comprehensive weight of each audio channel in the target mixing channel set, the set adaptive attenuation coefficient, and the weak alignment coefficient, perform weighted mixing of the pulse code modulation data of each audio channel in the target mixing channel set for the current acquisition period to obtain the mixed audio;

[0008] Determine the switched audio channels in the historical mixing channel set of the previous acquisition period relative to the target mixing channel set. During the process of outputting the mixed audio in the current acquisition period, mix and output the mixed audio of the previous set number of frames with the pulse code modulation data of the switched audio channels, and during the mixing output process, linearly adjust the gain of the switched audio channels from the historical value to 0.

[0009] Further, the signal-to-noise ratio calculation formula is expressed as:

[0010] SNR = 10·log 10 (P signal / P noise )

[0011] where SNR represents the signal-to-noise ratio of the audio channel, P signal represents the signal power of the audio channel, and P noise represents the noise power of the audio channel;

[0012] The spectral flatness calculation formula is expressed as:

[0013]

[0014] where Flatness Factor represents the spectral flatness of the audio channel, y i represents the energy value of the frequency point i in the pulse code modulation data of the audio channel, and N represents the number of frequency sampling points;

[0015] The channel coherence information represents the mean of the channel correlation between the corresponding audio channel and other audio channels, and the channel coherence calculation formula is expressed as:

[0016]

[0017] Among them, represents the channel coherence between two channels, C xy (f) represents the cross-spectrum between two channels, P xx (f)P yy (f) represents the power spectrum of two channels.

[0018] Furthermore, the calculation formula of the comprehensive weight is expressed as:

[0019] W = w1·SNR norm + w2·Flatness norm + w3·Coherence norm

[0020] Among them, W represents the comprehensive weight, w1, w2, and w3 represent the corresponding weight coefficients, SNR norm represents the normalized parameter of the signal-to-noise ratio, Flatness norm represents the normalized parameter of the spectral flatness, Coherence norm represents the normalized parameter of the channel coherence information.

[0021] Furthermore, screening several audio channels from multiple audio channels as the target mixing channel set for the current acquisition period based on the comprehensive weight includes:

[0022] Screening the audio channels with the comprehensive weight higher than the set threshold from multiple audio channels as the target mixing channel set for the current acquisition period; or,

[0023] Screening the set number of audio channels with the highest comprehensive weight from multiple audio channels as the target mixing channel set for the current acquisition period.

[0024] Furthermore, based on the comprehensive weight, the set adaptive attenuation coefficient, and the weak alignment coefficient of each audio channel in the target mixing channel set, performing weighted mixing on the pulse code modulation data of each audio channel in the target mixing channel set for the current acquisition period to obtain a mixed audio, includes:

[0025] Configuring the initial weight of each audio channel based on the comprehensive weight of each audio channel in the target mixing channel set, and the comprehensive weight is proportional to the initial weight;

[0026] Adjusting the initial weight based on the set adaptive attenuation coefficient and weak alignment coefficient of each audio channel in the target mixing channel set to obtain a mixing weight, and performing weighted mixing on the pulse code modulation data of each audio channel in the target mixing channel set for the current acquisition period according to the mixing weight to obtain a mixed audio.

[0027] Further, the adaptive attenuation coefficient is calculated according to the phase difference between audio channels and is used to adjust the attenuation degree of the corresponding audio channel.

[0028] Further, the weak alignment coefficient is calculated according to the time delay difference between audio channels and is used to compensate the signal time delay of the corresponding audio channel.

[0029] In a second aspect, an embodiment of the present application provides a real-time acquisition system for pulse code modulation data, including:

[0030] A calculation module, configured to obtain the pulse code modulation data of multiple audio channels in real time, calculate the signal-to-noise ratio, spectral flatness, and channel coherence information of each audio channel based on the pulse code modulation data, perform normalization processing on the signal-to-noise ratio, the spectral flatness, and the channel coherence information, and perform weighted summation on the normalized signal-to-noise ratio, the spectral flatness, and the channel coherence information to obtain the comprehensive weight of each audio channel;

[0031] A mixing module, configured to screen several audio channels from multiple audio channels as the target mixing channel set for the current acquisition period based on the comprehensive weight, and perform weighted mixing on the pulse code modulation data of each audio channel in the target mixing channel set based on the comprehensive weight of each audio channel in the target mixing channel set, the set adaptive attenuation coefficient, and the weak alignment coefficient to obtain a mixed audio;

[0032] An output module, configured to determine the switched audio channels in the historical mixing channel set of the previous acquisition period relative to the target mixing channel set, and during the process of outputting the mixed audio in the current acquisition period, mix and output the mixed audio of the previous set number of frames with the pulse code modulation data of the switched audio channels, and linearly adjust the gain of the switched audio channels from the historical value to 0 during the mixing and output process.

[0033] In a third aspect, an embodiment of the present application provides an electronic device, including:

[0034] A memory and one or more processors;

[0035] The memory is used to store one or more programs;

[0036] When the one or more programs are executed by the one or more processors, the one or more processors implement the real-time acquisition method for pulse code modulation data as described in the first aspect.

[0037] In a fourth aspect, an embodiment of the present application provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the method for real-time acquisition of pulse code modulation data as described in the first aspect when executed by a computer processor.

[0038] In the embodiment of the present application, by acquiring pulse code modulation data of multiple audio channels in real time, calculating the signal-to-noise ratio, spectral flatness, and channel coherence information of each audio channel based on the pulse code modulation data, normalizing the signal-to-noise ratio, spectral flatness, and channel coherence information, and performing weighted summation on the normalized signal-to-noise ratio, spectral flatness, and channel coherence information to obtain the comprehensive weight of each audio channel; screening several audio channels from the multiple audio channels based on the comprehensive weight as the target mixing channel set for the current acquisition period, and performing weighted mixing on the pulse code modulation data of each audio channel in the target mixing channel set based on the comprehensive weight of each audio channel in the target mixing channel set, the set adaptive attenuation coefficient, and the weak alignment coefficient to obtain the mixed audio; determining the switched audio channels in the historical mixing channel set of the previous acquisition period relative to the target mixing channel set, and during the process of outputting the mixed audio in the current acquisition period, mixing and outputting the mixed audio of the previous set number of frames with the pulse code modulation data of the switched audio channels, and during the mixing and output process, linearly adjusting the gain of the switched audio channel from the historical value to 0. By adopting the above technical means, dynamic screening of pulse code modulation data is realized by screening audio channels for weighted mixing, thereby improving the mixing effect and audio quality and enhancing the user's listening experience. And during the process of outputting the mixed audio, the switched audio channel is added to the mixed audio, and the gain of the switched audio channel is gradually linearly adjusted from the historical value to 0, so as to ensure smooth transition of the audio time-domain signal and volume during audio channel switching, make the audio time-domain waveform continuous, keep the audio volume stable, ensure the audio output effect, and further enhance the user's listening experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a flowchart of a method for real-time acquisition of pulse code modulation data provided in Embodiment 1 of the present application;

[0040] Figure 2 is a schematic diagram of screening and mixing of multiple audio channels in Embodiment 1 of the present application;

[0041] Figure 3 is a flowchart of enhanced mixing in Embodiment 1 of the present application;

[0042] Figure 4 is a flowchart of acquisition and output of pulse code modulation data of multiple audio channels in Embodiment 1 of the present application;

[0043] Figure 5It is a schematic structural diagram of a real-time acquisition system for pulse code modulation data provided by Embodiment 1 of the present application;

[0044] Figure 6 It is a schematic structural diagram of a real-time acquisition device for pulse code modulation data provided by Embodiment 1 of the present application Detailed implementation manners

[0045] In order to make the objectives, technical solutions, and advantages of the present application clearer, the following further describes specific embodiments of the present application in detail with reference to the accompanying drawings. It can be understood that the specific embodiments described herein are only for explaining the present application and not for limiting the present application. Additionally, it should be noted that for the sake of convenience of description, only parts related to the present application are shown in the drawings rather than all the content. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. When the operations are completed, the process can be terminated, but there may also be additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0046] Embodiment 1:

[0047] Figure 1 A flowchart of a real-time acquisition method for pulse code modulation data provided by Embodiment 1 of the present application is given. The real-time acquisition method for pulse code modulation data provided in this embodiment can be executed by a real-time acquisition device for pulse code modulation data. The real-time acquisition device for pulse code modulation data can be implemented in software and / or hardware. The real-time acquisition device for pulse code modulation data can be composed of two or more physical entities or one physical entity. Generally speaking, the real-time acquisition device for pulse code modulation data can be a processing device such as an audio coding server or an audio processor.

[0048] The following takes the real-time acquisition device for pulse code modulation data as the main body for executing the real-time acquisition method for pulse code modulation data for description. Refer to Figure 1 , the real-time acquisition method for pulse code modulation data specifically includes:

[0049] S110. Real-time obtain the pulse code modulation data of multiple audio channels, calculate the signal-to-noise ratio, spectral flatness, and channel coherence information of each audio channel based on the pulse code modulation data, perform normalization processing on the signal-to-noise ratio, spectral flatness, and channel coherence information, and perform weighted summation on the normalized signal-to-noise ratio, spectral flatness, and channel coherence information to obtain the comprehensive weight of each audio channel.

[0050] When this application collects and outputs pulse-code modulation data of multiple audio channels, it first captures audio information of different sound sources and spatial positions in real time through a multi-channel audio acquisition device (such as a microphone array). In audio frameworks such as ALSA, PCM data acquisition of 8 channels or more channels is supported, and each channel independently performs signal sampling, quantization, and encoding to generate PCM data. Multi-channel acquisition can capture sound field information more comprehensively, improve the audio spatial perception ability, and provide a data basis for subsequent dynamically screening high-quality channels.

[0051] Furthermore, for the pulse-code modulation data of each collected audio channel, the signal-to-noise ratio, spectral flatness, and channel coherence information of each audio channel are first calculated based on the pulse-code modulation data. Among them, the signal-to-noise ratio (SNR) is used to quantify the power ratio of the signal to the noise and measure the quality of the channel signal. The spectral flatness is used to evaluate the uniformity of the signal spectral energy distribution and reflect the audio frequency response characteristics. The channel coherence is used to measure the correlation of signals between channels in the frequency domain and judge the signal synchronization.

[0052] Specifically, the calculation formula for the signal-to-noise ratio is expressed as:

[0053] SNR = 10·log 10 (P signal / P noise )

[0054] where SNR represents the signal-to-noise ratio of the audio channel, P signal represents the signal power of the audio channel, and P noise represents the noise power of the audio channel;

[0055] The calculation formula for the spectral flatness is expressed as:

[0056]

[0057] where Flatness Factor represents the spectral flatness of the audio channel, y i represents the energy value of the frequency point i in the pulse-code modulation data of the audio channel, and N represents the number of frequency sampling points;

[0058] The channel coherence information represents the mean of the channel correlations between the corresponding audio channel and other audio channels, and the calculation formula for the channel coherence is expressed as:

[0059]

[0060] where represents the channel coherence between two channels, C xy (f) represents the cross-spectrum between two channels, and P xx (f)Pyy (f) represents the power spectra of two channels.

[0061] By calculating these parameters in real time, the audio quality of each channel can be comprehensively evaluated. It provides an objective quantitative index for dynamically screening channels, avoiding the subjectivity and uncertainty of manual intervention.

[0062] Based on the relevant parameters obtained from the above calculations, a comprehensive weight is obtained through normalization and weighted summation. Among them, the SNR, spectral flatness, and channel coherence information parameters are linearly mapped to the [0,1] interval to eliminate the dimension difference.

[0063] Furthermore, the calculation formula of the comprehensive weight is expressed as:

[0064] W = w1·SNR norm + w2·Flatness norm + w3·Coherence norm

[0065] Where, W represents the comprehensive weight, w1, w2, w3 represent the corresponding weight coefficients, SNR norm represents the normalized parameter of the signal-to-noise ratio, Flatness norm represents the normalized parameter of the spectral flatness, Coherence norm represents the normalized parameter of the channel coherence information.

[0066] By calculating the SNR, spectral flatness, and channel coherence information of each channel, and performing normalization and weighted summation on them, the comprehensive weight W can be obtained. This weight reflects the comprehensive performance of the audio channel in terms of signal quality, spectral uniformity, and correlation, providing a quantitative basis for dynamically selecting the mixing channel.

[0067] Optionally, after real-time obtaining the pulse code modulation data of multiple audio channels, an adaptive filtering step can be added to perform denoising and enhancement processing on the original audio data.

[0068] Among them, the least mean square (LMS) algorithm is used to estimate the background noise in the channel in real time. According to the noise estimation result, the coefficients of the adaptive filter are dynamically adjusted to filter out the noise components. The gain of the filtered signal is adjusted to improve the signal quality. In this way, the signal-to-noise ratio of the original audio signal can be improved, providing cleaner data for subsequent parameter calculations.

[0069] In one embodiment, when calculating the signal-to-noise ratio, spectral flatness, and channel coherence information, a machine learning model (such as LSTM or CNN) can be introduced for parameter optimization.

[0070] Among them, historical audio data is used to train the model to learn the relationship between parameters and audio quality in different scenarios. In the current acquisition cycle, the original audio data is input into the trained model to predict the optimized parameter values. Furthermore, the model-predicted parameters and traditional calculation parameters are weighted and fused to obtain the final parameters. Thereby improving the accuracy and robustness of parameter calculation and adapting to complex and changing audio environments.

[0071] Optionally, when performing normalization processing and weighted summation to obtain the comprehensive weight, a scene classifier is introduced to dynamically adjust the weight coefficient according to the scene type.

[0072] Among them, the current scene is classified (such as meeting, performance, voice call) using audio features (such as energy, spectral distribution). According to the scene type, the corresponding coefficients are selected from the predefined weight coefficient library for weighted summation. If the scene changes, reclassification and adjustment of the weight coefficient are performed. Thereby making the comprehensive weight more in line with the audio requirements of different scenarios and improving the mixing effect.

[0073] S120. Based on the comprehensive weight, several audio channels are selected from multiple audio channels as the target mixing channel set for the current acquisition cycle. Based on the comprehensive weights of the individual audio channels in the target mixing channel set, the set adaptive attenuation coefficient, and the weak alignment coefficient, weighted mixing of the pulse code modulation data of the individual audio channels in the target mixing channel set for the current acquisition cycle is performed to obtain the mixed audio.

[0074] Furthermore, based on this comprehensive weight, multiple audio channels can be screened. Selecting several audio channels from multiple audio channels as the target mixing channel set for the current acquisition cycle includes:

[0075] Selecting the audio channels with comprehensive weights higher than the set threshold from multiple audio channels as the target mixing channel set for the current acquisition cycle; or,

[0076] Selecting the set number of audio channels with the highest comprehensive weights from multiple audio channels as the target mixing channel set for the current acquisition cycle.

[0077] For example, in an 8-channel system, 4 - 6 high-quality channels can be dynamically selected. By selecting the channels with weights higher than the threshold or the top N channels with the highest weights as the target mixing channel set, low-quality channels (such as channels affected by noise) can be dynamically screened out, improving the overall audio quality.

[0078] As Figure 2 shown, in an 8-channel system, by determining the target mixing set through audio channel screening, mixing processing can be performed on the target mixing set, thereby outputting the mixed audio.

[0079] Among them, referring to Figure 3, based on the comprehensive weights of each audio channel in the target mixing channel set, the set adaptive attenuation coefficient, and the weak alignment coefficient, perform weighted mixing on the pulse code modulation data of each audio channel in the target mixing channel set for the current acquisition period to obtain a mixed audio, including:

[0080] S1201. Configure the initial weights of each audio channel based on the comprehensive weights of each audio channel in the target mixing channel set. The comprehensive weight and the initial weight are directly proportional.

[0081] S1202. Adjust the initial weights based on the set adaptive attenuation coefficient and weak alignment coefficient of each audio channel in the target mixing channel set to obtain the mixing weights, and perform weighted mixing on the pulse code modulation data of each audio channel in the target mixing channel set for the current acquisition period according to the mixing weights to obtain the mixed audio.

[0082] In the multi-channel weighted mixing process of audio processing, introducing the adaptive attenuation coefficient α and the weak alignment coefficient β can significantly improve the mixing quality and the alignment accuracy of signals between channels. Weighted mixing is to superimpose the signals of multiple audio channels according to the weights calculated in real time to generate a mixed output signal. The weight allocation needs to consider parameters such as the signal-to-noise ratio (SNR), spectral flatness, and channel coherence of each channel (i.e., the above-mentioned comprehensive weights) to ensure the quality and adaptability of the mixing effect.

[0083] Based on this, the present application first configures the initial weights of each audio channel based on the comprehensive weights of each audio channel in the target mixing channel set. The sum of the initial weights is 1. The larger the comprehensive weight, the larger the initial weight. The initial weights of each audio channel can be mapped by following the ratio of the comprehensive weights of each audio channel.

[0084] Furthermore, the adaptive attenuation coefficient is calculated based on the phase difference between audio channels and is used to adjust the attenuation degree of the corresponding audio channel. The weak alignment coefficient is calculated based on the time delay difference between audio channels and is used to compensate for the signal time delay of the corresponding audio channel.

[0085] The adaptive attenuation coefficient α is used to dynamically adjust the attenuation degree of the channel signal to prevent signal overflow and distortion. It is dynamically adjusted according to the phase difference between channels:

[0086] Among them, when the phase difference is large: the signal correlation is low, then increase α to reduce the contribution of the signal of this channel and avoid noise and distortion.

[0087] When the phase difference is small: the signal correlation is high, then decrease α to increase the contribution of the signal of this channel and enhance the mixing effect.

[0088] The calculation of α is usually based on the phase difference Δφ between channels:

[0089] α = e -k·|Δφ|

[0090] where k is the attenuation constant, which controls the sensitivity of α to the change in phase difference.

[0091] The weak alignment coefficient β is used to compensate for the time delay difference between channels, ensure signal alignment in time, and avoid echo and phase cancellation. It is dynamically adjusted according to the time delay difference Δt between channels:

[0092] where, if the time delay difference is large: the signal synchronization is poor, then β is increased to compensate for the signal of the channel with a large time delay.

[0093] If the time delay difference is small: the signal synchronization is good, then β is decreased to reduce the compensation and keep the mixing natural.

[0094] The calculation of β is usually based on the time delay difference Δt between channels:

[0095]

[0096] where c is the alignment constant, which controls the sensitivity of β to the change in time delay difference.

[0097] Through the above calculation method, the adaptive attenuation coefficient and weak alignment coefficient of each audio channel can be determined. Then, based on the adaptive attenuation coefficient and weak alignment coefficient set for each audio channel in the target mixing channel set, the initial weight is adjusted to obtain the mixing weight, and the pulse code modulation data of each audio channel in the target mixing channel set in the current acquisition period is weighted and mixed according to the mixing weight to obtain the mixed audio.

[0098] The formula for the mixed audio is expressed as:

[0099]

[0100] where x i (t) is the signal of the i-th channel, τ i is the time delay compensation of the i-th channel, and ω i is the initial weight of the i-th channel. By introducing the adaptive attenuation coefficient α and weak alignment coefficient β, the weighted mixing process can dynamically adapt to the phase difference and time delay difference between channels, significantly improving the quality and adaptability of the mixing effect.

[0101] Optionally, before screening the target mixing channel set, a channel quality prediction step can also be added to predict the change trend of the channel quality in the future period.

[0102] Among them, a time series model is established by analyzing the quality parameters of each channel within a historical period. The signal-to-noise ratio, spectral flatness, and channel coherence of each channel in the future period are predicted using the model. Based on the prediction results, the set of target mixed audio channels can be screened in advance. Thereby, frequent channel switching is avoided, and the system stability and audio quality are improved.

[0103] S130. Determine the audio channels to be switched in the historical mixed audio channel set of the previous acquisition period relative to the set of target mixed audio channels. During the process of outputting the mixed audio in the current acquisition period, mix the mixed audio of the previous set number of frames with the pulse code modulation data of the audio channels to be switched and output the mixture. During the mixing output process, linearly adjust the gain of the audio channels to be switched from the historical value to 0.

[0104] Furthermore, based on the obtained mixed audio above, considering that there may be switching of audio channels in different audio acquisition periods, that is, the audio channels screened in the previous and subsequent acquisition periods are different. In order to achieve smooth switching of audio channels, the present application further performs smooth processing of audio channel switching when outputting the mixed audio. By comparing the set of target mixed audio channels in the current and previous periods, find out the removed channels, that is, the audio channels to be switched.

[0105] When outputting the mixed audio, mix the mixed audio of the previous set number of frames (such as 100 ms) with the original signal of the channel to be switched and output the mixture, while linearly reducing the gain of the channel to be switched to 0. After the set number of frames, directly transmit the mixed audio. This can avoid audio mutations during channel switching, enhance the user experience, and ensure the continuity and naturalness of audio output.

[0106] Based on the above scheme, by dynamically screening channels, excluding low-quality channels based on the real-time calculated audio quality parameters, the overall audio quality is improved; through adaptive mixing, dynamically adjusted based on the attenuation and alignment coefficients, the mixing effect can be optimized, and distortion and echo can be avoided; through smooth transition of audio channel switching, seamless transition can be achieved during channel switching, and the user experience can be enhanced. In addition, by introducing objective parameters and weight calculation, the subjectivity of manual intervention can be avoided, and the system robustness can be improved.

[0107] Optionally, when performing weighted mixing, a non-linear mixing algorithm (such as a mixing model based on an artificial neural network) can also be introduced. Among them, the non-linear mixing model is trained using high-quality audio data to learn the mapping relationship from audio features to mixing output. In the current acquisition period, input the audio data of the set of target mixed audio channels into the trained model to generate the mixed audio. Then, post-process the model output according to the adaptive attenuation coefficient and weak alignment coefficient. Thereby, the naturalness and clarity of the mixed audio are enhanced, and the distortion and artificial traces brought by traditional linear mixing are avoided.

[0108] In addition, after determining the audio channel to be switched, a non-linear gain adjustment algorithm (such as exponential decay or logarithmic decay) can be used to replace the linear gain adjustment to achieve smoother gain changes.

[0109] Among them, a non-linear function is selected according to the audio characteristics, and then for each frame of the transition frame of the audio channel to be switched, the current gain value is calculated according to the non-linear function. The current gain value is applied to the signal of the audio channel to be switched and superimposed with the mixed audio for output. Non-linear gain adjustment can more naturally reduce the volume of the switched channel, avoid the abruptness caused by linear adjustment, and improve the user experience.

[0110] On the other hand, when switching channels, a phase alignment compensation step can also be added to compensate for the phase difference between the switched audio channel and the mixed audio.

[0111] Among them, the cross-correlation function or phase spectrum analysis is used to estimate the phase difference between the switched channel and the mixed audio. According to the estimated phase difference, the signal of the switched channel is phase-rotated or time-delay adjusted to align it with the mixed audio. In each frame of the transition frame, the phase difference estimation is updated in real time and compensated. Phase alignment compensation can avoid audio distortion or echo caused by phase difference during channel switching and improve the clarity and naturalness of the mixed audio.

[0112] As described above, with reference to Figure 4, by obtaining the pulse code modulation data of multiple audio channels in real time, calculating the signal-to-noise ratio, spectral flatness, and channel coherence information of each audio channel based on the pulse code modulation data, normalizing the signal-to-noise ratio, spectral flatness, and channel coherence information, and performing weighted summation on the normalized signal-to-noise ratio, spectral flatness, and channel coherence information to obtain the comprehensive weight of each audio channel; screening several audio channels from multiple audio channels as the target mixing channel set for the current acquisition period based on the comprehensive weight, and performing weighted mixing on the pulse code modulation data of each audio channel in the target mixing channel set based on the comprehensive weight of each audio channel in the target mixing channel set, the set adaptive attenuation coefficient, and the weak alignment coefficient to obtain the mixed audio; determining the switched audio channels in the historical mixing channel set of the previous acquisition period relative to the target mixing channel set, and during the process of outputting the mixed audio in the current acquisition period, mixing and outputting the mixed audio of the previous set number of frames with the pulse code modulation data of the switched audio channels, and during the mixing output process, linearly adjusting the gain of the switched audio channel from the historical value to 0. By adopting the above technical means, dynamic screening of the pulse code modulation data is achieved by screening audio channels for weighted mixing, thereby improving the mixing effect and audio quality and enhancing the user's listening experience. And during the process of outputting the mixed audio, the switched audio channel is added to the mixed audio, and the gain of the switched audio channel is gradually linearly adjusted from the historical value to 0, thereby ensuring smooth transition of the audio time-domain signal and volume during audio channel switching, making the audio time-domain waveform continuous and the audio volume stable, ensuring the audio output effect, and further enhancing the user's listening experience.

[0113] Embodiment 2:

[0114] Based on the above embodiment, Figure 5 is a schematic structural diagram of a real-time acquisition system for pulse code modulation data provided by Embodiment 2 of this application. Refer to Figure 5 , the real-time acquisition system for pulse code modulation data provided in this embodiment specifically includes:

[0115] Among them, the calculation module 21 is used to obtain the pulse code modulation data of multiple audio channels in real time, calculate the signal-to-noise ratio, spectral flatness, and channel coherence information of each audio channel based on the pulse code modulation data, normalize the signal-to-noise ratio, spectral flatness, and channel coherence information, and perform weighted summation on the normalized signal-to-noise ratio, spectral flatness, and channel coherence information to obtain the comprehensive weight of each audio channel;

[0116] The mixing module 22 is configured to screen several audio channels from multiple audio channels as the target mixing channel set for the current acquisition period based on the comprehensive weight, and perform weighted mixing on the pulse code modulation data of each audio channel in the target mixing channel set based on the comprehensive weight of each audio channel in the target mixing channel set, the set adaptive attenuation coefficient, and the weak alignment coefficient to obtain the mixed audio;

[0117] The output module 23 is configured to determine the switched audio channels in the historical mixing channel set of the previous acquisition period relative to the target mixing channel set, and during the process of outputting the mixed audio in the current acquisition period, mix and output the mixed audio of the previous set number of frames with the pulse code modulation data of the switched audio channels, and linearly adjust the gain of the switched audio channels from the historical value to 0 during the mixing and output process.

[0118] Specifically, the signal-to-noise ratio calculation formula is expressed as:

[0119] SNR = 10·log 10 (P signal / P noise )

[0120] where, SNR represents the signal-to-noise ratio of the audio channel, P signal represents the signal power of the audio channel, and P noise represents the noise power of the audio channel;

[0121] The spectral flatness calculation formula is expressed as:

[0122]

[0123] where, Flatness Factor represents the spectral flatness of the audio channel, y i represents the energy value of the frequency point i in the pulse code modulation data of the audio channel, and N represents the number of frequency sampling points;

[0124] The channel coherence information represents the mean of the channel correlations between the corresponding audio channel and other audio channels, and the channel coherence calculation formula is expressed as:

[0125]

[0126] where, represents the channel coherence between two channels, C xy (f) represents the cross-spectrum between two channels, and P xx (f)P yy (f) represents the power spectra of two channels.

[0127] Specifically, the calculation formula of the comprehensive weight is expressed as:

[0128] W = w1·SNR norm + w2·Flatness norm + w3·Coherence norm

[0129] Wherein, W represents the comprehensive weight, w1, w2, and w3 represent the corresponding weight coefficients, and SNR norm represents the normalized parameter of the signal-to-noise ratio, and Flatness norm represents the normalized parameter of the spectral flatness, and Coherence norm represents the normalized parameter of the channel coherence information.

[0130] Specifically, based on the comprehensive weight, several audio channels are selected from multiple audio channels as the target mixing channel set for the current acquisition period, including:

[0131] Selecting the audio channels with a comprehensive weight higher than the set threshold from multiple audio channels as the target mixing channel set for the current acquisition period; or,

[0132] Selecting a set number of audio channels with the highest comprehensive weight from multiple audio channels as the target mixing channel set for the current acquisition period.

[0133] Specifically, based on the comprehensive weight of each audio channel in the target mixing channel set, the set adaptive attenuation coefficient, and the weak alignment coefficient, weighted mixing of the pulse code modulation data of each audio channel in the target mixing channel set for the current acquisition period is performed to obtain the mixed audio, including:

[0134] Configuring the initial weight of each audio channel based on the comprehensive weight of each audio channel in the target mixing channel set, and the comprehensive weight is proportional to the initial weight;

[0135] Adjusting the initial weight based on the set adaptive attenuation coefficient and weak alignment coefficient of each audio channel in the target mixing channel set to obtain the mixing weight, and performing weighted mixing of the pulse code modulation data of each audio channel in the target mixing channel set for the current acquisition period according to the mixing weight to obtain the mixed audio.

[0136] Specifically, the adaptive attenuation coefficient is calculated based on the phase difference between audio channels and is used to adjust the attenuation degree of the corresponding audio channel.

[0137] Specifically, the weak alignment coefficient is calculated based on the time delay difference between audio channels and is used to compensate for the signal time delay of the corresponding audio channel.

[0138] As described above, by obtaining the pulse code modulation data of multiple audio channels in real time, calculating the signal-to-noise ratio, spectral flatness, and channel coherence information of each audio channel based on the pulse code modulation data, normalizing the signal-to-noise ratio, spectral flatness, and channel coherence information, and performing weighted summation on the normalized signal-to-noise ratio, spectral flatness, and channel coherence information to obtain the comprehensive weight of each audio channel; screening several audio channels from multiple audio channels as the target mixing channel set for the current acquisition period based on the comprehensive weight, and performing weighted mixing on the pulse code modulation data of each audio channel in the target mixing channel set based on the comprehensive weight of each audio channel in the target mixing channel set, the set adaptive attenuation coefficient, and the weak alignment coefficient to obtain the mixed audio; determining the switched audio channels in the historical mixing channel set of the previous acquisition period relative to the target mixing channel set, and during the process of outputting the mixed audio in the current acquisition period, mixing and outputting the mixed audio of the previous set number of frames with the pulse code modulation data of the switched audio channels, and during the mixing output process, linearly adjusting the gain of the switched audio channel from the historical value to 0. By adopting the above technical means, dynamic screening of the pulse code modulation data is realized by screening audio channels for weighted mixing, thereby improving the mixing effect and audio quality and enhancing the user's listening experience. And during the process of outputting the mixed audio, the switched audio channel is added to the mixed audio, and the gain of the switched audio channel is gradually linearly adjusted from the historical value to 0, thereby ensuring smooth transition of the audio time-domain signal and volume during the audio channel switching, making the audio time-domain waveform continuous and the audio volume stable, ensuring the audio output effect, and further enhancing the user's listening experience.

[0139] The real-time acquisition system for pulse code modulation data provided in Embodiment 2 of the present application can be used to execute the real-time acquisition method for pulse code modulation data provided in Embodiment 1 above, and has the corresponding functions and beneficial effects.

[0140] Embodiment 3:

[0141] Embodiment 3 of the present application provides an electronic device. Referring to Figure 6 , the electronic device includes: a processor 31, a memory 32, a communication module 33, an input device 34, and an output device 35. The number of processors in the electronic device can be one or more, and the number of memories in the electronic device can be one or more. The processor, memory, communication module, input device, and output device of the electronic device can be connected through a bus or other means.

[0142] As a computer-readable storage medium, the memory can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the real-time acquisition method of pulse code modulation data described in any embodiment of this application (for example, each module in the real-time acquisition system of pulse code modulation data). The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the device, etc. In addition, the memory can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory can further include a memory remotely set relative to the processor, and these remote memories can be connected to the device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0143] The communication module is used for data transmission.

[0144] By running the software programs, instructions, and modules stored in the memory, the processor executes various functional applications and data processing of the device, that is, implements the above-mentioned real-time acquisition method of pulse code modulation data.

[0145] The input device can be used to receive input digital or character information and generate key signal inputs related to the user settings and function control of the device. The output device can include display devices such as a display screen.

[0146] The above-provided electronic device can be used to execute the real-time acquisition method of pulse code modulation data provided in the first embodiment above, and has corresponding functions and beneficial effects.

[0147] Embodiment 4:

[0148] The embodiment of the present application further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute a real-time acquisition method for pulse code modulation data when executed by a computer processor. The real-time acquisition method for pulse code modulation data includes: acquiring in real time the pulse code modulation data of multiple audio channels, calculating the signal-to-noise ratio, spectral flatness, and channel coherence information of each audio channel based on the pulse code modulation data, performing normalization processing on the signal-to-noise ratio, spectral flatness, and channel coherence information, and performing weighted summation on the normalized signal-to-noise ratio, spectral flatness, and channel coherence information to obtain the comprehensive weight of each audio channel; screening several audio channels from the multiple audio channels as the target mixing channel set for the current acquisition period based on the comprehensive weight, and performing weighted mixing on the pulse code modulation data of each audio channel in the target mixing channel set based on the comprehensive weight of each audio channel in the target mixing channel set, the set adaptive attenuation coefficient, and the weak alignment coefficient to obtain the mixed audio; determining the switched audio channels in the historical mixing channel set of the previous acquisition period relative to the target mixing channel set, and during the process of outputting the mixed audio in the current acquisition period, mixing and outputting the mixed audio of the previous set number of frames with the pulse code modulation data of the switched audio channels, and during the mixing output process, linearly adjusting the gain of the switched audio channel from the historical value to 0.

[0149] Storage medium - any of various types of memory devices or storage devices. The term "storage medium" is intended to include: installation media such as CD-ROMs, floppy disks, or magnetic tape devices; computer system memory or random access memory such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory such as flash memory, magnetic media (such as hard disks or optical storage); registers or other similar types of memory elements, etc. The storage medium may also include other types of memory or combinations thereof. Additionally, the storage medium may be located in a first computer system in which the program is executed, or may be located in a different second computer system that is connected to the first computer system via a network (such as the Internet). The second computer system may provide program instructions to the first computer for execution. The term "storage medium" may include two or more storage media residing in different locations (such as in different computer systems connected via a network). The storage medium may store program instructions (such as embodied as a computer program) executable by one or more processors.

[0150] Of course, for a storage medium containing computer-executable instructions provided by the embodiment of the present application, the computer-executable instructions are not limited to the real-time acquisition method for pulse code modulation data as described above, and may also execute related operations in the real-time acquisition method for pulse code modulation data provided by any embodiment of the present application.

[0151] The real-time acquisition system, storage medium, and electronic device for pulse code modulation data provided in the above embodiments can execute the real-time acquisition method for pulse code modulation data provided in any embodiment of the present application. For technical details not described in detail in the above embodiments, reference can be made to the real-time acquisition method for pulse code modulation data provided in any embodiment of the present application.

[0152] The above is only the preferred embodiment of the present application and the technical principles applied. The present application is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions that can be made by those skilled in the art will not depart from the protection scope of the present application. Therefore, although the present application has been described in detail through the above embodiments, the present application is not limited to the above embodiments. Without departing from the concept of the present application, more other equivalent embodiments can be included, and the scope of the present application is determined by the scope of the claims.

Claims

1. A real-time acquisition method for pulse code modulation data, characterized in that Including: Obtaining pulse code modulation data of multiple audio channels in real time, calculating the signal-to-noise ratio, spectral flatness, and channel coherence information of each audio channel based on the pulse code modulation data, normalizing the signal-to-noise ratio, spectral flatness, and channel coherence information, and performing weighted summation on the normalized signal-to-noise ratio, spectral flatness, and channel coherence information to obtain the comprehensive weight of each audio channel; Based on the comprehensive weight, screening several audio channels from multiple audio channels as the target mixing channel set for the current acquisition period, and performing weighted mixing on the pulse code modulation data of each audio channel in the target mixing channel set for the current acquisition period based on the comprehensive weight of each audio channel in the target mixing channel set, the set adaptive attenuation coefficient, and the weak alignment coefficient to obtain the mixed audio; Determining the switched audio channels in the historical mixing channel set of the previous acquisition period relative to the target mixing channel set, during the process of outputting the mixed audio in the current acquisition period, mixing and outputting the mixed audio of the previous set number of frames with the pulse code modulation data of the switched audio channels, and linearly adjusting the gain of the switched audio channels from the historical value to 0 during the mixing and output process.

2. The real-time acquisition method of pulse-coded modulation data according to claim 1, characterized in that The formula for calculating the signal-to-noise ratio is expressed as: SNR = 10·log 10 (P signal / P noise ) Among them, SNR represents the signal-to-noise ratio of the audio channel, and P signal represents the signal power of the audio channel, and P noise represents the noise power of the audio channel; The formula for calculating the spectral flatness is expressed as: Among them, Flatness Factor represents the spectral flatness of the audio channel, and y i represents the energy value of the frequency point i in the pulse code modulation data of the audio channel, and N represents the number of frequency sampling points; The channel coherence information represents the mean of the channel correlations between the corresponding audio channel and other audio channels, and the formula for calculating the channel coherence is expressed as: Among them, represents the channel coherence between two channels, C xy (f) represents the cross-spectrum between two channels, P xx (f)P yy (f) represents the power spectra of two channels.

3. The real-time acquisition method of pulse code modulation data according to claim 2, characterized in that, The formula for calculating the comprehensive weight is expressed as: W = w1·SNR norm + w2·Flatness norm + w3·Coherence norm Among them, W represents the comprehensive weight, w1, w2, w3 represent the corresponding weight coefficients, SNR norm represents the normalized parameter of the signal-to-noise ratio, Flatness norm represents the normalized parameter of the spectral flatness, Coherence norm represents the normalized parameter of the channel coherence information.

4. The real-time acquisition method of pulse code modulation data according to claim 1, characterized in that The screening of several audio channels from multiple audio channels as the target mixing channel set for the current acquisition period based on the comprehensive weight includes: Screening the audio channels with a comprehensive weight higher than the set threshold from multiple audio channels as the target mixing channel set for the current acquisition period; or, Screening the set number of audio channels with the highest comprehensive weight from multiple audio channels as the target mixing channel set for the current acquisition period.

5. The real-time acquisition method of pulse code modulation data according to claim 1, characterized in that, The performing of weighted mixing on the pulse code modulation data of each audio channel in the target mixing channel set for the current acquisition period based on the comprehensive weight of each audio channel in the target mixing channel set, the set adaptive attenuation coefficient, and the weak alignment coefficient to obtain the mixed audio includes: Configuring the initial weight of each audio channel based on the comprehensive weight of each audio channel in the target mixing channel set, and the comprehensive weight is proportional to the initial weight; Adjusting the initial weight based on the set adaptive attenuation coefficient and weak alignment coefficient of each audio channel in the target mixing channel set to obtain the mixing weight, and performing weighted mixing on the pulse code modulation data of each audio channel in the target mixing channel set for the current acquisition period according to the mixing weight to obtain the mixed audio.

6. The real-time acquisition method of pulse code modulation data according to claim 5, characterized in that The adaptive attenuation coefficient is calculated based on the phase difference between audio channels and is used to adjust the attenuation degree of the corresponding audio channel.

7. The real-time acquisition method of pulse-coded modulation data according to claim 5, characterized in that, The weak alignment coefficient is calculated based on the time delay difference between audio channels and is used to compensate for the signal time delay of the corresponding audio channel.

8. A real-time acquisition system for pulse code modulation data, characterized in that, Including: A calculation module, configured to obtain pulse code modulation data of multiple audio channels in real time, calculate the signal-to-noise ratio, spectral flatness, and channel coherence information of each audio channel based on the pulse code modulation data, perform normalization processing on the signal-to-noise ratio, the spectral flatness, and the channel coherence information, and perform weighted summation on the normalized signal-to-noise ratio, the spectral flatness, and the channel coherence information to obtain the comprehensive weight of each audio channel; A mixing module, configured to screen several audio channels from the multiple audio channels as the target mixing channel set for the current acquisition period based on the comprehensive weight, and perform weighted mixing on the pulse code modulation data of each audio channel in the target mixing channel set based on the comprehensive weight of each audio channel in the target mixing channel set, a set adaptive attenuation coefficient, and a weak alignment coefficient to obtain a mixed audio; An output module, configured to determine the switched audio channels in the historical mixing channel set of the previous acquisition period relative to the target mixing channel set, and during the output of the mixed audio in the current acquisition period, mix and output the mixed audio of the previous set number of frames with the pulse code modulation data of the switched audio channels, and during the mixed output process, linearly adjust the gain of the switched audio channels from the historical value to 0.

9. An electronic device, characterized in that, Comprising: A memory and one or more processors; The memory is configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the real-time acquisition method of pulse code modulation data as described in any one of claims 1-7.

10. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions are used to execute the real-time acquisition method of pulse code modulation data as described in any one of claims 1-7 when executed by a computer processor.