Multi-channel audio transmission method and system and storage medium

Through the multi-channel audio transmission system, synchronous multi-channel ADC and deep audio analysis are used to solve the problem of dynamic adjustment in multiple audio transmission, achieving efficient audio signal recovery and stable transmission, and improving audio quality and user experience.

CN120452461AInactive Publication Date: 2025-08-08JIAXING WANSHENG ELECTRONICS TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510637568.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-08-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In multiple audio transmission, the prior art is difficult to dynamically adjust in complex network environments, resulting in loss or distortion of audio signals, affecting semantic coherence and recovery quality.

Method used

By synchronous multi-channel ADC system, audio signals are collected, short-time Fourier transform and feature extraction are performed, channel load data sets and damage data sets are obtained, channel load coupling score index and frame structure damage depth index are calculated, deep audio analysis instructions are performed, and audio frame recovery evaluation is carried out in a comprehensive scheduling and reconstruction index.

Benefits of technology

Improves the robustness and stability of audio transmission, ensures high-quality recovery and lossless transmission, reduces the risk of frame loss, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452461A_ABST
    Figure CN120452461A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-channel audio transmission method and system and a storage medium, and relates to the technical field of audio transmission, and the method comprises the steps: collecting voice signals of different audio sources on a plurality of input channels through a synchronous multi-channel ADC system, and obtaining an audio time-frequency spectrum through short-time Fourier transform; an audio feature data set is extracted and preprocessed, and a channel load data set, a damage data set and a recovery data set are generated. And calculating a channel load coupling scoring index tfo according to the channel load data set, performing transmission channel matching evaluation, and if a frame loss risk is detected, executing a deep audio analysis instruction. And further evaluating the recovery demand of the audio frame by calculating a comprehensive scheduling reconstruction index ddc, and selecting a recovery scheme, such as semantic structure recovery or interpolation recovery, according to an evaluation result. According to the method, the audio transmission quality and the recovery efficiency are optimized, and the frame loss risk is effectively reduced and the semantic consistency of the audio signals is improved in the multi-channel audio transmission process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio transmission, and in particular to a multi-channel audio transmission method, system and storage medium. Background Art

[0002] In the fields of modern communications and audio processing, with the continuous advancement of technology, multichannel transmission of audio signals has become a trend that cannot be ignored. The advantage of multichannel transmission is that it can process multiple signal sources simultaneously to meet the needs of different applications. However, with the increase in parallel processing of multiple signals, the stability and quality of audio transmission have become a major challenge. Especially in complex transmission environments, signals may be attenuated, lost, or distorted due to channel load, noise, or other factors. To ensure the integrity and quality of audio signals, how to effectively match channels and process audio data in multichannel environments is a key research topic. Therefore, designing a system that can assess the health of audio transmission channels and recover audio frames has become the core of solving this problem.

[0003] Although multi-channel audio transmission technology has made great progress, many problems still exist in real-world applications. Traditional audio recovery methods are often not easy to adjust dynamically when faced with complex network environments. For example, some methods rely on a fixed frame repair mechanism, which may not be easy to effectively adjust according to real-time network conditions in a dynamically changing network environment. When audio frames are lost or signal frames are lost, traditional methods often lead to poor recovery effects and may even lose the semantic coherence of the speech content. In addition, many audio recovery algorithms do not fully consider the characteristics of the audio signal, resulting in the recovered audio quality not being easy to achieve ideal, especially in terms of semantic consistency and audio coherence, which may affect the user's audio experience. Summary of the Invention

[0004] In view of the deficiencies in the prior art, the present invention provides a multi-channel audio transmission method, system and storage medium, which solve the problems in the above-mentioned background technology.

[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: A multi-channel audio transmission method, comprising the following steps:

[0006] S1. Speech signals from different audio sources are collected on multiple input channels through a synchronous multi-channel ADC system and divided into several audio frames. The time-domain signals are then converted into time-frequency domain signals through short-time Fourier transform to obtain the audio time-frequency spectrum.

[0007] S2. Extract features from the audio time-frequency spectrum to obtain an audio feature dataset, and then preprocess the audio dataset to obtain a channel load dataset, a damage dataset, and a restored dataset;

[0008] S3. Calculate based on the channel load data set to obtain a channel load coupling score index tfo for transmission channel matching assessment, and execute a deep audio analysis instruction if the assessment indicates that there is a risk of frame loss in the audio transmission channel;

[0009] S4, deep audio analysis instructions are used to calculate and obtain the frame structure damage depth index zss and semantic recovery credibility index yhf based on the damaged data group and the restored data group respectively;

[0010] S5. Based on the channel load coupling score index tfo, the frame structure damage depth index zss and the semantic recovery credibility index yhf, a summary calculation is performed to obtain the comprehensive scheduling reconstruction index ddc for audio frame recovery evaluation.

[0011] Preferably, S1 includes S11 and S12;

[0012] S11. The configured synchronous multi-channel ADC system collects voice signals from different audio sources on multiple input channels using a time-domain synchronous sampling method and divides the signals into several audio frames. The length of each audio frame is 20ms-40ms, and the specific length is determined by the audio processing requirements. The sampling rate of each audio frame is 48kHz, a common audio processing standard.

[0013] S12, denoising the acquired audio frame, dividing the denoised audio frame into frames according to the frame size, applying a window function to the audio frame signal of each frame to avoid spectrum leakage caused by signal truncation, and then converting the time domain signal into a time-frequency domain signal through short-time Fourier transform to obtain an audio time-frequency spectrum;

[0014] Denoising uses the wavelet denoising method to decompose the audio frame signal into sub-signals of different scales, suppress the noise in different frequency bands respectively, and filter out the noise influence in the audio frame.

[0015] Preferably, S2 includes S21 and S22;

[0016] S21, extracting features from the acquired audio time-frequency spectrum to obtain an audio feature dataset;

[0017] Perform a second-order difference calculation on the time domain waveform of each frame of audio signal based on the acquired audio time-frequency spectrum to obtain the second-order derivative energy density nm of the audio frame waveform;

[0018] The spectral entropy of each frame of the audio spectrum is calculated using the information entropy formula to obtain the spectral entropy of the audio frame. The spectral entropy change rate xb of adjacent frames is calculated based on the spectral entropy between each adjacent frame. The transmission and retransmission status of each audio frame is then tracked, and the entropy of the spectrum of each retransmitted frame is calculated to obtain the retransmission guided entropy loss rate ss.

[0019] The interference signal in the audio time spectrum is converted into a sparse signal using the compressed sensing method, and the local sparsity of the interference xs is obtained by calculating the sparsity of the sparse signal;

[0020] The phase of each frame signal of the audio time spectrum is extracted by Fourier transform, and the interference phase offset py is obtained by calculating the phase difference between adjacent audio frames;

[0021] Based on the amount of transmitted data in each frame in the audio time spectrum, the load similarity of the burst frame is calculated using correlation analysis to obtain the burst frame load coherence zf;

[0022] According to the transmission time domain of each audio frame, the difference in time intervals between adjacent frames is calculated to obtain the frame period non-uniformity zq;

[0023] The number of samples in each audio frame is obtained by multiplying the sampling rate of the audio frame and the duration of the audio frame. The square values of each audio frame are accumulated to obtain the audio frame energy. The difference in energy between adjacent frames is calculated based on the energy of each audio frame to obtain the frame energy gap qk.

[0024] Use natural language processing technology to extract the semantic information of each audio frame from the speech content in the audio, and obtain the semantic stability yw by comparing the semantic consistency of consecutive audio frames;

[0025] S22, performing outlier processing and dimensionless processing on the audio feature dataset to obtain a channel load data group, a damage data group, and a recovery data group;

[0026] Outlier processing is performed by using the interquartile range method to detect and process outliers in the audio feature dataset, and dimensionless processing is performed by using the Max-Min method to eliminate the dimensional influence of the audio feature dataset;

[0027] The channel load data set includes the second-order derivative energy density nm of the audio frame waveform, the spectral entropy change rate xb of adjacent frames, the interference local sparsity xs, the interference phase offset py and the burst frame load coherence zf;

[0028] The impairment data set includes frame period non-uniformity zq, retransmission guidance entropy loss rate ss and frame energy gap qk;

[0029] The restored data set includes the audio frame energy F and semantic stability yw.

[0030] Preferably, S3 includes S31;

[0031] S31, performing summary calculation based on the obtained channel load data group to obtain a channel load coupling score index tfo, evaluate the encoding adaptability of the audio stream in the current network environment, and select an appropriate encoding method;

[0032] ;

[0033] Where, e represents the exponential function and ln represents the logarithmic function.

[0034] Preferably, S3 further includes S32;

[0035] S32. Calculate the average of all historical channel load coupling score indices tfo using a statistical method based on the historical channel load coupling score indices tfo, and set a transmission channel risk threshold A based on the average. The threshold is then compared with the channel load coupling score indices tfo obtained in real time, and a transmission channel matching assessment is performed based on the comparison result. The specific assessment scheme is as follows;

[0036] When the channel load coupling score index tfo is greater than the transmission channel risk threshold A, it indicates that the audio and video transmission channel is healthy and Turbo coding is used for transmission, while real-time monitoring is maintained.

[0037] When the channel load coupling score index tfo ≤ the transmission channel risk threshold A, it means that there is a risk of frame loss in the audio transmission channel. At this time, low-density parity check code LDPC is used for transmission and deep audio analysis instructions are executed.

[0038] Preferably, S4, when the transmission channel matching assessment shows that there is a risk of frame loss in the audio transmission channel, executing the deep audio analysis instruction, specifically including S41 and S42;

[0039] S41, performing summary calculation based on the acquired damage data group to obtain a frame structure damage depth index zss;

[0040] ;

[0041] Where qk ref Represents the standard audio energy of each frame under the complete audio signal;

[0042] S42, performing summary calculation based on the obtained recovery data group to obtain a semantic recovery credibility index yhf;

[0043] ;

[0044] Where k represents the total number of context frames, F i represents the instantaneous energy of the i-th audio frame, Indicates the average energy of the audio frame within the acquisition interval, Var(F i ) represents the variance of the energy of the upper and lower audio frames of the i-th audio frame.

[0045] Preferably, S5 includes S51;

[0046] S51. Based on the obtained channel load coupling score index tfo, frame structure damage depth index zss and semantic recovery credibility index yhf, a summary calculation is performed to obtain a comprehensive scheduling reconstruction index ddc. The specific formula is as follows:

[0047] .

[0048] Preferably, S5 further includes S52;

[0049] S52: Preset an audio frame damage recovery threshold B according to the audio network interoperability industry standard and compare it with the obtained comprehensive scheduling reconstruction index ddc. Perform an audio frame recovery assessment based on the comparison result. The specific assessment scheme is as follows:

[0050] When the comprehensive scheduling reconstruction index ddc is greater than the audio frame damage recovery threshold B, it indicates that there is a risk of frame loss due to audio frame damage. In this case, the audio frame is classified as the first priority for recovery and the semantic structure recovery process is immediately executed.

[0051] When the comprehensive scheduling reconstruction index ddc ≤ the audio frame damage recovery threshold B, it means that there is no risk of frame loss due to audio frame damage. At this time, it is classified as the second priority recovery and interpolation repair is performed.

[0052] A multi-channel audio transmission system includes an audio acquisition module, a feature extraction module, a channel matching evaluation module, a depth analysis module and a comprehensive recovery evaluation module;

[0053] The audio acquisition module is used to collect voice signals from different audio sources on multiple input channels based on a synchronous multi-channel ADC system, divide them into several audio frames, and then convert the time domain signals into time-frequency domain signals through short-time Fourier transform to obtain the audio time-frequency spectrum;

[0054] The feature extraction module is used to extract features from the audio time-frequency spectrum to obtain an audio feature data set, and then preprocess the audio data set to obtain a channel load data set, a damage data set, and a recovery data set;

[0055] The channel matching evaluation module is used to calculate based on the channel load data group to obtain the channel load coupling score index TFO for transmission channel matching evaluation, and execute deep audio analysis instructions when it is assessed that there is a risk of frame loss in the audio transmission channel;

[0056] The deep analysis module is used to execute deep audio analysis instructions, and calculate the frame structure damage depth index zss and the semantic recovery credibility index yhf according to the damaged data group and the restored data group respectively;

[0057] The comprehensive recovery evaluation module is used to perform summary calculations based on the channel load coupling score index tfo, the frame structure damage depth index zss and the semantic recovery credibility index yhf, and obtain the comprehensive scheduling reconstruction index ddc for audio frame recovery evaluation.

[0058] A multi-channel audio transmission storage medium stores a computer program. When the computer program is executed, a multi-channel audio transmission method is implemented when the program is executed by a processor.

[0059] The present invention provides a multi-channel audio transmission method, system and storage medium. It has the following beneficial effects:

[0060] (1) This method uses a synchronous multi-channel ADC system to synchronously sample the audio signals of multiple input channels. After denoising, the time-domain signals are converted into time-frequency domain signals through short-time Fourier transform to obtain the audio time-frequency spectrum. Feature extraction is performed on the audio time-frequency spectrum to obtain an audio feature dataset. Then, outlier processing and dimensionless processing are performed on the audio feature dataset to obtain channel load data sets, damage data sets, and recovery data sets, providing an accurate data basis for subsequent analysis.

[0061] (2) This method obtains the channel load coupling score index tfo by summarizing and calculating the channel load data group, and uses this to evaluate the encoding adaptability of the current audio stream. It then performs a matching evaluation within the transmission channel based on the preset transmission channel risk threshold A, selects an appropriate encoding method, and evaluates the health of the audio transmission channel by comparing historical data to determine whether there is a risk of frame loss. When the evaluation results show that there is a risk of frame loss, the deep audio analysis instruction is triggered, and the frame structure damage depth index zss is calculated based on the damage data group. At the same time, the semantic recovery credibility index yhf is calculated based on the recovery data group, thereby determining the degree of damage to the audio frame and its recovery potential, providing a reference for further recovery.

[0062] (3) The method uses the comprehensive scheduling reconstruction index DDC to evaluate the recovery of audio frames. The comprehensive scheduling reconstruction index DDC is obtained by summarizing and calculating the channel load coupling score index TFO, the frame structure damage depth index ZSS, and the semantic recovery credibility index YHF. It can make nonlinear adjustments to the recovery of audio frames to ensure stable recovery under channel fluctuations. According to the preset audio frame damage recovery threshold B, the frame loss risk of the audio frame is evaluated. If the comprehensive scheduling reconstruction index DDC exceeds the audio frame damage recovery threshold B, the semantic structure recovery process is immediately executed. If the comprehensive scheduling reconstruction index DDC is lower than or equal to the audio frame damage recovery threshold B, interpolation repair is performed. Through this complete set of systematic processes, this method effectively improves the reliability and robustness of audio transmission, ensuring high-quality recovery and lossless transmission of audio signals. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 A schematic diagram of the steps of a multi-channel audio transmission method of the present invention;

[0064] Figure 2 This is a schematic diagram of a multi-channel audio transmission system flow of the present invention;

[0065] Figure 3 Schematic diagram of the audio frame recovery evaluation line of the present invention. DETAILED DESCRIPTION

[0066] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0067] Example 1

[0068] See also Figure 1 The present invention provides a multi-channel audio transmission method. To achieve the above purpose, the present invention is implemented through the following technical solutions: comprising the following steps:

[0069] S1. Speech signals from different audio sources are collected on multiple input channels through a synchronous multi-channel ADC system and divided into several audio frames. The time-domain signals are then converted into time-frequency domain signals through short-time Fourier transform to obtain the audio time-frequency spectrum.

[0070] S2. Extract features from the audio time-frequency spectrum to obtain an audio feature dataset, and then preprocess the audio dataset to obtain a channel load dataset, a damage dataset, and a restored dataset;

[0071] S3. Calculate based on the channel load data set to obtain a channel load coupling score index tfo for transmission channel matching assessment, and execute a deep audio analysis instruction if the assessment indicates that there is a risk of frame loss in the audio transmission channel;

[0072] S4, deep audio analysis instructions are used to calculate and obtain the frame structure damage depth index zss and semantic recovery credibility index yhf based on the damaged data group and the restored data group respectively;

[0073] S5. Based on the channel load coupling score index tfo, the frame structure damage depth index zss and the semantic recovery credibility index yhf, a summary calculation is performed to obtain the comprehensive scheduling reconstruction index ddc for audio frame recovery evaluation.

[0074] In this embodiment, in S1, a synchronous multi-channel ADC system acquires signals from multiple audio sources and converts the time-domain signals into time-frequency domain signals, obtaining a precise audio time-frequency spectrum. This provides a high-quality data foundation for subsequent feature extraction and signal processing. In S2, feature extraction is performed on the audio time-frequency spectrum to obtain an audio feature dataset. This audio feature dataset is then processed for outliers and dimensionless, resulting in a channel load dataset, a damage dataset, and a restored dataset, providing stable and reliable feature data support. In S3 and S4, the channel load coupling score (TFO) is calculated from the channel load dataset to assess transmission channel matching and accurately evaluate the health of the audio transmission channel. When the channel is at risk of frame loss, deep audio analysis is promptly performed to calculate the frame structure damage depth index (ZSS) and the semantic restoration confidence index (YHF) based on the damage dataset and the restored dataset, respectively, effectively detecting the degree of audio signal damage and its potential for recovery. This process addresses the lack of dynamic evaluation and adjustment mechanisms in traditional audio transmission methods. Through real-time feedback and deep analysis, it ensures high audio signal quality and low frame loss during transmission. S5 comprehensively calculates the channel load coupling score index (TFO), the frame structure damage depth index (ZSS), and the semantic recovery confidence index (YHF) to generate the comprehensive scheduling reconstruction index (DDC) for audio frame recovery assessment, providing a comprehensive evaluation of audio frame recovery. Compared to existing technologies, this method can intelligently select a recovery strategy based on real-time network conditions when the audio signal is lost or damaged, avoiding the unstable recovery results caused by the static repair strategies of traditional methods. This innovation not only improves the robustness of audio transmission, but also achieves more efficient audio recovery in dynamic network environments, significantly optimizing audio transmission quality and user experience.

[0075] Example 2

[0076] Please refer to Figure 1 ,Specifically: S1 includes S11 and S12;

[0077] S11. The configured synchronous multi-channel ADC system collects voice signals from different audio sources on multiple input channels using a time-domain synchronous sampling method and divides the signals into several audio frames. The length of each audio frame is 20ms-40ms, and the specific length is determined by the audio processing requirements. The sampling rate of each audio frame is 48kHz, a common audio processing standard.

[0078] S12, denoising the acquired audio frame, dividing the denoised audio frame into frames according to the frame size, applying a window function to the audio frame signal of each frame to avoid spectrum leakage caused by signal truncation, and then converting the time domain signal into a time-frequency domain signal through short-time Fourier transform to obtain an audio time-frequency spectrum;

[0079] Denoising uses the wavelet denoising method to decompose the audio frame signal into sub-signals of different scales, suppress the noise in different frequency bands respectively, and filter out the noise influence in the audio frame.

[0080] In this embodiment, through the configured synchronous multi-channel ADC system, based on the time domain synchronous sampling method, the voice signals of different audio sources are accurately collected on multiple input channels, and the signals are divided into audio frames according to the set frame length of 20ms-40ms and the standard sampling rate of 48kHz. This process lays a high-quality foundation for subsequent feature extraction and analysis. Through the denoising processing and wavelet denoising method in S12, the audio frame signal is decomposed into sub-signals of different frequency bands, effectively suppressing the influence of noise and avoiding the occurrence of spectrum leakage. The window function is applied to perform signal smoothing processing, and then the time domain signal is converted into a time-frequency domain signal through short-time Fourier transform, and the audio time-frequency spectrum is successfully obtained. These optimization measures significantly improve the clarity and stability of the audio signal, reduce the interference of noise on signal analysis, provide accurate and reliable data support for subsequent audio feature extraction and signal recovery, and effectively enhance the transmission quality and recovery capability of the system.

[0081] Example 3

[0082] Please refer to Figure 1 , specifically: S2 includes S21 and S22;

[0083] S21, extracting features from the acquired audio time-frequency spectrum to obtain an audio feature dataset;

[0084] Perform a second-order difference calculation on the time domain waveform of each frame of audio signal based on the acquired audio time-frequency spectrum to obtain the second-order derivative energy density nm of the audio frame waveform;

[0085] The spectral entropy of each frame of the audio spectrum is calculated using the information entropy formula to obtain the spectral entropy of the audio frame. The spectral entropy change rate xb of adjacent frames is calculated based on the spectral entropy between each adjacent frame. The transmission and retransmission status of each audio frame is then tracked, and the entropy of the spectrum of each retransmitted frame is calculated to obtain the retransmission guided entropy loss rate ss.

[0086] The interference signal in the audio time spectrum is converted into a sparse signal using the compressed sensing method, and the local sparsity of the interference xs is obtained by calculating the sparsity of the sparse signal;

[0087] The phase of each frame signal of the audio time spectrum is extracted by Fourier transform, and the interference phase offset py is obtained by calculating the phase difference between adjacent audio frames;

[0088] Based on the amount of transmitted data in each frame in the audio time spectrum, the load similarity of the burst frame is calculated using correlation analysis to obtain the burst frame load coherence zf;

[0089] According to the transmission time domain of each audio frame, the difference in time intervals between adjacent frames is calculated to obtain the frame period non-uniformity zq;

[0090] The number of samples in each audio frame is obtained by multiplying the sampling rate of the audio frame and the duration of the audio frame. The square values of each audio frame are accumulated to obtain the audio frame energy. The difference in energy between adjacent frames is calculated based on the energy of each audio frame to obtain the frame energy gap qk.

[0091] Use natural language processing technology to extract the semantic information of each audio frame from the speech content in the audio, and obtain the semantic stability yw by comparing the semantic consistency of consecutive audio frames;

[0092] S22, performing outlier processing and dimensionless processing on the audio feature dataset to obtain a channel load data group, a damage data group, and a recovery data group;

[0093] Outlier processing is performed by using the interquartile range method to detect and process outliers in the audio feature dataset, and dimensionless processing is performed by using the Max-Min method to eliminate the dimensional influence of the audio feature dataset;

[0094] The channel load data set includes the second-order derivative energy density nm of the audio frame waveform, the spectral entropy change rate xb of adjacent frames, the interference local sparsity xs, the interference phase offset py and the burst frame load coherence zf;

[0095] The impairment data set includes frame period non-uniformity zq, retransmission guidance entropy loss rate ss and frame energy gap qk;

[0096] The restored data set includes the audio frame energy F and semantic stability yw.

[0097] In this embodiment, through precise audio time-frequency spectrum analysis and feature extraction, combined with a variety of advanced data processing technologies, the transmission quality and recovery capability of the audio signal are significantly improved. First, through second-order difference calculation, spectral entropy calculation, compressed sensing and Fourier transform methods, the multi-dimensional features of the audio frame are extracted to obtain an audio feature data set, which fully reflects the time-frequency characteristics of the audio signal. Pre-processing is performed through outlier processing and dimensionless technology to accurately standardize the data, and stable channel load data groups, damage data groups and recovery data groups are obtained, laying a solid foundation for subsequent audio transmission and recovery evaluation. This solution not only effectively reduces the risk of frame loss and signal damage by performing fine time-frequency analysis and dynamic adjustment of audio signals, but also improves the robustness of audio transmission through precise recovery strategies, ensures high-quality audio recovery in complex environments, and significantly improves the stability of audio transmission and user experience.

[0098] Example 4

[0099] Please refer to Figure 1 , specifically: S3 includes S31;

[0100] S31, performing summary calculation based on the obtained channel load data group to obtain a channel load coupling score index tfo, evaluate the encoding adaptability of the audio stream in the current network environment, and select an appropriate encoding method;

[0101] ;

[0102] Where, e represents the exponential function and ln represents the logarithmic function.

[0103] S3 also includes S32;

[0104] S32. Calculate the average of all historical channel load coupling score indices tfo using a statistical method based on the historical channel load coupling score indices tfo, and set a transmission channel risk threshold A based on the average. The threshold is then compared with the channel load coupling score indices tfo obtained in real time, and a transmission channel matching assessment is performed based on the comparison result. The specific assessment scheme is as follows;

[0105] When the channel load coupling score index tfo is greater than the transmission channel risk threshold A, it indicates that the audio and video transmission channel is healthy and Turbo coding is used for transmission, while real-time monitoring is maintained.

[0106] When the channel load coupling score index tfo ≤ the transmission channel risk threshold A, it means that there is a risk of frame loss in the audio transmission channel. At this time, low-density parity check code LDPC is used for transmission and deep audio analysis instructions are executed.

[0107] In this embodiment, by calculating the channel load coupling score index tfo and comparing it with the preset transmission channel risk threshold A, the health status of the audio transmission channel can be evaluated in real time. During the evaluation process, a summary calculation is performed based on the channel load data group to obtain the channel load coupling score index tfo, which is compared with the preset transmission channel risk threshold A to determine whether the current channel has a risk of frame loss. When the evaluation result shows that the channel is healthy, turbo coding Turbo is selected for efficient transmission, and the transmission process is continuously monitored; when the evaluation result shows that there is a risk of frame loss, low-density parity check coding LDPC is switched to for more robust transmission, and deep audio analysis instructions are enabled. This implementation plan effectively enhances the adaptability and robustness of the audio transmission process, ensures that the encoding method can be flexibly adjusted in different network environments, reduces frame loss, improves the quality and stability of audio transmission, and ultimately optimizes the audio recovery effect and transmission efficiency, thereby improving user experience and system performance.

[0108] Example 5

[0109] Please refer to Figure 1 Specifically: S4, when the transmission channel matching assessment shows that there is a risk of frame loss in the audio transmission channel, execute the deep audio analysis instruction, specifically including S41 and S42;

[0110] S41. Perform summary calculation based on the acquired damage data group to obtain a frame structure damage depth index zss, which is used to analyze whether the received audio frame is damaged and the severity of the damage;

[0111] ;

[0112] Where qk ref Represents the standard audio energy of each frame under the complete audio signal;

[0113] S42, performing summary calculation based on the obtained recovery data group to obtain a semantic recovery credibility index yhf, measuring whether the current frame can be recovered based on the context content, and analyzing the credibility of the recovery;

[0114] ;

[0115] Where k represents the total number of context frames, F i represents the instantaneous energy of the i-th audio frame, Indicates the average energy of the audio frame within the acquisition interval, Var(F i ) represents the variance of the energy of the upper and lower audio frames of the i-th audio frame, which is calculated by statistical methods and is used to measure the variation of the context frame features.

[0116] In this embodiment, when the transmission channel is assessed as being at risk of frame loss, a deep audio analysis instruction is executed, including two sub-steps, S41 and S42, to conduct a comprehensive analysis of the damage and recovery potential of the audio frame. S41 calculates the frame structure damage depth index zss based on the damage data set, quantifies the degree of damage to the audio frame, and ensures accurate identification of lost frames or damaged audio; S42 calculates the semantic recovery credibility index yhf based on the recovery data set to evaluate whether the audio frame can be recovered based on contextual information, thereby determining the feasibility of recovery. This process avoids the audio quality degradation caused by the inability to dynamically adapt to the risk of frame loss in traditional methods by accurately assessing the degree of damage and recovery potential of the audio signal, effectively improving the robustness and recovery efficiency of the audio signal in complex transmission environments. Through these implementations, the recovery strategy can be dynamically selected based on real-time evaluation, significantly improving the stability and high-quality recovery effect during the audio transmission process, ensuring the integrity and semantic coherence of the audio, and ultimately improving the user experience and the reliability of audio transmission.

[0117] Example 6

[0118] Please refer to Figure 1 and Figure 3, specifically: S5 includes S51;

[0119] S51. Based on the obtained channel load coupling score index tfo, frame structure damage depth index zss and semantic recovery credibility index yhf, a summary calculation is performed to obtain a comprehensive scheduling reconstruction index ddc. The specific formula is as follows:

[0120] ;

[0121] Where, By using the frame structure damage depth index zss to indicate the severity of the recovery required, the overall damage is weighted according to the semantic recovery credibility index yhf. By introducing the inverse square root of the load coupling score index TFO, the channel quality is converted into a recovery modulation factor, and nonlinear adjustment is made to the comprehensive scheduling reconstruction index DDC to avoid drastic jumps in the recovery logic due to small channel fluctuations.

[0122] S5 also includes S52;

[0123] S52: Preset an audio frame damage recovery threshold B according to the audio network interoperability industry standard and compare it with the obtained comprehensive scheduling reconstruction index ddc. Perform an audio frame recovery assessment based on the comparison result. The specific assessment scheme is as follows:

[0124] When the comprehensive scheduling reconstruction index ddc is greater than the audio frame damage recovery threshold B, it indicates that there is a risk of frame loss due to audio frame damage. In this case, the audio frame is classified as the first priority for recovery and the semantic structure recovery process is immediately executed.

[0125] When the comprehensive scheduling reconstruction index ddc ≤ the audio frame damage recovery threshold B, it means that there is no risk of frame loss due to audio frame damage. At this time, it is classified as the second priority recovery and interpolation repair is performed.

[0126] In this embodiment, based on the comprehensive calculation of the channel load coupling score index TFO, the frame structure damage depth index ZSS and the semantic recovery credibility index YHF, the comprehensive scheduling reconstruction index DDC is obtained, which provides an accurate evaluation mechanism for the recovery of the audio signal. By combining the frame structure damage depth index ZSS with the semantic recovery credibility index YHF, and introducing the inverse square root of the channel load score TFO, the recovery strategy can be adjusted intelligently, effectively avoiding the instability of the recovery effect caused by network fluctuations. In addition, based on the audio frame damage recovery threshold B preset in the audio network interoperability industry standard, the risk of frame loss of audio frames can be flexibly evaluated to ensure that the semantic structure recovery process is executed immediately when the audio frame is severely damaged, and interpolation repair is used when the damage is relatively minor. This dynamic and adaptive recovery strategy greatly improves the reliability of audio signal transmission, ensures high-quality transmission of audio signals in complex network environments, and optimizes the efficiency and effect of audio recovery.

[0127] Example 7

[0128] Please refer to Figure 2 , a multi-channel audio transmission system, including an audio acquisition module, a feature extraction module, a channel matching evaluation module, a depth analysis module and a comprehensive recovery evaluation module;

[0129] The audio acquisition module is used to collect voice signals from different audio sources on multiple input channels based on a synchronous multi-channel ADC system, divide them into several audio frames, and then convert the time domain signals into time-frequency domain signals through short-time Fourier transform to obtain the audio time-frequency spectrum;

[0130] The feature extraction module is used to extract features from the audio time-frequency spectrum to obtain an audio feature data set, and then preprocess the audio data set to obtain a channel load data set, a damage data set, and a recovery data set;

[0131] The channel matching evaluation module is used to calculate based on the channel load data group to obtain the channel load coupling score index TFO for transmission channel matching evaluation, and execute deep audio analysis instructions when it is assessed that there is a risk of frame loss in the audio transmission channel;

[0132] The deep analysis module is used to execute deep audio analysis instructions, and calculate the frame structure damage depth index zss and the semantic recovery credibility index yhf according to the damaged data group and the restored data group respectively;

[0133] The comprehensive recovery evaluation module is used to perform summary calculations based on the channel load coupling score index tfo, the frame structure damage depth index zss and the semantic recovery credibility index yhf, and obtain the comprehensive scheduling reconstruction index ddc for audio frame recovery evaluation.

[0134] A multi-channel audio transmission storage medium stores a computer program. When the computer program is executed, a multi-channel audio transmission method is implemented when the program is executed by a processor.

[0135] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A multi-channel audio transmission method, characterized in that: The following steps are involved: S1. Speech signals from different audio sources are collected on multiple input channels through a synchronous multi-channel ADC system and divided into several audio frames. The time-domain signals are then converted into time-frequency domain signals through short-time Fourier transform to obtain the audio time-frequency spectrum. S2. Extract features from the audio time-frequency spectrum to obtain an audio feature dataset, and then preprocess the audio dataset to obtain a channel load dataset, a damage dataset, and a restored dataset; S3. Calculate based on the channel load data set to obtain a channel load coupling score index tfo for transmission channel matching assessment, and execute a deep audio analysis instruction if the assessment indicates that there is a risk of frame loss in the audio transmission channel; S4, deep audio analysis instructions are used to calculate and obtain the frame structure damage depth index zss and semantic recovery credibility index yhf based on the damaged data group and the restored data group respectively; S5. Based on the channel load coupling score index tfo, the frame structure damage depth index zss and the semantic recovery credibility index yhf, a summary calculation is performed to obtain the comprehensive scheduling reconstruction index ddc for audio frame recovery evaluation.

2. The multi-channel audio transmission method according to claim 1, wherein: S1 includes S11 and S12; S11. The configured synchronous multi-channel ADC system collects voice signals from different audio sources on multiple input channels using a time-domain synchronous sampling method and divides the signals into several audio frames. The length of each audio frame is 20ms-40ms, and the specific length is determined by the audio processing requirements. The sampling rate of each audio frame is 48kHz, a common audio processing standard. S12, denoising the acquired audio frame, dividing the denoised audio frame into frames according to the frame size, applying a window function to the audio frame signal of each frame to avoid spectrum leakage caused by signal truncation, and then converting the time domain signal into a time-frequency domain signal through short-time Fourier transform to obtain an audio time-frequency spectrum; Denoising uses the wavelet denoising method to decompose the audio frame signal into sub-signals of different scales, suppress the noise in different frequency bands respectively, and filter out the noise influence in the audio frame.

3. The multi-channel audio transmission method according to claim 2, wherein: S2 includes S21 and S22; S21, extracting features from the acquired audio time-frequency spectrum to obtain an audio feature dataset; Perform a second-order difference calculation on the time domain waveform of each frame of audio signal based on the acquired audio time-frequency spectrum to obtain the second-order derivative energy density nm of the audio frame waveform; The spectral entropy of each frame of the audio spectrum is calculated using the information entropy formula to obtain the spectral entropy of the audio frame. The spectral entropy change rate xb of adjacent frames is calculated based on the spectral entropy between each adjacent frame. The transmission and retransmission status of each audio frame is then tracked, and the entropy of the spectrum of each retransmitted frame is calculated to obtain the retransmission guided entropy loss rate ss. The interference signal in the audio time spectrum is converted into a sparse signal using the compressed sensing method, and the local sparsity of the interference xs is obtained by calculating the sparsity of the sparse signal; The phase of each frame signal of the audio time spectrum is extracted by Fourier transform, and the interference phase offset py is obtained by calculating the phase difference between adjacent audio frames; Based on the amount of transmitted data in each frame in the audio time spectrum, the load similarity of the burst frame is calculated using correlation analysis to obtain the burst frame load coherence zf; According to the transmission time domain of each audio frame, the difference in time intervals between adjacent frames is calculated to obtain the frame period non-uniformity zq; The number of samples in each audio frame is obtained by multiplying the sampling rate of the audio frame and the duration of the audio frame. The square values of each audio frame are accumulated to obtain the audio frame energy. The difference in energy between adjacent frames is calculated based on the energy of each audio frame to obtain the frame energy gap qk. Use natural language processing technology to extract the semantic information of each audio frame from the speech content in the audio, and obtain the semantic stability yw by comparing the semantic consistency of consecutive audio frames; S22, performing outlier processing and dimensionless processing on the audio feature dataset to obtain a channel load data group, a damage data group, and a recovery data group; Outlier processing is performed by using the interquartile range method to detect and process outliers in the audio feature dataset, and dimensionless processing is performed by using the Max-Min method to eliminate the dimensional influence of the audio feature dataset; The channel load data set includes the second-order derivative energy density nm of the audio frame waveform, the spectral entropy change rate xb of adjacent frames, the interference local sparsity xs, the interference phase offset py and the burst frame load coherence zf; The impairment data set includes frame period non-uniformity zq, retransmission guidance entropy loss rate ss and frame energy gap qk; The restored data set includes the audio frame energy F and semantic stability yw.

4. The multi-channel audio transmission method according to claim 3, wherein: S3 includes S31; S31, performing summary calculation based on the obtained channel load data group to obtain a channel load coupling score index tfo, evaluate the encoding adaptability of the audio stream in the current network environment, and select an appropriate encoding method; ; Where, e represents the exponential function and ln represents the logarithmic function.

5. The multi-channel audio transmission method according to claim 4, characterized in that: S3 also includes S32; S32. Calculate the average of all historical channel load coupling score indices tfo using a statistical method based on the historical channel load coupling score indices tfo, and set a transmission channel risk threshold A based on the average. The threshold is then compared with the channel load coupling score indices tfo obtained in real time, and a transmission channel matching assessment is performed based on the comparison result. The specific assessment scheme is as follows; When the channel load coupling score index tfo is greater than the transmission channel risk threshold A, it indicates that the audio and video transmission channel is healthy and Turbo coding is used for transmission, while real-time monitoring is maintained. When the channel load coupling score index tfo ≤ the transmission channel risk threshold A, it means that there is a risk of frame loss in the audio transmission channel. At this time, low-density parity check code LDPC is used for transmission and deep audio analysis instructions are executed.

6. The multi-channel audio transmission method according to claim 5, characterized in that: S4. When the transmission channel matching assessment indicates that there is a risk of frame loss in the audio transmission channel, execute the deep audio analysis instruction, specifically including S41 and S42; S41, performing summary calculation based on the acquired damage data group to obtain a frame structure damage depth index zss; ; Where qk ref Represents the standard audio energy of each frame under the complete audio signal; S42, performing summary calculation based on the obtained recovery data group to obtain a semantic recovery credibility index yhf; ; Where k represents the total number of context frames, F i represents the instantaneous energy of the i-th audio frame, Indicates the average energy of the audio frame within the acquisition interval, Var(F i ) represents the variance of the energy of the upper and lower audio frames of the i-th audio frame.

7. The multi-channel audio transmission method according to claim 6, characterized in that: S5 includes S51; S51. Based on the obtained channel load coupling score index tfo, frame structure damage depth index zss and semantic recovery credibility index yhf, a summary calculation is performed to obtain a comprehensive scheduling reconstruction index ddc. The specific formula is as follows: 。 8. The multi-channel audio transmission method according to claim 7, wherein: S5 also includes S52; S52: Preset an audio frame damage recovery threshold B according to the audio network interoperability industry standard and compare it with the obtained comprehensive scheduling reconstruction index ddc. Perform an audio frame recovery assessment based on the comparison result. The specific assessment scheme is as follows: When the comprehensive scheduling reconstruction index ddc is greater than the audio frame damage recovery threshold B, it indicates that there is a risk of frame loss due to audio frame damage. In this case, the audio frame is classified as the first priority for recovery and the semantic structure recovery process is immediately executed. When the comprehensive scheduling reconstruction index ddc ≤ the audio frame damage recovery threshold B, it means that there is no risk of frame loss due to audio frame damage. At this time, it is classified as the second priority recovery and interpolation repair is performed.

9. A multi-channel audio transmission system, comprising a multi-channel audio transmission method according to any one of claims 1 to 8, characterized in that: It includes audio acquisition module, feature extraction module, channel matching evaluation module, depth analysis module and comprehensive recovery evaluation module; The audio acquisition module is used to collect voice signals from different audio sources on multiple input channels based on a synchronous multi-channel ADC system, divide them into several audio frames, and then convert the time domain signals into time-frequency domain signals through short-time Fourier transform to obtain the audio time-frequency spectrum; The feature extraction module is used to extract features from the audio time-frequency spectrum to obtain an audio feature data set, and then preprocess the audio data set to obtain a channel load data set, a damage data set, and a recovery data set; The channel matching evaluation module is used to calculate based on the channel load data group to obtain the channel load coupling score index TFO for transmission channel matching evaluation, and execute deep audio analysis instructions when it is assessed that there is a risk of frame loss in the audio transmission channel; The deep analysis module is used to execute deep audio analysis instructions, and calculate the frame structure damage depth index zss and the semantic recovery credibility index yhf according to the damaged data group and the restored data group respectively; The comprehensive recovery evaluation module is used to perform summary calculations based on the channel load coupling score index tfo, the frame structure damage depth index zss and the semantic recovery credibility index yhf, and obtain the comprehensive scheduling reconstruction index ddc for audio frame recovery evaluation.

10. A multi-channel audio transmission storage medium, characterized by: The storage medium stores a computer program. When the computer program is executed, the program is executed by a processor to implement a multi-channel audio transmission method as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Key frame determination method, electronic equipment, storage medium and computer program product

    CN121415320A