Voice data loss recovery processing method, device and equipment and storage medium

By performing frame processing and sparse reconstruction on voice message data in VoIP communication, and combining it with a dynamic filtering network to recover lost voice data, the problems of bandwidth waste caused by redundant transmission and poor traditional filtering effect are solved, and efficient and stable voice data recovery and quality improvement are achieved.

CN120544589BActive Publication Date: 2025-10-10SHENZHEN DINSTAR TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510992158.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-10-10
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

Existing technologies compensate for message loss in VoIP communications by redundantly transmitting voice data, which increases bandwidth requirements, exacerbates network congestion, and fails to guarantee voice transmission rate and quality.

Method used

The received voice message data is framed, sparsely reconstructed using compressed sensing technology, and dynamically filtered in combination with a preset filtering network to restore the lost voice data.

Benefits of technology

It saves bandwidth while ensuring the quality of voice transmission, efficiently recovers lost voice frames, improves the quality of voice communication, and adapts to different noise environments and differences in the voice characteristics of speakers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544589B_ABST
    Figure CN120544589B_ABST
Patent Text Reader

Abstract

The application discloses a voice data loss recovery processing method and device, equipment and a storage medium, and relates to the technical field of voice transmission. The method comprises the following steps: performing frame processing on received voice message data to obtain frame voice data; when signal loss is detected, performing voice reconstruction on the frame voice data based on compressed sensing to obtain initial voice recovery data; and performing dynamic voice filtering on the initial voice recovery data through a preset filtering network to obtain target voice recovery data. The voice message data is first subjected to frame processing, the voice is reconstructed based on compressed sensing by using the sparsity of the voice signal when the signal is lost, and then dynamic filtering is performed through the preset filtering network, so that high-quality voice data is finally output. The method solves the problems of bandwidth waste caused by redundant transmission and poor effect of traditional interpolation filtering in the prior art, realizes the effects of saving bandwidth, efficiently recovering lost voice frames and improving voice quality, and meets the demand for high-quality voice communication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of voice transmission technology, and in particular to a method, apparatus, device, and storage medium for recovering lost voice data. Background Art

[0002] In the field of Voice over Internet Protocol (VoIP) communications, voice packet loss is a common problem. Currently, redundancy is commonly used to ensure voice quality, specifically by transmitting redundant voice data. While this approach can compensate for packet loss to a certain extent, it has a serious drawback: it significantly increases bandwidth requirements and leads to increased network congestion. Given limited network bandwidth resources, excessive redundant data transmission consumes significant bandwidth, disrupting the normal operation of other network services and degrading overall network performance.

[0003] Therefore, how to ensure the voice transmission quality while ensuring the voice transmission rate has become an urgent problem to be solved. Summary of the Invention

[0004] The main purpose of this application is to provide a voice data loss recovery processing method, device, equipment and storage medium, aiming to solve the technical problem of how to ensure the voice transmission quality while ensuring the voice transmission rate.

[0005] To achieve the above objectives, the present application proposes a method for recovering lost voice data, which includes:

[0006] Performing frame processing on the received voice message data to obtain framed voice data;

[0007] When signal loss is detected, performing speech sparse reconstruction based on compressed sensing on the framed speech data to obtain initial speech recovery data;

[0008] The initial speech restoration data is subjected to dynamic speech filtering through a preset filtering network to obtain target speech restoration data.

[0009] In one embodiment, when signal loss is detected, the step of performing compressed sensing-based speech sparse reconstruction on the framed speech data to obtain initial speech recovery data includes:

[0010] When signal loss is detected, converting the detected lost frame speech data into a sparse speech signal;

[0011] Obtaining a lost frame sparse coefficient corresponding to the sparse speech signal;

[0012] The framed speech data is reconstructed using a preset speech configuration according to the lost frame sparse coefficient to obtain initial speech recovery data.

[0013] In one embodiment, the step of converting the detected lost frame speech data into a sparse speech signal includes:

[0014] Obtaining a target sparse transform basis corresponding to the detected lost frame speech data;

[0015] The lost frame speech data is projected onto a sparse domain according to the local sparsity of the lost frame speech data and the target sparse transformation basis to obtain a sparse speech signal.

[0016] In one embodiment, the step of obtaining the lost frame sparse coefficient corresponding to the sparse speech signal includes:

[0017] Constructing a target measurement matrix corresponding to the lost frame speech data based on adjacent valid frame data of the lost frame speech data and the target sparse transformation basis;

[0018] Signal solution is performed based on the target observation matrix and the sparse speech signal through a preset optimization algorithm to obtain the lost frame sparse coefficients.

[0019] In one embodiment, the step of constructing a target measurement matrix corresponding to the lost frame speech data based on adjacent valid frame data of the lost frame speech data and the target sparse transform basis includes:

[0020] Obtaining a voice activity detection result of the lost frame voice data;

[0021] Determining an initial observation matrix based on the time-frequency characteristics of the lost frame speech data;

[0022] The initial measurement matrix is ​​optimized according to the voice activity detection result, the adjacent valid frame data of the lost frame voice data and the target sparse transformation basis to obtain a target measurement matrix.

[0023] In one embodiment, the step of reconstructing the framed speech data according to the lost frame sparse coefficient to obtain initial speech recovery data includes:

[0024] Obtaining an inverse transformation matrix corresponding to the target sparse transformation basis;

[0025] Reconstructing a preset voice based on the lost frame sparse coefficient and the inverse transformation matrix to obtain a lost reconstructed voice frame;

[0026] Initial speech recovery data is generated based on the lost reconstructed speech frame and the framed speech data.

[0027] In one embodiment, the step of performing dynamic speech filtering on the initial speech recovery data through a preset filtering network to obtain target speech recovery data includes:

[0028] Inputting the initial speech recovery data into a preset filter network to obtain adaptive filter parameters;

[0029] Dynamic speech filtering is performed on the initial speech restoration data based on the adaptive filtering parameters and the dynamic filter to obtain target speech restoration data.

[0030] In addition, to achieve the above-mentioned purpose, the present application also proposes a voice data loss recovery processing device, which includes:

[0031] The voice framing module is used to perform framing processing on the received voice message data to obtain framed voice data;

[0032] A speech reconstruction module is used to perform speech sparse reconstruction based on compressed sensing on the framed speech data to obtain initial speech recovery data when signal loss is detected;

[0033] The filtering module is used to perform dynamic speech filtering on the initial speech recovery data through a preset filtering network to obtain target speech recovery data.

[0034] In addition, to achieve the above-mentioned purpose, the present application also proposes a voice data loss recovery processing device, which includes: a memory, a processor, and a voice data loss recovery processing program stored in the memory and executable on the processor, wherein the voice data loss recovery processing program is configured to implement the steps of the voice data loss recovery processing method as described above.

[0035] In addition, to achieve the above-mentioned objectives, the present application also proposes a storage medium, which is a storage medium, on which is stored a voice data loss recovery processing program, and when the voice data loss recovery processing program is executed by a processor, the steps of the voice data loss recovery processing method described above are implemented. The present application provides a voice data loss recovery processing method, apparatus, device, and storage medium.

[0036] The present application discloses a method, apparatus, device and storage medium for voice data loss recovery processing, the method comprising: performing frame processing on received voice message data to obtain framed voice data; when signal loss is detected, performing voice reconstruction based on compressed sensing on the framed voice data to obtain initial voice recovery data; and performing dynamic voice filtering on the initial voice recovery data through a preset filtering network to obtain target voice recovery data. The present application first performs frame processing on the voice message data, and when the signal is lost, reconstructs the voice based on compressed sensing using the sparsity of the voice signal, and then dynamically filters the voice through a preset filtering network to finally output high-quality voice data. This solves the problems in the prior art of redundant transmission leading to bandwidth waste and poor traditional interpolation filtering effects, and achieves the effects of saving bandwidth, efficiently recovering lost voice frames and improving voice quality, meeting the needs of high-quality voice communication. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0038] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0039] Figure 1 This is a first flow chart of the first embodiment of the voice data loss recovery method of the present application;

[0040] Figure 2 This is a second flow chart of the first embodiment of the voice data loss recovery method of the present application;

[0041] Figure 3 This is a first flow chart of the second embodiment of the voice data loss recovery method of the present application;

[0042] Figure 4 This is a second flow chart of the second embodiment of the voice data loss recovery method of the present application;

[0043] Figure 5 This is a third flow chart of the second embodiment of the voice data loss recovery method of the present application;

[0044] Figure 6 This is a schematic diagram of the module structure of the voice data loss recovery processing device according to an embodiment of the present application;

[0045] Figure 7A device structure schematic diagram of a hardware running environment involved in a voice data loss recovery processing method in the embodiments of the present application.

[0046] The object implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0047] It should be understood that the specific embodiments described herein are merely intended to explain the technical solutions of the present application, and are not intended to limit the present application.

[0048] In order to better understand the technical solutions of the present application, the specific embodiments will be described in detail below with reference to the drawings and the accompanying drawings.

[0049] The main solution of the present application is: performing frame processing on the received voice message data to obtain frame voice data; when signal loss is detected, performing voice sparse reconstruction based on compressed sensing on the frame voice data to obtain initial voice recovery data; and performing dynamic voice filtering on the initial voice recovery data through a preset filtering network to obtain target voice recovery data.

[0050] The method of currently adopted additional transmission of redundant voice data to guarantee voice quality, although to some extent, makes up for the impact of message loss, has serious defects: greatly increases bandwidth demand, leading to network congestion intensification. In the case of limited network bandwidth resources, too much redundant data transmission will occupy a large amount of bandwidth, making other network services unable to develop normally, and reducing the overall network performance.

[0051] The present application provides a voice data loss recovery processing method, which first performs frame processing on voice message data, reconstructs voice based on compressed sensing using the sparsity of voice signal when signal loss occurs, and then performs dynamic filtering through a preset filtering network, and finally outputs high-quality voice data. The method solves the problems of bandwidth waste caused by redundant transmission and poor effect of traditional interpolation filtering in the prior art, realizes the effects of saving bandwidth, efficiently recovering lost voice frames and improving voice quality, and meets the demand for high-quality VoIP voice communication.

[0052] It should be noted that the execution subject of the present embodiment can be a voice data loss recovery processing system, or a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or a voice data loss recovery processing device capable of realizing the above functions, etc., and the present embodiment does not make specific limitation thereon. The present embodiment and the following embodiments will be described below by taking a voice data loss recovery processing device (referred to as a processing device for short) as an execution subject.

[0053] Based on this, the present embodiment provides a voice data loss recovery processing method, which refers to Figure 1 , Figure 1 This is a first flow chart of the first embodiment of the voice data loss recovery processing method of the present application.

[0054] In this embodiment, the voice data loss recovery method includes steps S10 to S30:

[0055] Step S10, performing frame processing on the received voice message data to obtain framed voice data;

[0056] Step S20, when signal loss is detected, performing speech sparse reconstruction based on compressed sensing on the framed speech data to obtain initial speech recovery data;

[0057] Step S30 , performing dynamic speech filtering on the initial speech recovery data through a preset filtering network to obtain target speech recovery data.

[0058] It is understood that the aforementioned voice message data may be data packets carrying voice information transmitted in VoIP communications, comprising data obtained through sampling, encoding, and other processing of voice signals, and used to transmit voice content between different communication devices. To facilitate subsequent rapid search and recovery of lost frames, this embodiment may segment the continuous voice signal in the voice message data into multiple short-term frames at regular time intervals, such as one frame every 20 ms, to obtain a voice signal comprising multiple voice frames, i.e., the aforementioned framed voice signal.

[0059] It should be understood that the processing device can monitor the framed voice signal in real time during the communication process. Furthermore, this embodiment utilizes compressed sensing technology. Based on the sparse nature of the voice signal and known partial data, upon detecting signal loss in the framed voice signal (e.g., RTP sequence number (RTP) sequence number (RTP) sequence number) interruption), it can quickly reconstruct and recover the lost voice frames in the framed voice signal using a preset mathematical model and algorithm to ensure the integrity of the voice message. In comparison, the amount of data in the lost voice frame is much smaller than the data in the framed voice signal. Therefore, this embodiment can effectively improve the speed of voice data restoration.

[0060] It can be understood that the preliminary recovered speech data obtained after the speech reconstruction based on the compressed sensing may have some noise or inaccuracy. To ensure the quality of speech transmission, the initial speech recovery data can be further processed in the embodiment. The preset filtering network can be a network structure designed in advance for filtering processing of speech signals, which is usually constructed based on a deep learning model. The network can process the input speech data and output filtering parameters dynamically adjusted according to the real-time characteristics of the input speech signals, so as to realize adaptive filtering processing of the speech signals, remove noise, enhance the key features of the speech signals, and thus obtain high-quality target speech recovery data.

[0061] In a possible implementation, with reference to Figure 2 , Figure 2 FIG. 2 is a second flowchart of a speech data loss recovery processing method according to an embodiment of the present application. In the embodiment, step S30 can include steps A1-A2.

[0062] In step A1, the initial speech recovery data is input into the preset filtering network to obtain adaptive filtering parameters.

[0063] In step A2, the initial speech recovery data is dynamically filtered based on the adaptive filtering parameters and a dynamic filter to obtain target speech recovery data.

[0064] It can be understood that the adaptive filtering parameters can be filtering parameters generated by the preset filtering network in real time according to the input initial speech recovery data. These parameters can be dynamically adjusted according to the changes of the speech signals and the noise environment, so as to achieve the best filtering effect of the speech signals. For example, under different noise interference conditions, the adaptive filtering parameters can change the frequency response characteristics and gain of the filter, so as to better remove noise and retain the key features of the speech signals.

[0065] It will be readily understood that in this embodiment, the preset filtering network may employ an adaptive filtering model based on deep learning, such as a recurrent neural network (RNN), a gated recurrent unit (GRU), or a long short-term memory network (LSTM). For example, when the preset filtering network is an LSTM model, it can learn the long-term dependencies and noise characteristics of speech signals. When the acoustic environment of the input speech signal changes, the LSTM model dynamically adjusts its internal parameters, such as the weight matrix and bias vector, based on the real-time characteristics of the input signal (such as spectral characteristics and time-domain waveform). Therefore, this embodiment enables the preset filtering network to process sequence data and capture the temporal characteristics of speech signals. During the model training phase, a large amount of speech data from different acoustic environments can be pre-trained, allowing the iterative preset filtering network to effectively learn the mapping relationship between different noise environments and filtering parameters. Furthermore, the preset filtering network also features a real-time feedback mechanism that can further adjust model parameters based on the quality of the filtered speech signal, forming a closed-loop optimization loop. This allows the network to quickly adapt to different acoustic environments and effectively filter speech signals.

[0066] In practical applications, the potentially noisy initial speech data is fed into an iteratively pre-set filtering network. The pre-set filtering network then generates filter parameters in real time based on the characteristics of the initial speech data (such as spectral characteristics and time-domain waveform). When the acoustic environment changes, the initial speech data can rapidly adjust its parameters to adapt to the new noise conditions. For example, when a burst of noise is detected in the speech signal, the model instantly updates the filter's frequency response, enhancing its ability to suppress noise at specific frequencies. Furthermore, the model possesses online learning capabilities, enabling it to further optimize model parameters based on real-time feedback on filtering performance, thereby better adapting to changing acoustic environments.

[0067] It is understandable that the above-mentioned dynamic filter can be a filter that can adjust its filtering characteristics in real time based on the input signal and adaptive filtering parameters, such as an FIR filter (Finite Impulse Response Filter), which can be used to perform time-domain convolution filtering on the reconstructed frame. Compared with traditional fixed parameter filters, dynamic filters have greater adaptability and flexibility and can adapt to complex and changing voice communication environments and noise conditions. This embodiment can utilize a dynamic filter to adjust the filter characteristics in real time based on the generated adaptive filtering parameters, and perform adaptive filtering on the initial voice recovery data, thereby effectively removing noise interference and enhancing the key information of the voice signal, thereby obtaining high-precision target voice recovery data.

[0068] For example, when a speech signal characteristic changes (such as a sudden change in fundamental frequency), for example, the transition tone from "ah" (ah) to "oh" (oo) appears in the current speech frame in the initial speech recovery data, the fundamental frequency suddenly changes from 120 Hz to 150 Hz, the formant F1 rises from 500 Hz to 700 Hz, and F2 drops from 1200 Hz to 1000 Hz. However, the noise environment is a stable white background noise with a signal-to-noise ratio (SNR) of 15 dB (decibel).

[0069] At this time, if the preset filtering network is an LSTM model, the changing trends of the fundamental frequency and formant can be learned through historical speech frames (such as the "ah" sound in the previous 50ms) during the training process, and the 32nd-order FIR filter coefficient Wold=[0.12,-0.05,0.23,...,0.08] (suitable for the low-frequency "ah" sound) can be output.

[0070] When the LSTM hidden layer captures sudden changes in the fundamental frequency and formants of the current frame in the input feature vector (such as MFCC (Mel-frequency cepstral coefficient) features), for example, a rising trend in the fundamental frequency (ΔF0 = +30 Hz) and a formant shift (ΔF1 = +200 Hz, ΔF2 = -200 Hz) are detected, the LSTM then maps the hidden layer state into filter coefficients through a fully connected layer, updating the coefficients in real time to [-0.03, 0.18, 0.28, ..., 0.15]. The output new filter parameters increase the weights of the 10th to 15th taps (corresponding to the 500-1000 Hz frequency band) to enhance the mid-frequency formants, and decrease the weights of the first five taps (corresponding to the 0-300 Hz frequency band) to suppress the impact noise caused by the sudden change in the low-frequency fundamental frequency. This enhances the mid-frequency components and effectively adapts to the spectral characteristics of the "oh" sound.

[0071] In addition, if burst noise is detected in the speech signal, the preset filtering network can also instantly update the frequency response of the filter to enhance the ability to suppress noise of specific frequencies, thereby achieving the purpose of dynamically adjusting the model parameters to cope with different acoustic environments.

[0072] For example, if a sudden change in the noise environment occurs (such as a sudden car horn), the current frame of the normal speech content "Please wait" contains the character "děng" (wait), and the spectrum may be concentrated between 400 and 800 Hz. However, a sudden car horn superimposes on the background white noise (a sudden narrowband noise burst at 1500 Hz, causing the SNR to drop sharply to 5 dB). In this case, a preset filtering network, such as an RNN model, can output initial coefficients of Wold = [0.05, 0.15, 0.20, ..., 0.04] (a passband covering 200-1000 Hz to suppress high-frequency noise) for the first 100 ms of stable speech data. Correspondingly, after detecting the 1500 Hz noise burst, the RNN model can update the coefficients to Wnew = [0.03, 0.12, 0.18, ..., -0.10] within 5 ms (i.e., introducing a notch filter at 1500 Hz), thereby achieving noise suppression.

[0073] Therefore, the final speech restoration result obtained after dynamic speech filtering processing, that is, the above-mentioned target speech restoration data, has good speech quality, clarity and intelligibility, and can be directly used for playback, providing users with a high-quality and clear voice communication experience.

[0074] This embodiment proposes a method for recovering from voice data loss. The method first frames the voice message data. When the signal is lost, the voice is reconstructed using compressed sensing, leveraging the sparsity of the voice signal. The voice is then adaptively filtered using a preset filtering network and a dynamic filter, ultimately outputting high-quality voice data. Compressed sensing-based voice reconstruction fully exploits the sparsity of the voice signal to accurately reconstruct lost voice frames, preserving the original characteristics of the voice to the greatest extent possible. A deep learning-based adaptive filtering algorithm dynamically adjusts filter parameters based on real-time changes in the voice signal and noise environment, effectively removing noise and significantly improving voice clarity and intelligibility, providing users with a superior voice communication experience.

[0075] At the same time, the powerful adaptive capabilities of the preset filtering network enable the processing device to quickly adapt to the varying speech characteristics of different speakers, complex and variable network noise environments, and bandwidth-sensitive and noise-variable scenarios such as mines and vehicles, maintaining stable and efficient voice processing performance. Therefore, this embodiment solves the existing problems of redundant transmission leading to bandwidth waste and poor traditional interpolation filtering, achieving bandwidth savings, efficient recovery of lost voice frames, and improved voice quality, meeting the needs of high-quality voice communication.

[0076] The embodiment provides a voice data loss recovery processing method, which comprises the following steps: performing frame processing on received voice message data to obtain frame voice data; when signal loss is detected, performing voice sparse reconstruction on the frame voice data based on compressed sensing to obtain initial voice recovery data; inputting the initial voice recovery data into a preset filter network to obtain adaptive filter parameters; and performing dynamic voice filtering on the initial voice recovery data based on the adaptive filter parameters and a dynamic filter to obtain target voice recovery data. The voice message data is first frame processed in the embodiment, and when signal loss occurs, the voice is reconstructed based on compressed sensing according to the sparsity of the voice signal, and then the voice is adaptively filtered through the preset filter network and the dynamic filter, so that high-quality voice data is finally output. The powerful adaptive capability of the preset filter network enables the processing device to quickly adapt to different voice characteristics of different speakers and complex and changeable network noise environments, and to maintain stable and efficient voice processing performance. Therefore, the embodiment solves the problems of bandwidth waste caused by redundant transmission and poor effect of traditional interpolation filtering in the prior art, achieves the effects of saving bandwidth, efficiently recovering lost voice frames and improving voice quality, and meets the high-quality voice communication demand.

[0077] Based on the first embodiment of the application, the same or similar contents as the above embodiment one can be referred to the above description, and will not be described hereinafter.

[0078] Based on the first embodiment, please refer to Figure 3 , Figure 3 is a first flowchart of the second embodiment of the voice data loss recovery processing method of the application. In the embodiment, step S20 comprises steps B1-B3.

[0079] Step B1, when signal loss is detected, the detected lost frame voice data is converted into sparse voice signals.

[0080] It is easy to understand that the above-mentioned lost frame voice data is voice data frame in the frame voice data which is not successfully received or transmitted in the voice communication process due to network and other reasons, and affects the integrity and continuity of the voice, and the voice reconstruction technology can be used for recovery in the embodiment. The voice signal has sparsity in a specific transform domain (such as a discrete cosine transform domain, a wavelet transform domain, etc.), that is, most of the energy is concentrated in a small number of coefficients, only a small number of non-zero coefficients or coefficients with large amplitude, and these coefficients can represent the main characteristics of the voice signal, which are obtained by sparse representation of the voice signal. Therefore, in view of the problem that the existing lost frame recovery technology cannot fully utilize the complex characteristics of the voice signal, and it is difficult to restore the real characteristics of the voice when recovering the lost voice frame, the embodiment is used for high-precision recovery of the lost frame voice data, and the lost frame voice data can be converted into sparse domain for processing.

[0081] In a possible implementation, please refer to Figure 4 , Figure 4 This is a second flow chart of the second embodiment of the voice data loss recovery method of the present application. In this embodiment, step B1 may include steps B11 to B12:

[0082] Step B11, obtaining a target sparse transform basis corresponding to the detected lost frame speech data;

[0083] Step B12: Projecting the lost frame speech data into a sparse domain according to the local sparsity of the lost frame speech data and the target sparse transformation basis to obtain a sparse speech signal.

[0084] It is understood that this embodiment can perform a sparse transform on the lost frame speech data, converting it to a sparse domain to obtain a sparse speech signal, which can then be used to reconstruct the speech using its sparse characteristics. The target sparse transform basis can be a transform basis suitable for sparse representation, selected based on the characteristics of the lost frame speech data. Common examples include the discrete cosine transform (DCT), wavelet transform, and Fourier transform. Different transform bases can convert the speech signal into different transform domains, highlighting its sparse characteristics.

[0085] Exemplarily, the processing device may first determine that the target sparse transform basis of the lost frame speech data is a wavelet transform basis. Then, the lost frame speech data is subjected to wavelet transform to convert the lost frame speech data from the time domain to the wavelet domain. In the wavelet domain, the energy of the lost frame speech data is mainly concentrated in a small number of wavelet coefficients. By setting a threshold, coefficients with amplitudes greater than the threshold are retained, and coefficients with amplitudes less than the threshold are discarded, thereby achieving sparse representation of the lost frame speech data. At the same time, appropriate wavelet basis functions and decomposition layers can be selected according to the characteristics of the speech signal and the application scenario to improve the effect of sparse representation. For example, for speech signals with different frequency components, orthogonal wavelet basis functions with different characteristics can be selected for transformation to better capture the sparse characteristics of the lost frame speech data.

[0086] Taking the Fourier transform as an example, after converting the lost frame speech data into the frequency domain, the energy of the speech signal is primarily concentrated in certain frequency components. In this case, thresholding the signal in the frequency domain can be performed to retain frequency coefficients with larger amplitudes and discard those with smaller amplitudes, thereby achieving a sparse representation of the speech signal.

[0087] In addition, in the sparse representation process, the time-frequency characteristics of the speech signal can be combined to select appropriate time windows and transformation parameters to better capture the local characteristics of the speech signal and improve the effect of sparse representation.

[0088] It needs to be understood that the above local sparsity can refer to the characteristic that the missing frame speech data has sparsity in a short time window (such as the time range of a frame of speech signal), such as the periodicity of voiced sound and the wideband characteristic of unvoiced sound. Therefore, in the process of projecting the missing frame speech data into the sparse domain based on the target sparse transform basis, the embodiment can concentrate most of the energy of the speech signal on a small number of coefficients based on the local sparsity, which can reflect the main characteristics of the speech signal in the local time range, such as the fundamental frequency and the harmonic structure, so as to maximize the sparse characteristic of the speech signal, and enable more effective signal processing and recovery based on the sparse representation subsequently.

[0089] In the embodiment, by selecting a suitable target sparse transform basis and combining the domain conversion of the local sparsity of the missing frame speech data, the sparse speech signal can be more accurately obtained, which provides a more reliable basis for subsequent speech reconstruction, solves the problem of inaccurate sparse representation and affects the quality of speech recovery caused by directly performing sparse representation, and further improves the accuracy of the speech data loss recovery processing.

[0090] Step B2, obtaining the missing frame sparse coefficient corresponding to the sparse speech signal;

[0091] It is easy to understand that the above missing frame sparse coefficient can be a sparse representation coefficient representing the characteristics of the missing frame speech data in the sparse domain. According to the known speech frame and the compressive sensing theory, the embodiment can construct an observation matrix based on the time-space correlation of adjacent frames, capture the time sequence dependence (such as the continuity of the speech signal) between the missing frame and the adjacent frames, and process the sparse speech signal through a specific algorithm (such as the orthogonal matching pursuit algorithm), to extract the sparse coefficient representing the characteristics of the missing frame speech data.

[0092] In a feasible embodiment, in the embodiment, step B2 can include steps B21-B22:

[0093] Step B21, constructing a target observation matrix corresponding to the missing frame speech data based on the adjacent valid frame data of the missing frame speech data and the target sparse transform basis;

[0094] It needs to be noted that the above adjacent valid frame data can be those successfully received or transmitted speech frame data adjacent to the missing frame, which contains speech information having continuity and correlation with the missing frame in time, and can provide an important reference for recovering the missing frame and help infer the speech characteristics and change trend of the missing frame.

[0095] The target measurement matrix is ​​a matrix used to connect the original signal (i.e., the lost frame speech data) and the measured values ​​(i.e., known adjacent valid frame data). Its design and construction directly impact the quality of signal reconstruction. This embodiment uses the target measurement matrix to capture relevant information, such as the temporal dependencies between the lost frame and its preceding and subsequent frames, and the sparsity characteristics of the speech signal. This is crucial for signal analysis and recovery.

[0096] In one possible embodiment, referring to Figure 5 , Figure 5 This is a third flow chart of the second embodiment of the voice data loss recovery method of the present application. In this embodiment, step B21 may include steps B211 to B213:

[0097] Step B211, obtaining a voice activity detection result of the lost frame voice data;

[0098] Step B212, determining an initial measurement matrix based on the time-frequency characteristics of the lost frame speech data;

[0099] Step B213: Optimize the initial measurement matrix according to the voice activity detection result, the valid frame data adjacent to the lost frame voice data, and the target sparse transformation basis to obtain a target measurement matrix.

[0100] It is easy to understand that the above-mentioned voice activity detection result can be used to determine whether there is valid voice activity in the speech signal, for example, to distinguish between speech segments and silence or background noise segments. In speech data loss recovery, the voice activity detection result can help determine the speech characteristics of the lost frame (such as whether it is a speech-active frame or a non-speech-active frame, or a voiced or unvoiced frame), thereby providing a basis for constructing a more accurate measurement matrix. For example, the corresponding measurement matrix construction methods may differ for speech-active frames and non-speech-active frames.

[0101] The time-frequency characteristics described above represent the characteristic representations of the lost frame's speech data in the time and frequency domains, including the speech signal's energy distribution and how its frequency components change over time. By analyzing the time-frequency characteristics of the lost frame's speech data, we can understand the signal's primary characteristics and changing patterns, providing a reference for determining the initial observation matrix.

[0102] Illustratively, in this embodiment, the initial measurement matrix may include a random matrix, such as a Gaussian random matrix or a Bernoulli random matrix; and a structured matrix, such as a partial Fourier matrix or a Toeplitz matrix.

[0103] The elements in the Gaussian random matrix obey the independent and identically distributed Gaussian distribution N(0, 1 / m) (m is the number of measurements). It satisfies the restricted isometry property (RIP) with high probability, theoretically ensuring the accurate reconstruction of sparse signals and is suitable for sparse representation of general speech without prior knowledge. The elements in the Bernoulli random matrix have equal probability of taking values ​​of Its calculations only require addition, subtraction, and scaling operations, making it highly efficient and low-complexity in hardware. However, it requires more measurements to achieve reconstruction performance similar to that of a Gaussian matrix.

[0104] Some Fourier matrices in the structured matrix class can be constructed by randomly selecting m rows from the Discrete Fourier Transform (DFT) matrix. This exploits the frequency-domain sparsity of speech signals, making measurements equivalent to frequency-domain sampling and highly compatible with frequency-domain sparse bases such as DCT and wavelets. Toeplitz matrices can be generated from random vectors, forming a block circulant matrix that satisfies A_{i,j} = a_{ij}. This approach can be used to rapidly reduce computational complexity through the Fast Fourier Transform (FFT), making it suitable for large-scale data processing.

[0105] Therefore, in the specific implementation, when the theoretical optimal performance is adopted and the computing resources are sufficient, the initial measurement matrix can be a Gaussian random matrix; when the transmission system hardware complexity is required to be low, the initial measurement matrix can be a Bernoulli matrix; when the amount of voice transmission data is large and the memory is limited, the initial measurement matrix can be a Toeplitz matrix; when a frequency domain sparse signal is used, the initial measurement matrix can preferentially select a partial Fourier matrix.

[0106] Furthermore, this embodiment can comprehensively consider factors such as voice activity detection results, adjacent valid frame data of the lost frame voice data, and target sparse transformation basis to adjust and optimize the initial measurement matrix to obtain a target measurement matrix that better meets actual signal reconstruction requirements.

[0107] For example, this embodiment can optimize the initial measurement matrix by minimizing the cross-correlation μ between the measurement matrix and the target sparse transform basis, enabling it to work in conjunction with the target sparse transform basis. For example, when using the discrete cosine transform (DCT) as the target sparse transform basis, a portion of the Fourier matrix can be selected as the measurement matrix. This matrix has low cross-correlation with the DCT sparse basis, effectively improving the reconstruction of lost speech frames and reducing reconstruction errors.

[0108] In addition to minimizing the cross-correlation μ mentioned above, the measurement matrix can also be adaptively adjusted based on the characteristics of the speech signal. For example, the measurement matrix parameters can be dynamically adjusted for different speech frame types (voiced, unvoiced, or silent). For voiced frames, the measurement density in the low-frequency region can be increased; for unvoiced frames, the measurement density in the high-frequency region can be enhanced. For another example, if the voice activity detection results indicate that the lost frame is a voice activity frame, the sampling density of the measurement matrix in the voice activity frequency band can be increased to better capture useful information in the speech signal.

[0109] In addition, this embodiment can also adjust the weights based on the energy distribution of adjacent valid frame data. For example, if the low-frequency ratio of adjacent valid frame data is greater than 60%, the low-frequency measurement weight can be increased in the initial measurement matrix. At the same time, an iterative optimization algorithm is used to update the measurement matrix. For example, according to the matrix parameter update formula, the measurement matrix is ​​continuously optimized to better cooperate with the sparse basis and improve the reconstruction quality. The matrix parameter update formula can be expressed as:

[0110] ; (1)

[0111] in, represents the updated observation matrix, Represents the observation matrix at the tth iteration before the update; η is the learning rate, which plays a role in controlling the update step size during the optimization process. By adjusting the value of η, the convergence speed and stability of the algorithm can be balanced; e represents the reconstruction error, which is given by the formula Calculated, is the original signal, is the initial observation matrix, α is the sparse coefficient vector, and the reconstruction error reflects the degree of deviation of the reconstruction of the original signal by the current observation matrix and the sparse coefficient vector; It is the transpose of the sparse coefficient vector α, which participates in the operation in the formula and is the same as Together, the initial observation matrix can be iteratively updated in the direction of reducing the reconstruction error.

[0112] Furthermore, this embodiment can also use machine learning algorithms to optimize the combination of the observation matrix and the sparse basis. Through a large number of speech data samples, the optimal observation matrix parameters that match various types of sparse transformation bases can be learned, thereby ensuring that the collaborative work of the observation matrix and the sparse transformation basis achieves the optimal effect.

[0113] In this embodiment, by comprehensively considering multiple factors to optimize the initial measurement matrix, the target measurement matrix can more accurately reflect the characteristics and correlation of the lost frame speech data, thereby improving the accuracy of signal solution and reconstruction, solving the problem that the initial measurement matrix may deviate from the actual signal reconstruction requirements, and further improving the effect and quality of speech data loss recovery processing.

[0114] Step B22: performing signal solution based on the target measurement matrix and the sparse speech signal by a preset optimization algorithm to obtain the lost frame sparse coefficients.

[0115] It is understood that the aforementioned preset optimization algorithm can be a pre-defined algorithm for solving signal sparse coefficients, such as the orthogonal matching pursuit algorithm and the compressed sampling matching pursuit algorithm. These algorithms can gradually approximate and solve the sparse coefficients corresponding to the lost frame speech data through iterative calculations based on the observation matrix and the sparse speech signal, providing key parameters for speech reconstruction.

[0116] In this implementation, to obtain the lost frame sparse coefficients corresponding to the sparse speech signal, a target measurement matrix is ​​constructed based on the adjacent valid frame data of the lost frame speech data and a target sparse transform basis. Then, using a preset optimization algorithm, this target measurement matrix and the sparse speech signal are used for signal solution to obtain the lost frame sparse coefficients. By fully utilizing the adjacent valid frame data and the target sparse transform basis to construct the target measurement matrix and combining it with the optimization algorithm for signal solution, this solution can more accurately obtain the lost frame sparse coefficients, thereby improving the accuracy of speech recovery. This solves the problem of traditional methods' difficulty in accurately solving the lost frame sparse coefficients, further enhancing the effectiveness and quality of speech data recovery.

[0117] Step B3: reconstructing the framed speech data into preset speech data according to the lost frame sparse coefficient to obtain initial speech recovery data.

[0118] It is easy to understand that this embodiment can use the obtained lost frame sparse coefficients to reconstruct and restore the lost part of the framed speech data according to a preset speech reconstruction method (such as a reconstruction algorithm based on compressed sensing) to generate initial speech recovery data.

[0119] In a feasible implementation manner, in this embodiment, step B3 may include steps B31 to B33:

[0120] Step B31, obtaining the inverse transformation matrix corresponding to the target sparse transformation basis;

[0121] Step B32, reconstructing a preset speech based on the lost frame sparse coefficient and the inverse transformation matrix to obtain a lost reconstructed speech frame;

[0122] Step B33: generating initial speech recovery data based on the lost reconstructed speech frame and the framed speech data.

[0123] It should be noted that the aforementioned inverse transform matrix can be the inverse transform matrix corresponding to the target sparse transform basis. This matrix is ​​used to map the "compressed" speech features in the sparse domain (retaining only a small number of non-zero coefficients) back to the time domain, restoring the complete speech signal waveform. For example, if the target sparse transform basis is the discrete cosine transform basis, then the corresponding inverse transform matrix is ​​the inverse discrete cosine transform matrix. By multiplying the sparse coefficients with the inverse transform matrix, the sparsely represented speech signal can be restored to the time domain speech waveform.

[0124] Then, according to a pre-set speech reconstruction method and algorithm, the lost speech frame can be reconstructed and restored using information such as the lost frame sparse coefficients and the inverse transformation matrix to generate a lost reconstructed speech frame. For example, this embodiment can perform a matrix multiplication operation on the lost frame sparse coefficients and the inverse transformation matrix to convert the signal in the sparse domain back into a time domain speech signal to obtain a lost reconstructed speech frame. The lost reconstructed speech frame restores the speech characteristics of the lost speech frame to a certain extent, but may still contain certain errors or imperfections. It needs to be combined with the originally received framed speech data, for example, by replacing the lost frame position in the original signal with the reconstructed time domain speech frame to generate initial speech recovery data containing the restored lost portion, thereby completing the integrity restoration of the speech signal and providing a basis for subsequent filtering processing.

[0125] Furthermore, because signal noise is typically distributed throughout the frequency domain, while the sparse coefficients of speech are concentrated in specific regions, threshold filtering can effectively separate signal from noise. For example, in noisy speech, the signal-to-noise ratio (SNR) of sparse coefficients is typically 10-15 dB higher than that of non-sparse coefficients. In this embodiment, the aforementioned preset filtering network can further learn the distribution patterns of sparse coefficients. This predicts the sparse coefficient distribution of lost frames, assists in compressed sensing reconstruction, and improves recovery accuracy under low signal-to-noise ratio conditions.

[0126] Therefore, in this embodiment, the sparse coefficients of the above-mentioned speech signal can not only be used to achieve high-quality speech reconstruction under low bandwidth through compressed sensing, breaking through the traditional interpolation method's reliance on redundant data, but can also further provide efficient feature representation for subsequent deep learning-based speech filtering, further improving noise suppression and adaptive processing capabilities.

[0127] In this implementation, the sparse representation in the frequency / transform domain is "decompressed" back to the time domain through an inverse transform of the sparse basis. The sparse coefficients preserve the core energy distribution of the speech signal, while the inverse transform restores it to a perceptible speech waveform, enabling efficient recovery of lost frames at low measurement rates.

[0128] Therefore, the embodiment directly aims at the packet loss scene, reconstructs the lost frame from the sparse domain through compressed sensing, recovers the original speech characteristics, utilizes the speech signal sparsity, accurately reconstructs through a small amount of observation data, reduces the calculation redundancy, infers the characteristics of the lost frame from multiple frames of effective data through compressed sensing, and relies on the space-time correlation and dynamic sparsity of adjacent frames. Through the sparse representation and reconstruction algorithm, the embodiment more effectively utilizes the sparse characteristics of the speech signal, improves the accuracy of the lost speech frame recovery, solves the problem of poor recovery effect of the traditional interpolation method, and further improves the effect and quality of the speech data loss recovery processing.

[0129] In the embodiment, when signal loss is detected, a target sparse transform basis corresponding to the detected lost frame speech data is acquired. The lost frame speech data is projected to a sparse domain according to the local sparsity of the lost frame speech data and the target sparse transform basis, and a sparse speech signal is obtained. A voice activity detection result of the lost frame speech data is acquired. An initial observation matrix is determined based on the time-frequency characteristics of the lost frame speech data. The initial observation matrix is optimized based on the voice activity detection result, adjacent effective frame data of the lost frame speech data, and the target sparse transform basis, and a target observation matrix is obtained. Signal solving is performed based on the target observation matrix and the sparse speech signal through a preset optimization algorithm, and lost frame sparse coefficients are obtained. An inverse transform matrix corresponding to the target sparse transform basis is acquired. A preset speech reconstruction is performed based on the lost frame sparse coefficients and the inverse transform matrix, and a lost reconstructed speech frame is obtained. Initial speech recovery data is generated based on the lost reconstructed speech frame and the frame speech data. The embodiment directly aims at the packet loss scene, reconstructs the lost frame from the sparse domain through compressed sensing, recovers the original speech characteristics, utilizes the speech signal sparsity, accurately reconstructs through a small amount of observation data, reduces the calculation redundancy, infers the characteristics of the lost frame from multiple frames of effective data through compressed sensing, and relies on the space-time correlation and dynamic sparsity of adjacent frames. Therefore, the embodiment more effectively utilizes the sparse characteristics of the speech signal, improves the accuracy of the lost speech frame recovery, solves the problem of poor recovery effect of the traditional interpolation method, and further improves the effect and quality of the speech data loss recovery processing.

[0130] It should be noted that the above examples are only used for understanding the present application and do not constitute a limitation on the speech data loss recovery processing method of the present application. More forms of simple transformation based on this technical concept are within the protection scope of the present application.

[0131] The present application also provides a speech data loss recovery processing device, which is described with reference to Figure 6 , Figure 6 The present application also provides a speech data loss recovery processing device, which is described with reference to

[0132] The voice framing module 601 is used to perform framing processing on the received voice message data to obtain framed voice data;

[0133] The speech reconstruction module 602 is configured to perform speech sparse reconstruction based on compressed sensing on the framed speech data to obtain initial speech recovery data when signal loss is detected;

[0134] The filtering module 603 is configured to perform dynamic speech filtering on the initial speech recovery data through a preset filtering network to obtain target speech recovery data.

[0135] As an implementable method, in this embodiment, the speech reconstruction module 602 is also used to convert the detected lost frame speech data into a sparse speech signal when signal loss is detected; obtain the lost frame sparse coefficient corresponding to the sparse speech signal; and perform preset speech reconstruction on the framed speech data according to the lost frame sparse coefficient to obtain initial speech recovery data.

[0136] As an implementable method, in this embodiment, the speech reconstruction module 602 is also used to obtain a target sparse transformation basis corresponding to the detected lost frame speech data; and project the lost frame speech data into a sparse domain according to the local sparsity of the lost frame speech data and the target sparse transformation basis to obtain a sparse speech signal.

[0137] As an implementable method, in this embodiment, the speech reconstruction module 602 is also used to construct a target observation matrix corresponding to the lost frame speech data based on the adjacent valid frame data of the lost frame speech data and the target sparse transformation basis; and perform signal solution based on the target observation matrix and the sparse speech signal through a preset optimization algorithm to obtain the lost frame sparse coefficient.

[0138] As an implementable embodiment, in this embodiment, the speech reconstruction module 602 is further used to obtain the speech activity detection result of the lost frame speech data; determine the initial observation matrix based on the time-frequency characteristics of the lost frame speech data; and optimize the initial observation matrix according to the speech activity detection result, the adjacent valid frame data of the lost frame speech data, and the target sparse transformation basis to obtain the target observation matrix.

[0139] As an implementable method, in this embodiment, the speech reconstruction module 602 is also used to obtain the inverse transformation matrix corresponding to the target sparse transformation basis; perform preset speech reconstruction based on the lost frame sparse coefficients and the inverse transformation matrix to obtain the lost reconstructed speech frame; and generate initial speech recovery data based on the lost reconstructed speech frame and the framed speech data.

[0140] As an implementable embodiment, in this embodiment, the filtering module 603 is also used to input the initial speech recovery data into a preset filtering network to obtain adaptive filtering parameters; based on the adaptive filtering parameters and the dynamic filter, the initial speech recovery data is dynamically filtered to obtain target speech recovery data.

[0141] The voice data loss recovery device provided in this application utilizes the voice data loss recovery method described in the aforementioned embodiments to address the issue of voice quality degradation caused by message loss in VoIP communications. Compared to the prior art, the voice data loss recovery device provided in this application achieves the same beneficial effects as the voice data loss recovery method described in the aforementioned embodiments. Other technical features of the voice data loss recovery device are the same as those disclosed in the aforementioned embodiments and are not further detailed here.

[0142] The present application provides a voice data loss recovery processing device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the voice data loss recovery processing method of the above-mentioned embodiment 1.

[0143] Reference below Figure 7 , which shows a schematic diagram of the structure of a voice data loss recovery processing device suitable for implementing the embodiments of the present application. The voice data loss recovery processing device in the embodiments of the present application can include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The voice data loss recovery processing device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0144] like Figure 7As shown, the voice data loss recovery device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. RAM 1004 also stores various programs and data required for the operation of the voice data loss recovery device. Processing device 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the voice data loss recovery processing device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a voice data loss recovery processing device with various systems, it should be understood that implementation or presence of all the illustrated systems is not required. More or fewer systems may alternatively be implemented or present.

[0145] In particular, according to the embodiments disclosed herein, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed herein include a voice data loss recovery program product, which includes a voice data loss recovery program carried on a computer-readable medium, the voice data loss recovery program containing program code for executing the method illustrated in the flowchart. In such an embodiment, the voice data loss recovery program can be downloaded and installed from a network via a communication device, or installed from storage device 1003 or read-only memory 1002. When the voice data loss recovery program is executed by processing device 1001, the aforementioned functions defined in the method of the embodiments disclosed herein are performed.

[0146] The voice data loss recovery device provided in this application utilizes the voice data loss recovery method described in the aforementioned embodiment, resolving the technical issue of existing voice data loss recovery processes affecting voice transmission rates. Compared to the prior art, the beneficial effects of the voice data loss recovery device provided in this application are the same as those of the voice data loss recovery method described in the aforementioned embodiment. Other technical features of the voice data loss recovery device are the same as those disclosed in the aforementioned embodiment and are not further elaborated here.

[0147] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0148] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0149] The present application provides a storage medium having computer-readable program instructions (ie, a voice data loss recovery processing program) stored thereon, wherein the computer-readable program instructions are used to execute the voice data loss recovery processing method in the above-mentioned embodiment.

[0150] The storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination thereof. More specific examples of storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this embodiment, the storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system or device. The program code contained on the storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0151] The above storage medium may be included in the voice data loss recovery processing device; or may exist independently without being assembled into the voice data loss recovery processing device.

[0152] The storage medium carries one or more programs. When the one or more programs are executed by the voice data loss recovery processing device, the voice data loss recovery processing device performs voice data loss recovery processing.

[0153] The voice data loss recovery processing program code for performing the operations of the present application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, via the Internet using an Internet service provider).

[0154] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the systems, methods, and voice data loss recovery processing program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in an order different from that marked in the accompanying drawings. For example, two blocks shown in succession may actually be executed substantially in parallel, or they may sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or may be implemented using a combination of dedicated hardware and computer instructions.

[0155] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.

[0156] The readable storage medium provided in this application is a storage medium storing computer-readable program instructions (i.e., a voice data loss recovery program) for executing the aforementioned voice data loss recovery method. This storage medium can address the technical issue of existing voice data loss recovery processes affecting voice transmission rates. Compared to the prior art, the beneficial effects of the storage medium provided in this application are similar to those of the voice data loss recovery method provided in the aforementioned embodiments, and are not further elaborated here.

[0157] The above are only some embodiments of the present application and are not intended to limit the patent scope of the present application. All equivalent structural transformations made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A method for recovering voice data loss, characterized in that: The method comprises: Performing frame processing on the received voice message data to obtain framed voice data; When signal loss is detected, performing speech sparse reconstruction based on compressed sensing on the framed speech data to obtain initial speech recovery data; Performing dynamic speech filtering on the initial speech recovery data through a preset filtering network to obtain target speech recovery data; When signal loss is detected, the step of performing speech sparse reconstruction based on compressed sensing on the framed speech data to obtain initial speech recovery data includes: When signal loss is detected, converting the detected lost frame speech data into a sparse speech signal; Obtaining a lost frame sparse coefficient corresponding to the sparse speech signal; Reconstructing the framed speech data into a preset speech according to the lost frame sparse coefficient to obtain initial speech recovery data; The step of converting the detected lost frame voice data into a sparse voice signal comprises: Obtaining a target sparse transform basis corresponding to the detected lost frame speech data; Projecting the lost frame speech data into a sparse domain according to the local sparsity of the lost frame speech data and the target sparse transformation basis to obtain a sparse speech signal; The step of obtaining the lost frame sparse coefficient corresponding to the sparse speech signal includes: Constructing a target measurement matrix corresponding to the lost frame speech data based on adjacent valid frame data of the lost frame speech data and the target sparse transformation basis; Performing signal solution based on the target measurement matrix and the sparse speech signal by a preset optimization algorithm to obtain sparse coefficients of lost frames; The step of constructing a target observation matrix corresponding to the lost frame speech data based on adjacent valid frame data of the lost frame speech data and the target sparse transform basis includes: Obtaining a voice activity detection result of the lost frame voice data; Determining an initial observation matrix based on the time-frequency characteristics of the lost frame speech data; The initial measurement matrix is ​​optimized according to the voice activity detection result, the adjacent valid frame data of the lost frame voice data and the target sparse transformation basis to obtain a target measurement matrix.

2. The voice data loss recovery processing method according to claim 1, wherein: The step of reconstructing the framed speech data according to the lost frame sparse coefficient to obtain initial speech recovery data includes: Obtaining an inverse transformation matrix corresponding to the target sparse transformation basis; Reconstructing a preset voice based on the lost frame sparse coefficient and the inverse transformation matrix to obtain a lost reconstructed voice frame; Initial speech recovery data is generated based on the lost reconstructed speech frame and the framed speech data.

3. The voice data loss recovery processing method according to claim 1, wherein: The step of performing dynamic speech filtering on the initial speech recovery data through a preset filtering network to obtain target speech recovery data includes: Inputting the initial speech recovery data into a preset filter network to obtain adaptive filter parameters; Dynamic speech filtering is performed on the initial speech restoration data based on the adaptive filtering parameters and the dynamic filter to obtain target speech restoration data.

4. A voice data loss recovery processing device, characterized in that: The voice data loss recovery processing device comprises: The voice framing module is used to perform framing processing on the received voice message data to obtain framed voice data; A speech reconstruction module is used to perform speech sparse reconstruction based on compressed sensing on the framed speech data to obtain initial speech recovery data when signal loss is detected; A filtering module, configured to perform dynamic speech filtering on the initial speech recovery data through a preset filtering network to obtain target speech recovery data; The speech reconstruction module is further configured to, when signal loss is detected, convert the detected lost frame speech data into a sparse speech signal; obtain a lost frame sparse coefficient corresponding to the sparse speech signal; and perform a preset speech reconstruction on the framed speech data according to the lost frame sparse coefficient to obtain initial speech recovery data; The speech reconstruction module is further configured to obtain a target sparse transformation basis corresponding to the detected lost frame speech data; project the lost frame speech data into a sparse domain based on the local sparsity of the lost frame speech data and the target sparse transformation basis to obtain a sparse speech signal; The speech reconstruction module is further configured to construct a target measurement matrix corresponding to the lost frame speech data based on adjacent valid frame data of the lost frame speech data and the target sparse transformation basis; perform signal solution based on the target measurement matrix and the sparse speech signal using a preset optimization algorithm to obtain sparse coefficients of the lost frame; The speech reconstruction module is further configured to obtain a speech activity detection result of the lost frame speech data; determine an initial measurement matrix based on the time-frequency characteristics of the lost frame speech data; and optimize the initial measurement matrix according to the speech activity detection result, adjacent valid frame data of the lost frame speech data, and the target sparse transformation basis to obtain a target measurement matrix.

5. A voice data loss recovery processing device, characterized in that: The device includes: a memory, a processor, and a voice data loss recovery processing program stored in the memory and executable on the processor, wherein the voice data loss recovery processing program is configured to implement the steps of the voice data loss recovery processing method according to any one of claims 1 to 3.

6. A storage medium, characterized in that The storage medium stores a voice data loss recovery processing program, which, when executed by a processor, implements the steps of the voice data loss recovery processing method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Speech coding method based on compressed sensing and sparse representation

    CN103778919A

  • Audio restoration method and system based on low consistency dictionary and sparse expression

    CN107039042A