Voice data loss recovery processing method and device, equipment and storage medium

The lost voice data is recovered through frame processing and compression perception technology, and dynamic filtering is used to solve the bandwidth waste and voice quality problems caused by redundant transmission in VoIP communication, achieving efficient recovery and improvement of voice quality.

CN120544589AActive Publication Date: 2025-08-26SHENZHEN DINSTAR TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510992158.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-08-26
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

The prior art ensures voice quality by redundant transmission of redundant voice data in VoIP communication, resulting in an increase in bandwidth demand and intensified network congestion, which cannot meet the needs of high-quality voice transmission rates.

Method used

Framework processing and speech sparse reconstruction technology based on compression perception are adopted, and dynamic speech filtering is performed in combination with a preset filtering network to recover lost speech data.

Benefits of technology

Effectively save bandwidth, improve voice quality and transmission rate, and meet high-quality VoIP communication needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544589A_ABST
    Figure CN120544589A_ABST
Patent Text Reader

Abstract

The invention discloses a voice data loss recovery processing method, device and equipment and a storage medium, and relates to the technical field of voice transmission, and the method comprises the steps: carrying out the framing processing of received voice message data, and obtaining the framed voice data; when detecting that the signal is lost, performing voice reconstruction based on compressed sensing on the framed voice data to obtain initial voice recovery data; and performing dynamic voice filtering on the initial voice recovery data through a preset filtering network to obtain target voice recovery data. According to the method, the voice message data is framed, the voice is reconstructed based on compressed sensing by using the sparsity of the voice signal when the signal is lost, then dynamic filtering is performed through the preset filtering network, and finally the high-quality voice data is output. The problems of bandwidth waste caused by redundant transmission and poor traditional interpolation filtering effect in the prior art are solved, the effects of saving bandwidth, efficiently recovering lost voice frames and improving voice quality are achieved, and the high-quality voice communication requirement is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of voice transmission technology, and in particular to a method, apparatus, device, and storage medium for recovering lost voice data. Background Art

[0002] In the field of Voice over Internet Protocol (VoIP) communications, voice packet loss is a common problem. Currently, redundancy is commonly used to ensure voice quality, specifically by transmitting redundant voice data. While this approach can compensate for packet loss to a certain extent, it has a serious drawback: it significantly increases bandwidth requirements and leads to increased network congestion. Given limited network bandwidth resources, excessive redundant data transmission consumes significant bandwidth, disrupting the normal operation of other network services and degrading overall network performance.

[0003] Therefore, how to ensure the voice transmission quality while ensuring the voice transmission rate has become an urgent problem to be solved. Summary of the Invention

[0004] The main purpose of this application is to provide a voice data loss recovery processing method, device, equipment and storage medium, aiming to solve the technical problem of how to ensure the voice transmission quality while ensuring the voice transmission rate.

[0005] To achieve the above objectives, the present application proposes a method for recovering lost voice data, which includes: Performing frame processing on the received voice message data to obtain framed voice data; When signal loss is detected, performing speech sparse reconstruction based on compressed sensing on the framed speech data to obtain initial speech recovery data; The initial speech restoration data is subjected to dynamic speech filtering through a preset filtering network to obtain target speech restoration data.

[0006] In one embodiment, when signal loss is detected, the step of performing compressed sensing-based speech sparse reconstruction on the framed speech data to obtain initial speech recovery data includes: When signal loss is detected, converting the detected lost frame speech data into a sparse speech signal; Obtaining a lost frame sparse coefficient corresponding to the sparse speech signal; The framed speech data is reconstructed using a preset speech configuration according to the lost frame sparse coefficient to obtain initial speech recovery data.

[0007] In one embodiment, the step of converting the detected lost frame speech data into a sparse speech signal includes: Obtaining a target sparse transform basis corresponding to the detected lost frame speech data; The lost frame speech data is projected onto a sparse domain according to the local sparsity of the lost frame speech data and the target sparse transformation basis to obtain a sparse speech signal.

[0008] In one embodiment, the step of obtaining the lost frame sparse coefficient corresponding to the sparse speech signal includes: Constructing a target measurement matrix corresponding to the lost frame speech data based on adjacent valid frame data of the lost frame speech data and the target sparse transformation basis; Signal solution is performed based on the target observation matrix and the sparse speech signal through a preset optimization algorithm to obtain the lost frame sparse coefficients.

[0009] In one embodiment, the step of constructing a target measurement matrix corresponding to the lost frame speech data based on adjacent valid frame data of the lost frame speech data and the target sparse transform basis includes: Obtaining a voice activity detection result of the lost frame voice data; Determining an initial observation matrix based on the time-frequency characteristics of the lost frame speech data; The initial measurement matrix is ​​optimized according to the voice activity detection result, the adjacent valid frame data of the lost frame voice data and the target sparse transformation basis to obtain a target measurement matrix.

[0010] In one embodiment, the step of reconstructing the framed speech data according to the lost frame sparse coefficient to obtain initial speech recovery data includes: Obtaining an inverse transformation matrix corresponding to the target sparse transformation basis; Reconstructing a preset voice based on the lost frame sparse coefficient and the inverse transformation matrix to obtain a lost reconstructed voice frame; Initial speech recovery data is generated based on the lost reconstructed speech frame and the framed speech data.

[0011] In one embodiment, the step of performing dynamic speech filtering on the initial speech recovery data through a preset filtering network to obtain target speech recovery data includes: Inputting the initial speech recovery data into a preset filter network to obtain adaptive filter parameters; Dynamic speech filtering is performed on the initial speech restoration data based on the adaptive filtering parameters and the dynamic filter to obtain target speech restoration data.

[0012] In addition, to achieve the above-mentioned purpose, the present application also proposes a voice data loss recovery processing device, which includes: The voice framing module is used to perform framing processing on the received voice message data to obtain framed voice data; A speech reconstruction module is used to perform speech sparse reconstruction based on compressed sensing on the framed speech data to obtain initial speech recovery data when signal loss is detected; The filtering module is used to perform dynamic speech filtering on the initial speech recovery data through a preset filtering network to obtain target speech recovery data.

[0013] In addition, to achieve the above-mentioned purpose, the present application also proposes a voice data loss recovery processing device, which includes: a memory, a processor, and a voice data loss recovery processing program stored in the memory and executable on the processor, wherein the voice data loss recovery processing program is configured to implement the steps of the voice data loss recovery processing method as described above.

[0014] In addition, to achieve the above-mentioned objectives, the present application also proposes a storage medium, which is a storage medium, on which is stored a voice data loss recovery processing program, and when the voice data loss recovery processing program is executed by a processor, the steps of the voice data loss recovery processing method described above are implemented. The present application provides a voice data loss recovery processing method, apparatus, device, and storage medium.

[0015] The present application discloses a method, apparatus, device and storage medium for voice data loss recovery processing, the method comprising: performing frame processing on received voice message data to obtain framed voice data; when signal loss is detected, performing voice reconstruction based on compressed sensing on the framed voice data to obtain initial voice recovery data; and performing dynamic voice filtering on the initial voice recovery data through a preset filtering network to obtain target voice recovery data. The present application first performs frame processing on the voice message data, and when the signal is lost, reconstructs the voice based on compressed sensing using the sparsity of the voice signal, and then dynamically filters the voice through a preset filtering network to finally output high-quality voice data. This solves the problems in the prior art of redundant transmission leading to bandwidth waste and poor traditional interpolation filtering effects, and achieves the effects of saving bandwidth, efficiently recovering lost voice frames and improving voice quality, meeting the needs of high-quality voice communication. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0018] Figure 1 This is a first flow chart of the first embodiment of the voice data loss recovery method of the present application; Figure 2 This is a second flow chart of the first embodiment of the voice data loss recovery method of the present application; Figure 3 This is a first flow chart of the second embodiment of the voice data loss recovery method of the present application; Figure 4 This is a second flow chart of the second embodiment of the voice data loss recovery method of the present application; Figure 5 This is a third flow chart of the second embodiment of the voice data loss recovery method of the present application; Figure 6 This is a schematic diagram of the module structure of the voice data loss recovery processing device according to an embodiment of the present application; Figure 7 This is a schematic diagram of the device structure of the hardware operating environment involved in the voice data loss recovery processing method in the embodiment of the present application.

[0019] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0020] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0021] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0022] The main solution of this application is: to perform frame processing on the received voice message data to obtain framed voice data; when signal loss is detected, to perform voice sparse reconstruction based on compressed sensing on the framed voice data to obtain initial voice recovery data; to perform dynamic voice filtering on the initial voice recovery data through a preset filtering network to obtain target voice recovery data.

[0023] While the current method of transmitting redundant voice data to ensure voice quality can compensate for packet loss to some extent, it has a serious drawback: it significantly increases bandwidth requirements and exacerbates network congestion. Given limited network bandwidth resources, excessive redundant data transmission consumes significant bandwidth, disrupting other network services and degrading overall network performance.

[0024] This application proposes a method for recovering lost voice data. This method first frames voice message data, then exploits the sparsity of voice signals in the event of signal loss to reconstruct the voice using compressed sensing. The data is then dynamically filtered through a preset filtering network, ultimately outputting high-quality voice data. This method addresses the existing issues of redundant transmission leading to bandwidth waste and poor performance of traditional interpolation filtering. It achieves bandwidth savings, efficiently recovers lost voice frames, and improves voice quality, meeting the needs of high-quality VoIP communication.

[0025] It should be noted that the execution entity of this embodiment may be a voice data loss recovery processing system, a computing service device with data processing, network communication, and program execution capabilities, such as a tablet computer, personal computer, or mobile phone, or a voice data loss recovery processing device capable of performing the aforementioned functions, and this embodiment does not specifically limit this. The following description of this embodiment and the following embodiments uses a voice data loss recovery processing device (hereinafter referred to as the processing device) as an example execution entity.

[0026] Based on this, the embodiment of the present application provides a method for recovering lost voice data. Figure 1 , Figure 1 This is a first flow chart of the first embodiment of the voice data loss recovery processing method of the present application.

[0027] In this embodiment, the voice data loss recovery method includes steps S10 to S30: Step S10, performing frame processing on the received voice message data to obtain framed voice data; Step S20, when signal loss is detected, performing speech sparse reconstruction based on compressed sensing on the framed speech data to obtain initial speech recovery data; Step S30 , performing dynamic speech filtering on the initial speech recovery data through a preset filtering network to obtain target speech recovery data.

[0028] It is understood that the aforementioned voice message data may be data packets carrying voice information transmitted in VoIP communications, comprising data obtained through sampling, encoding, and other processing of voice signals, and used to transmit voice content between different communication devices. To facilitate subsequent rapid search and recovery of lost frames, this embodiment may segment the continuous voice signal in the voice message data into multiple short-term frames at regular time intervals, such as one frame every 20 ms, to obtain a voice signal comprising multiple voice frames, i.e., the aforementioned framed voice signal.

[0029] It should be understood that the processing device can monitor the framed voice signal in real time during the communication process. Furthermore, this embodiment utilizes compressed sensing technology. Based on the sparse nature of the voice signal and known partial data, upon detecting signal loss in the framed voice signal (e.g., RTP sequence number (RTP) sequence number (RTP) sequence number) interruption), it can quickly reconstruct and recover the lost voice frames in the framed voice signal using a preset mathematical model and algorithm to ensure the integrity of the voice message. In comparison, the amount of data in the lost voice frame is much smaller than the data in the framed voice signal. Therefore, this embodiment can effectively improve the speed of voice data restoration.

[0030] It is understandable that the initially recovered speech data obtained after compressed sensing-based speech reconstruction may contain certain noise or inaccuracies. To ensure the quality of speech transmission, this embodiment may further process the initial speech recovery data. The aforementioned preset filtering network can be a pre-designed network structure for filtering speech signals, typically constructed based on a deep learning model. It can process the input speech data and output filtering parameters that are dynamically adjusted based on the real-time characteristics of the input speech signal to achieve adaptive filtering of the speech signal, remove noise, and enhance the key features of the speech signal, thereby obtaining high-quality target speech recovery data.

[0031] In one possible implementation, refer to Figure 2 , Figure 2 This is a second flow chart of the first embodiment of the voice data loss recovery method of the present application. In this embodiment, step S30 may include steps A1 to A2: Step A1: inputting the initial speech recovery data into a preset filtering network to obtain adaptive filtering parameters; Step A2: performing dynamic speech filtering on the initial speech recovery data based on the adaptive filtering parameters and the dynamic filter to obtain target speech recovery data.

[0032] It will be readily understood that the adaptive filtering parameters described above may be generated in real time by a preset filtering network based on the input initial speech recovery data. These parameters can be dynamically adjusted based on changes in the speech signal and noise environment to achieve optimal filtering of the speech signal. For example, under varying noise interference conditions, the adaptive filtering parameters may alter the filter's frequency response characteristics, gain, and other parameters accordingly to better remove noise while preserving the key characteristics of the speech signal.

[0033] It will be readily understood that in this embodiment, the preset filtering network may employ an adaptive filtering model based on deep learning, such as a recurrent neural network (RNN), a gated recurrent unit (GRU), or a long short-term memory network (LSTM). For example, when the preset filtering network is an LSTM model, it can learn the long-term dependencies and noise characteristics of speech signals. When the acoustic environment of the input speech signal changes, the LSTM model dynamically adjusts its internal parameters, such as the weight matrix and bias vector, based on the real-time characteristics of the input signal (such as spectral characteristics and time-domain waveform). Therefore, this embodiment enables the preset filtering network to process sequence data and capture the temporal characteristics of speech signals. During the model training phase, a large amount of speech data from different acoustic environments can be pre-trained, allowing the iterative preset filtering network to effectively learn the mapping relationship between different noise environments and filtering parameters. Furthermore, the preset filtering network also features a real-time feedback mechanism that can further adjust model parameters based on the quality of the filtered speech signal, forming a closed-loop optimization loop. This allows the network to quickly adapt to different acoustic environments and effectively filter speech signals.

[0034] In practical applications, the potentially noisy initial speech data is fed into an iteratively pre-set filtering network. The pre-set filtering network then generates filter parameters in real time based on the characteristics of the initial speech data (such as spectral characteristics and time-domain waveform). When the acoustic environment changes, the initial speech data can rapidly adjust its parameters to adapt to the new noise conditions. For example, when a burst of noise is detected in the speech signal, the model instantly updates the filter's frequency response, enhancing its ability to suppress noise at specific frequencies. Furthermore, the model possesses online learning capabilities, enabling it to further optimize model parameters based on real-time feedback on filtering performance, thereby better adapting to changing acoustic environments.

[0035] It is understandable that the above-mentioned dynamic filter can be a filter that can adjust its filtering characteristics in real time based on the input signal and adaptive filtering parameters, such as an FIR filter (Finite Impulse Response Filter), which can be used to perform time-domain convolution filtering on the reconstructed frame. Compared with traditional fixed parameter filters, dynamic filters have greater adaptability and flexibility and can adapt to complex and changing voice communication environments and noise conditions. This embodiment can utilize a dynamic filter to adjust the filter characteristics in real time based on the generated adaptive filtering parameters, and perform adaptive filtering on the initial voice recovery data, thereby effectively removing noise interference and enhancing the key information of the voice signal, thereby obtaining high-precision target voice recovery data.

[0036] For example, when a speech signal characteristic changes (such as a sudden change in fundamental frequency), for example, the transition tone from "ah" (ah) to "oh" (oo) appears in the current speech frame in the initial speech recovery data, the fundamental frequency suddenly changes from 120 Hz to 150 Hz, the formant F1 rises from 500 Hz to 700 Hz, and F2 drops from 1200 Hz to 1000 Hz. However, the noise environment is a stable white background noise with a signal-to-noise ratio (SNR) of 15 dB (decibel).

[0037] At this time, if the preset filtering network is an LSTM model, the changing trends of the fundamental frequency and formant can be learned through historical speech frames (such as the "ah" sound in the previous 50ms) during the training process, and the 32nd-order FIR filter coefficient Wold=[0.12,-0.05,0.23,...,0.08] (suitable for the low-frequency "ah" sound) can be output.

[0038] When the LSTM hidden layer captures sudden changes in the fundamental frequency and formants of the current frame in the input feature vector (such as MFCC (Mel-frequency cepstral coefficient) features), for example, a rising trend in the fundamental frequency (ΔF0 = +30 Hz) and a formant shift (ΔF1 = +200 Hz, ΔF2 = -200 Hz) are detected, the LSTM then maps the hidden layer state into filter coefficients through a fully connected layer, updating the coefficients in real time to [-0.03, 0.18, 0.28, ..., 0.15]. The output new filter parameters increase the weights of the 10th to 15th taps (corresponding to the 500-1000 Hz frequency band) to enhance the mid-frequency formants, and decrease the weights of the first five taps (corresponding to the 0-300 Hz frequency band) to suppress the impact noise caused by the sudden change in the low-frequency fundamental frequency. This enhances the mid-frequency components and effectively adapts to the spectral characteristics of the "oh" sound.

[0039] In addition, if burst noise is detected in the speech signal, the preset filtering network can also instantly update the frequency response of the filter to enhance the ability to suppress noise of specific frequencies, thereby achieving the purpose of dynamically adjusting the model parameters to cope with different acoustic environments.

[0040] For example, if a sudden change in the noise environment occurs (such as a sudden car horn), the current frame of the normal speech content "Please wait" contains the character "děng" (wait), and the spectrum may be concentrated between 400 and 800 Hz. However, a sudden car horn superimposes on the background white noise (a sudden narrowband noise burst at 1500 Hz, causing the SNR to drop sharply to 5 dB). In this case, a preset filtering network, such as an RNN model, can output initial coefficients of Wold = [0.05, 0.15, 0.20, ..., 0.04] (a passband covering 200-1000 Hz to suppress high-frequency noise) for the first 100 ms of stable speech data. Correspondingly, after detecting the 1500 Hz noise burst, the RNN model can update the coefficients to Wnew = [0.03, 0.12, 0.18, ..., -0.10] within 5 ms (i.e., introducing a notch filter at 1500 Hz), thereby achieving noise suppression.

[0041] Therefore, the final speech restoration result obtained after dynamic speech filtering processing, that is, the above-mentioned target speech restoration data, has good speech quality, clarity and intelligibility, and can be directly used for playback, providing users with a high-quality and clear voice communication experience.

[0042] This embodiment proposes a method for recovering from voice data loss. The method first frames the voice message data. When the signal is lost, the voice is reconstructed using compressed sensing, leveraging the sparsity of the voice signal. The voice is then adaptively filtered using a preset filtering network and a dynamic filter, ultimately outputting high-quality voice data. Compressed sensing-based voice reconstruction fully exploits the sparsity of the voice signal to accurately reconstruct lost voice frames, preserving the original characteristics of the voice to the greatest extent possible. A deep learning-based adaptive filtering algorithm dynamically adjusts filter parameters based on real-time changes in the voice signal and noise environment, effectively removing noise and significantly improving voice clarity and intelligibility, providing users with a superior voice communication experience.

[0043] At the same time, the powerful adaptive capabilities of the preset filtering network enable the processing device to quickly adapt to the varying speech characteristics of different speakers, complex and variable network noise environments, and bandwidth-sensitive and noise-variable scenarios such as mines and vehicles, maintaining stable and efficient voice processing performance. Therefore, this embodiment solves the existing problems of redundant transmission leading to bandwidth waste and poor traditional interpolation filtering, achieving bandwidth savings, efficient recovery of lost voice frames, and improved voice quality, meeting the needs of high-quality voice communication.

[0044] This embodiment provides a method for recovering voice data loss. The method comprises: framing received voice message data to obtain framed voice data; when signal loss is detected, performing compressed sensing-based voice sparse reconstruction on the framed voice data to obtain initial voice recovery data; inputting the initial voice recovery data into a preset filtering network to obtain adaptive filtering parameters; and performing dynamic voice filtering on the initial voice recovery data based on the adaptive filtering parameters and a dynamic filter to obtain target voice recovery data. This embodiment first frames the voice message data, then reconstructs the voice data based on compressed sensing using the sparsity of the voice signal when signal loss occurs. The data is then adaptively filtered using a preset filtering network and a dynamic filter, ultimately outputting high-quality voice data. The powerful adaptive capabilities of the preset filtering network enable the processing device to quickly adapt to varying voice characteristics between speakers and complex and variable network noise environments, maintaining stable and efficient voice processing performance. Therefore, this embodiment solves the existing problems of redundant transmission leading to bandwidth waste and poor performance of traditional interpolation filtering, saving bandwidth, efficiently recovering lost voice frames, and improving voice quality, meeting the needs of high-quality voice communication.

[0045] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction and will not be repeated later.

[0046] Based on the first embodiment, please refer to Figure 3 , Figure 3 This is a first flow chart of the second embodiment of the voice data loss recovery method of the present application. In this embodiment, step S20 includes steps B1 to B3: Step B1, when signal loss is detected, converting the detected lost frame voice data into a sparse voice signal; It is easy to understand that the lost frame voice data mentioned above refers to voice data frames within the framed voice data that were not successfully received or transmitted during voice communication due to network issues or other reasons, affecting the integrity and continuity of the voice. This embodiment can restore the lost frame voice data through voice reconstruction technology. However, voice signals in specific transform domains (such as the discrete cosine transform domain or the wavelet transform domain) exhibit sparseness, meaning that most of the energy is concentrated in a small number of coefficients, with only a small number of non-zero coefficients or coefficients with large amplitudes. These coefficients can represent the key characteristics of the voice signal and are obtained through a sparse representation of the voice signal. Therefore, to address the problem that existing lost frame recovery technologies fail to fully utilize the complex characteristics of voice signals and struggle to restore the true characteristics of the voice when recovering lost voice frames, this embodiment achieves high-precision recovery of lost frame voice data by converting it to a sparse domain for processing.

[0047] In one possible implementation, refer to Figure 4 , Figure 4This is a second flow chart of the second embodiment of the voice data loss recovery method of the present application. In this embodiment, step B1 may include steps B11 to B12: Step B11, obtaining a target sparse transform basis corresponding to the detected lost frame speech data; Step B12: Projecting the lost frame speech data into a sparse domain according to the local sparsity of the lost frame speech data and the target sparse transformation basis to obtain a sparse speech signal.

[0048] It is understood that this embodiment can perform a sparse transform on the lost frame speech data, converting it to a sparse domain to obtain a sparse speech signal, which can then be used to reconstruct the speech using its sparse characteristics. The target sparse transform basis can be a transform basis suitable for sparse representation, selected based on the characteristics of the lost frame speech data. Common examples include the discrete cosine transform (DCT), wavelet transform, and Fourier transform. Different transform bases can convert the speech signal into different transform domains, highlighting its sparse characteristics.

[0049] Exemplarily, the processing device may first determine that the target sparse transform basis of the lost frame speech data is a wavelet transform basis. Then, the lost frame speech data is subjected to wavelet transform to convert the lost frame speech data from the time domain to the wavelet domain. In the wavelet domain, the energy of the lost frame speech data is mainly concentrated in a small number of wavelet coefficients. By setting a threshold, coefficients with amplitudes greater than the threshold are retained, and coefficients with amplitudes less than the threshold are discarded, thereby achieving sparse representation of the lost frame speech data. At the same time, appropriate wavelet basis functions and decomposition layers can be selected according to the characteristics of the speech signal and the application scenario to improve the effect of sparse representation. For example, for speech signals with different frequency components, orthogonal wavelet basis functions with different characteristics can be selected for transformation to better capture the sparse characteristics of the lost frame speech data.

[0050] Taking the Fourier transform as an example, after converting the lost frame speech data into the frequency domain, the energy of the speech signal is primarily concentrated in certain frequency components. In this case, thresholding the signal in the frequency domain can be performed to retain frequency coefficients with larger amplitudes and discard those with smaller amplitudes, thereby achieving a sparse representation of the speech signal.

[0051] In addition, in the sparse representation process, the time-frequency characteristics of the speech signal can be combined to select appropriate time windows and transformation parameters to better capture the local characteristics of the speech signal and improve the effect of sparse representation.

[0052] It should be understood that the aforementioned local sparsity may refer to the sparse nature of the lost frame speech data within a short time window (e.g., the time span of a single frame of speech signal), such as the periodicity of voiced sounds or the broadband nature of unvoiced sounds. Therefore, in this embodiment, during the process of projecting the lost frame speech data into a sparse domain based on a target sparse transform basis, the energy of the speech signal can be concentrated on a small number of coefficients based on local sparsity. These coefficients can reflect the primary characteristics of the speech signal within a local time span, such as the fundamental frequency and harmonic structure, thereby maximizing the sparse nature of the speech signal and enabling more efficient signal processing and recovery based on the sparse representation.

[0053] In this embodiment, by selecting a suitable target sparse transformation basis and performing domain conversion in combination with the local sparsity of the lost frame speech data, the sparse speech signal can be obtained more accurately, providing a more reliable basis for subsequent speech reconstruction. This solves the problem that direct sparse representation may lead to inaccurate sparse representation and affect the quality of speech recovery, and further improves the accuracy of speech data loss recovery processing.

[0054] Step B2, obtaining the lost frame sparse coefficient corresponding to the sparse speech signal; It is readily understood that the aforementioned lost frame sparse coefficients may be sparse representation coefficients that characterize the characteristics of the lost frame's speech data in a sparse domain. This embodiment, based on known speech frames and compressed sensing theory, constructs an observation matrix based on the spatiotemporal correlation of multiple adjacent frames, capturing the temporal dependencies between the lost frame and preceding and following frames (e.g., speech signal coherence). The sparse speech signal is then processed using a specific algorithm (e.g., an orthogonal matching pursuit algorithm) to extract the sparse coefficients that characterize the lost frame's speech data.

[0055] In a feasible implementation manner, in this embodiment, step B2 may include steps B21 to B22: Step B21, constructing a target measurement matrix corresponding to the lost frame speech data based on adjacent valid frame data of the lost frame speech data and the target sparse transformation basis; It should be noted that the above-mentioned adjacent valid frame data can be the voice frame data that are successfully received or transmitted before and after the lost frame voice data. These data contain voice information that is continuous and relevant in time with the lost frame, and can provide an important reference basis for recovering the lost frame and help infer the voice characteristics and change trends of the lost frame.

[0056] The target measurement matrix is ​​a matrix used to connect the original signal (i.e., the lost frame speech data) and the measured values ​​(i.e., known adjacent valid frame data). Its design and construction directly impact the quality of signal reconstruction. This embodiment uses the target measurement matrix to capture relevant information, such as the temporal dependencies between the lost frame and its preceding and subsequent frames, and the sparsity characteristics of the speech signal. This is crucial for signal analysis and recovery.

[0057] In one possible implementation, refer to Figure 5 , Figure 5 This is a third flow chart of the second embodiment of the voice data loss recovery method of the present application. In this embodiment, step B21 may include steps B211 to B213: Step B211, obtaining a voice activity detection result of the lost frame voice data; Step B212, determining an initial measurement matrix based on the time-frequency characteristics of the lost frame speech data; Step B213: Optimize the initial measurement matrix according to the voice activity detection result, the valid frame data adjacent to the lost frame voice data, and the target sparse transformation basis to obtain a target measurement matrix.

[0058] It is easy to understand that the above-mentioned voice activity detection result can be used to determine whether there is valid voice activity in the speech signal, for example, to distinguish between speech segments and silence or background noise segments. In speech data loss recovery, the voice activity detection result can help determine the speech characteristics of the lost frame (such as whether it is a speech-active frame or a non-speech-active frame, or a voiced or unvoiced frame), thereby providing a basis for constructing a more accurate measurement matrix. For example, the corresponding measurement matrix construction methods may differ for speech-active frames and non-speech-active frames.

[0059] The time-frequency characteristics described above represent the characteristic representations of the lost frame's speech data in the time and frequency domains, including the speech signal's energy distribution and how its frequency components change over time. By analyzing the time-frequency characteristics of the lost frame's speech data, we can understand the signal's primary characteristics and changing patterns, providing a reference for determining the initial observation matrix.

[0060] Illustratively, in this embodiment, the initial measurement matrix may include a random matrix, such as a Gaussian random matrix or a Bernoulli random matrix; and a structured matrix, such as a partial Fourier matrix or a Toeplitz matrix.

[0061] The elements in the Gaussian random matrix obey the independent and identically distributed Gaussian distribution N(0, 1 / m) (m is the number of measurements). It satisfies the restricted isometry property (RIP) with high probability, theoretically ensuring the accurate reconstruction of sparse signals and is suitable for sparse representation of general speech without prior knowledge. The elements in the Bernoulli random matrix have equal probability of taking values ​​of Its calculations only require addition, subtraction, and scaling operations, making it highly efficient and low-complexity in hardware. However, it requires more measurements to achieve reconstruction performance similar to that of a Gaussian matrix.

[0062] Some Fourier matrices in the structured matrix class can be constructed by randomly selecting m rows from the Discrete Fourier Transform (DFT) matrix. This exploits the frequency-domain sparsity of speech signals, making measurements equivalent to frequency-domain sampling and highly compatible with frequency-domain sparse bases such as DCT and wavelets. Toeplitz matrices can be generated from random vectors, forming a block circulant matrix that satisfies A_{i,j} = a_{ij}. This approach can be used to rapidly reduce computational complexity through the Fast Fourier Transform (FFT), making it suitable for large-scale data processing.

[0063] Therefore, in the specific implementation, when the theoretical optimal performance is adopted and the computing resources are sufficient, the initial measurement matrix can be a Gaussian random matrix; when the transmission system hardware complexity is required to be low, the initial measurement matrix can be a Bernoulli matrix; when the amount of voice transmission data is large and the memory is limited, the initial measurement matrix can be a Toeplitz matrix; when a frequency domain sparse signal is used, the initial measurement matrix can preferentially select a partial Fourier matrix.

[0064] Furthermore, this embodiment can comprehensively consider factors such as voice activity detection results, adjacent valid frame data of the lost frame voice data, and target sparse transformation basis to adjust and optimize the initial measurement matrix to obtain a target measurement matrix that better meets actual signal reconstruction requirements.

[0065] For example, this embodiment can optimize the initial measurement matrix by minimizing the cross-correlation μ between the measurement matrix and the target sparse transform basis, enabling it to work in conjunction with the target sparse transform basis. For example, when using the discrete cosine transform (DCT) as the target sparse transform basis, a portion of the Fourier matrix can be selected as the measurement matrix. This matrix has low cross-correlation with the DCT sparse basis, effectively improving the reconstruction of lost speech frames and reducing reconstruction errors.

[0066] In addition to minimizing the cross-correlation μ mentioned above, the measurement matrix can also be adaptively adjusted based on the characteristics of the speech signal. For example, the measurement matrix parameters can be dynamically adjusted for different speech frame types (voiced, unvoiced, or silent). For voiced frames, the measurement density in the low-frequency region can be increased; for unvoiced frames, the measurement density in the high-frequency region can be enhanced. For another example, if the voice activity detection results indicate that the lost frame is a voice activity frame, the sampling density of the measurement matrix in the voice activity frequency band can be increased to better capture useful information in the speech signal.

[0067] In addition, this embodiment can also adjust the weights based on the energy distribution of adjacent valid frame data. For example, if the low-frequency ratio of adjacent valid frame data is greater than 60%, the low-frequency measurement weight can be increased in the initial measurement matrix. At the same time, an iterative optimization algorithm is used to update the measurement matrix. For example, according to the matrix parameter update formula, the measurement matrix is ​​continuously optimized to better cooperate with the sparse basis and improve the reconstruction quality. The matrix parameter update formula can be expressed as: ; (1) in, represents the updated observation matrix, Represents the observation matrix at the tth iteration before the update; η is the learning rate, which plays a role in controlling the update step size during the optimization process. By adjusting the value of η, the convergence speed and stability of the algorithm can be balanced; e represents the reconstruction error, which is given by the formula Calculated, is the original signal, is the initial observation matrix, α is the sparse coefficient vector, and the reconstruction error reflects the degree of deviation of the reconstruction of the original signal by the current observation matrix and the sparse coefficient vector; It is the transpose of the sparse coefficient vector α, which participates in the operation in the formula and is the same as Together, the initial observation matrix can be iteratively updated in the direction of reducing the reconstruction error.

[0068] Furthermore, this embodiment can also use machine learning algorithms to optimize the combination of the observation matrix and the sparse basis. Through a large number of speech data samples, the optimal observation matrix parameters that match various types of sparse transformation bases can be learned, thereby ensuring that the collaborative work of the observation matrix and the sparse transformation basis achieves the optimal effect.

[0069] In this embodiment, by comprehensively considering multiple factors to optimize the initial measurement matrix, the target measurement matrix can more accurately reflect the characteristics and correlation of the lost frame speech data, thereby improving the accuracy of signal solution and reconstruction, solving the problem that the initial measurement matrix may deviate from the actual signal reconstruction requirements, and further improving the effect and quality of speech data loss recovery processing.

[0070] Step B22: performing signal solution based on the target measurement matrix and the sparse speech signal by a preset optimization algorithm to obtain the lost frame sparse coefficients.

[0071] It is understood that the aforementioned preset optimization algorithm can be a pre-defined algorithm for solving signal sparse coefficients, such as the orthogonal matching pursuit algorithm and the compressed sampling matching pursuit algorithm. These algorithms can gradually approximate and solve the sparse coefficients corresponding to the lost frame speech data through iterative calculations based on the observation matrix and the sparse speech signal, providing key parameters for speech reconstruction.

[0072] In this implementation, to obtain the lost frame sparse coefficients corresponding to the sparse speech signal, a target measurement matrix is ​​constructed based on the adjacent valid frame data of the lost frame speech data and a target sparse transform basis. Then, using a preset optimization algorithm, this target measurement matrix and the sparse speech signal are used for signal solution to obtain the lost frame sparse coefficients. By fully utilizing the adjacent valid frame data and the target sparse transform basis to construct the target measurement matrix and combining it with the optimization algorithm for signal solution, this solution can more accurately obtain the lost frame sparse coefficients, thereby improving the accuracy of speech recovery. This solves the problem of traditional methods' difficulty in accurately solving the lost frame sparse coefficients, further enhancing the effectiveness and quality of speech data recovery.

[0073] Step B3: reconstructing the framed speech data into preset speech data according to the lost frame sparse coefficient to obtain initial speech recovery data.

[0074] It is easy to understand that this embodiment can use the obtained lost frame sparse coefficients to reconstruct and restore the lost part of the framed speech data according to a preset speech reconstruction method (such as a reconstruction algorithm based on compressed sensing) to generate initial speech recovery data.

[0075] In a feasible implementation manner, in this embodiment, step B3 may include steps B31 to B33: Step B31, obtaining the inverse transformation matrix corresponding to the target sparse transformation basis; Step B32, reconstructing a preset speech based on the lost frame sparse coefficient and the inverse transformation matrix to obtain a lost reconstructed speech frame; Step B33: generating initial speech recovery data based on the lost reconstructed speech frame and the framed speech data.

[0076] It should be noted that the aforementioned inverse transform matrix can be the inverse transform matrix corresponding to the target sparse transform basis. This matrix is ​​used to map the "compressed" speech features in the sparse domain (retaining only a small number of non-zero coefficients) back to the time domain, restoring the complete speech signal waveform. For example, if the target sparse transform basis is the discrete cosine transform basis, then the corresponding inverse transform matrix is ​​the inverse discrete cosine transform matrix. By multiplying the sparse coefficients with the inverse transform matrix, the sparsely represented speech signal can be restored to the time domain speech waveform.

[0077] Then, according to a pre-set speech reconstruction method and algorithm, the lost speech frame can be reconstructed and restored using information such as the lost frame sparse coefficients and the inverse transformation matrix to generate a lost reconstructed speech frame. For example, this embodiment can perform a matrix multiplication operation on the lost frame sparse coefficients and the inverse transformation matrix to convert the signal in the sparse domain back into a time domain speech signal to obtain a lost reconstructed speech frame. The lost reconstructed speech frame restores the speech characteristics of the lost speech frame to a certain extent, but may still contain certain errors or imperfections. It needs to be combined with the originally received framed speech data, for example, by replacing the lost frame position in the original signal with the reconstructed time domain speech frame to generate initial speech recovery data containing the restored lost portion, thereby completing the integrity restoration of the speech signal and providing a basis for subsequent filtering processing.

[0078] Furthermore, because signal noise is typically distributed throughout the frequency domain, while the sparse coefficients of speech are concentrated in specific regions, threshold filtering can effectively separate signal from noise. For example, in noisy speech, the signal-to-noise ratio (SNR) of sparse coefficients is typically 10-15 dB higher than that of non-sparse coefficients. In this embodiment, the aforementioned preset filtering network can further learn the distribution patterns of sparse coefficients. This predicts the sparse coefficient distribution of lost frames, assists in compressed sensing reconstruction, and improves recovery accuracy under low signal-to-noise ratio conditions.

[0079] Therefore, in this embodiment, the sparse coefficients of the above-mentioned speech signal can not only be used to achieve high-quality speech reconstruction under low bandwidth through compressed sensing, breaking through the traditional interpolation method's reliance on redundant data, but can also further provide efficient feature representation for subsequent deep learning-based speech filtering, further improving noise suppression and adaptive processing capabilities.

[0080] In this implementation, the sparse representation in the frequency / transform domain is "decompressed" back to the time domain through an inverse transform of the sparse basis. The sparse coefficients preserve the core energy distribution of the speech signal, while the inverse transform restores it to a perceptible speech waveform, enabling efficient recovery of lost frames at low measurement rates.

[0081] Therefore, this embodiment directly targets message loss scenarios, using compressed sensing to reconstruct lost frames from a sparse domain and restore the original speech features. It leverages the sparsity of speech signals to accurately reconstruct using a small amount of observation data, reducing computational redundancy. Compressed sensing is used to infer the features of lost frames from multiple frames of valid data, relying on the spatiotemporal correlation and dynamic sparsity of adjacent frames. This embodiment utilizes sparse representation and reconstruction algorithms to more effectively utilize the sparse nature of speech signals, improving the accuracy of lost speech frame recovery and addressing the poor recovery effects of traditional interpolation methods. This further enhances the effectiveness and quality of speech data loss recovery.

[0082] This embodiment discloses that when signal loss is detected, a target sparse transformation basis corresponding to the detected lost frame speech data is obtained; the lost frame speech data is projected into a sparse domain based on the local sparsity of the lost frame speech data and the target sparse transformation basis to obtain a sparse speech signal. A speech activity detection result of the lost frame speech data is obtained; an initial measurement matrix is ​​determined based on the time-frequency characteristics of the lost frame speech data; the initial measurement matrix is ​​optimized based on the speech activity detection result, adjacent valid frame data of the lost frame speech data, and the target sparse transformation basis to obtain a target measurement matrix; a signal solution is performed based on the target measurement matrix and the sparse speech signal using a preset optimization algorithm to obtain lost frame sparse coefficients; an inverse transformation matrix corresponding to the target sparse transformation basis is obtained; a preset speech reconstruction is performed based on the lost frame sparse coefficients and the inverse transformation matrix to obtain a lost reconstructed speech frame; and initial speech recovery data is generated based on the lost reconstructed speech frame and the framed speech data. This embodiment directly targets message loss scenarios, using compressed sensing to reconstruct lost frames from a sparse domain and restore the original speech features. It leverages the sparsity of speech signals to accurately reconstruct with a small amount of observation data, reducing computational redundancy. Compressed sensing is used to infer the features of lost frames from multiple frames of valid data, relying on the spatiotemporal correlation and dynamic sparsity of adjacent frames. Therefore, through sparse representation and reconstruction algorithms, this embodiment more effectively utilizes the sparse nature of speech signals, improving the accuracy of lost speech frame recovery, addressing the poor recovery results of traditional interpolation methods, and further enhancing the effectiveness and quality of speech data loss recovery.

[0083] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the voice data loss recovery processing method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.

[0084] This application also provides a voice data loss recovery processing device, please refer to Figure 6 , Figure 6 This is a schematic diagram of the module structure of the voice data loss recovery processing device according to an embodiment of the present application. In this embodiment, the voice data loss recovery processing device includes: The voice framing module 601 is used to perform framing processing on the received voice message data to obtain framed voice data; The speech reconstruction module 602 is configured to perform speech sparse reconstruction based on compressed sensing on the framed speech data to obtain initial speech recovery data when signal loss is detected; The filtering module 603 is configured to perform dynamic speech filtering on the initial speech recovery data through a preset filtering network to obtain target speech recovery data.

[0085] As an implementable method, in this embodiment, the speech reconstruction module 602 is also used to convert the detected lost frame speech data into a sparse speech signal when signal loss is detected; obtain the lost frame sparse coefficient corresponding to the sparse speech signal; and perform preset speech reconstruction on the framed speech data according to the lost frame sparse coefficient to obtain initial speech recovery data.

[0086] As an implementable method, in this embodiment, the speech reconstruction module 602 is also used to obtain a target sparse transformation basis corresponding to the detected lost frame speech data; and project the lost frame speech data into a sparse domain according to the local sparsity of the lost frame speech data and the target sparse transformation basis to obtain a sparse speech signal.

[0087] As an implementable method, in this embodiment, the speech reconstruction module 602 is also used to construct a target observation matrix corresponding to the lost frame speech data based on the adjacent valid frame data of the lost frame speech data and the target sparse transformation basis; and perform signal solution based on the target observation matrix and the sparse speech signal through a preset optimization algorithm to obtain the lost frame sparse coefficient.

[0088] As an implementable embodiment, in this embodiment, the speech reconstruction module 602 is further used to obtain the speech activity detection result of the lost frame speech data; determine the initial observation matrix based on the time-frequency characteristics of the lost frame speech data; and optimize the initial observation matrix according to the speech activity detection result, the adjacent valid frame data of the lost frame speech data, and the target sparse transformation basis to obtain the target observation matrix.

[0089] As an implementable method, in this embodiment, the speech reconstruction module 602 is also used to obtain the inverse transformation matrix corresponding to the target sparse transformation basis; perform preset speech reconstruction based on the lost frame sparse coefficients and the inverse transformation matrix to obtain the lost reconstructed speech frame; and generate initial speech recovery data based on the lost reconstructed speech frame and the framed speech data.

[0090] As an implementable embodiment, in this embodiment, the filtering module 603 is also used to input the initial speech recovery data into a preset filtering network to obtain adaptive filtering parameters; based on the adaptive filtering parameters and the dynamic filter, the initial speech recovery data is dynamically filtered to obtain target speech recovery data.

[0091] The voice data loss recovery device provided in this application utilizes the voice data loss recovery method described in the aforementioned embodiments to address the issue of voice quality degradation caused by message loss in VoIP communications. Compared to the prior art, the voice data loss recovery device provided in this application achieves the same beneficial effects as the voice data loss recovery method described in the aforementioned embodiments. Other technical features of the voice data loss recovery device are the same as those disclosed in the aforementioned embodiments and are not further detailed here.

[0092] The present application provides a voice data loss recovery processing device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the voice data loss recovery processing method of the above-mentioned embodiment 1.

[0093] Reference below Figure 7 , which shows a schematic diagram of the structure of a voice data loss recovery processing device suitable for implementing the embodiments of the present application. The voice data loss recovery processing device in the embodiments of the present application can include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The voice data loss recovery processing device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0094] like Figure 7As shown, the voice data loss recovery device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. RAM 1004 also stores various programs and data required for the operation of the voice data loss recovery device. Processing device 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the voice data loss recovery processing device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a voice data loss recovery processing device with various systems, it should be understood that implementation or presence of all the illustrated systems is not required. More or fewer systems may alternatively be implemented or present.

[0095] In particular, according to the embodiments disclosed herein, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed herein include a voice data loss recovery program product, which includes a voice data loss recovery program carried on a computer-readable medium, the voice data loss recovery program containing program code for executing the method illustrated in the flowchart. In such an embodiment, the voice data loss recovery program can be downloaded and installed from a network via a communication device, or installed from storage device 1003 or read-only memory 1002. When the voice data loss recovery program is executed by processing device 1001, the aforementioned functions defined in the method of the embodiments disclosed herein are performed.

[0096] The voice data loss recovery device provided in this application utilizes the voice data loss recovery method described in the aforementioned embodiment, resolving the technical issue of existing voice data loss recovery processes affecting voice transmission rates. Compared to the prior art, the beneficial effects of the voice data loss recovery device provided in this application are the same as those of the voice data loss recovery method described in the aforementioned embodiment. Other technical features of the voice data loss recovery device are the same as those disclosed in the aforementioned embodiment and are not further elaborated here.

[0097] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0098] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0099] The present application provides a storage medium having computer-readable program instructions (ie, a voice data loss recovery processing program) stored thereon, wherein the computer-readable program instructions are used to execute the voice data loss recovery processing method in the above-mentioned embodiment.

[0100] The storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination thereof. More specific examples of storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this embodiment, the storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system or device. The program code contained on the storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0101] The above storage medium may be included in the voice data loss recovery processing device; or may exist independently without being assembled into the voice data loss recovery processing device.

[0102] The storage medium carries one or more programs. When the one or more programs are executed by the voice data loss recovery processing device, the voice data loss recovery processing device performs voice data loss recovery processing.

[0103] The voice data loss recovery processing program code for performing the operations of the present application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, via the Internet using an Internet service provider).

[0104] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the systems, methods, and voice data loss recovery processing program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in an order different from that marked in the accompanying drawings. For example, two blocks shown in succession may actually be executed substantially in parallel, or they may sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or may be implemented using a combination of dedicated hardware and computer instructions.

[0105] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.

[0106] The readable storage medium provided in this application is a storage medium storing computer-readable program instructions (i.e., a voice data loss recovery program) for executing the aforementioned voice data loss recovery method. This storage medium can address the technical issue of existing voice data loss recovery processes affecting voice transmission rates. Compared to the prior art, the beneficial effects of the storage medium provided in this application are similar to those of the voice data loss recovery method provided in the aforementioned embodiments, and are not further elaborated here.

[0107] The above are only some embodiments of the present application and are not intended to limit the patent scope of the present application. All equivalent structural transformations made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A method for recovering voice data loss, characterized in that: The method comprises: Performing frame processing on the received voice message data to obtain framed voice data; When signal loss is detected, performing speech sparse reconstruction based on compressed sensing on the framed speech data to obtain initial speech recovery data; The initial speech restoration data is subjected to dynamic speech filtering through a preset filtering network to obtain target speech restoration data.

2. The voice data loss recovery processing method according to claim 1, wherein: When signal loss is detected, the step of performing speech sparse reconstruction based on compressed sensing on the framed speech data to obtain initial speech recovery data includes: When signal loss is detected, converting the detected lost frame speech data into a sparse speech signal; Obtaining a lost frame sparse coefficient corresponding to the sparse speech signal; The framed speech data is reconstructed using a preset speech configuration according to the lost frame sparse coefficient to obtain initial speech recovery data.

3. The voice data loss recovery processing method according to claim 2, wherein: The step of converting the detected lost frame voice data into a sparse voice signal comprises: Obtaining a target sparse transform basis corresponding to the detected lost frame speech data; The lost frame speech data is projected onto a sparse domain according to the local sparsity of the lost frame speech data and the target sparse transformation basis to obtain a sparse speech signal.

4. The voice data loss recovery method according to claim 3, wherein: The step of obtaining the lost frame sparse coefficient corresponding to the sparse speech signal includes: Constructing a target measurement matrix corresponding to the lost frame speech data based on adjacent valid frame data of the lost frame speech data and the target sparse transformation basis; Signal solution is performed based on the target observation matrix and the sparse speech signal through a preset optimization algorithm to obtain the lost frame sparse coefficients.

5. The voice data loss recovery processing method according to claim 4, wherein: The step of constructing a target observation matrix corresponding to the lost frame speech data based on adjacent valid frame data of the lost frame speech data and the target sparse transform basis includes: Obtaining a voice activity detection result of the lost frame voice data; Determining an initial observation matrix based on the time-frequency characteristics of the lost frame speech data; The initial measurement matrix is ​​optimized according to the voice activity detection result, the adjacent valid frame data of the lost frame voice data and the target sparse transformation basis to obtain a target measurement matrix.

6. The voice data loss recovery method according to claim 5, wherein: The step of reconstructing the framed speech data according to the lost frame sparse coefficient to obtain initial speech recovery data includes: Obtaining an inverse transformation matrix corresponding to the target sparse transformation basis; Reconstructing a preset voice based on the lost frame sparse coefficient and the inverse transformation matrix to obtain a lost reconstructed voice frame; Initial speech recovery data is generated based on the lost reconstructed speech frame and the framed speech data.

7. The voice data loss recovery method according to claim 1, wherein: The step of performing dynamic speech filtering on the initial speech recovery data through a preset filtering network to obtain target speech recovery data includes: Inputting the initial speech recovery data into a preset filter network to obtain adaptive filter parameters; Dynamic speech filtering is performed on the initial speech restoration data based on the adaptive filtering parameters and the dynamic filter to obtain target speech restoration data.

8. A voice data loss recovery processing device, characterized in that: The voice data loss recovery processing device comprises: The voice framing module is used to perform framing processing on the received voice message data to obtain framed voice data; A speech reconstruction module is used to perform speech sparse reconstruction based on compressed sensing on the framed speech data to obtain initial speech recovery data when signal loss is detected; The filtering module is used to perform dynamic speech filtering on the initial speech recovery data through a preset filtering network to obtain target speech recovery data.

9. A voice data loss recovery processing device, characterized in that: The device includes: a memory, a processor, and a voice data loss recovery processing program stored in the memory and executable on the processor, wherein the voice data loss recovery processing program is configured to implement the steps of the voice data loss recovery processing method according to any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium stores a voice data loss recovery processing program, which, when executed by a processor, implements the steps of the voice data loss recovery processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Speech coding method based on compressed sensing and sparse representation

    CN103778919A

  • Audio restoration method and system based on low consistency dictionary and sparse expression

    CN107039042A

  • Voice line spectrum frequency coding based on compression sensing and self-adaption quick reconstructing method

    CN109545234A

  • Anti-packet-loss compressed sensing base audio stream encoding and decoding method and system

    CN111404639A

  • Device for correcting voice signal loss of internet telephone

    KR1020020036592A