Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

144 results about "Speech coding" patented technology

Speech coding is an application of data compression of digital audio signals containing speech. Speech coding uses speech-specific parameter estimation using audio signal processing techniques to model the speech signal, combined with generic data compression algorithms to represent the resulting modeled parameters in a compact bitstream.

Speech translation model training method, speech translation method and device based on cross-modal attention, global memory and dynamic convolution

The invention discloses a cross-modal attention, global memory and dynamic convolution-based speech translation model training method and device, and a speech translation method and device, and relates to the technical field of speech processing and machine translation. A speech translation model is designed to comprise a speech encoder, a text embedding layer, a cross-modal attention adapter, a large language model decoder, a global memory network, a dynamic convolution decoder and an output layer. The cross-modal attention adapter projects audio features and performs multi-head cross attention fusion with text embedding; the global memory network updates and enhances historical memory on the basis of a gating mechanism and a Transform Encoder; and the dynamic convolution decoder performs multi-scale convolution extraction on the decoded hidden representation and fuses with the memory, so that the translation quality is improved. According to the method, deep fusion of voice and text, continuous memory with contextual coherence and high-quality translation generation can be realized, the end-to-end voice translation performance is remarkably improved, and the actual requirements of real-time and high-quality end-to-end voice translation in a complex scene are met.
Owner:BEIJING YUNSHANG TECH CO LTD

Signal encoding using latent feature prediction

Techniques and solutions are described for encoding and decoding signals, such as audio data. Disclosed innovations can find particular use in speech coding applications, such as for real time communications. Using a neural network, contextual coding can be used to encode latent features for a current frame using a prediction from reconstructed latent features of past frames as a context. An extractor learns a residual-like feature based on such prediction and latent features of the current frame obtained using an encoder. The residual-like feature is then quantized. At a decoder portion of a coding framework, the quantized feature is dequantized and then combined with a prediction from prior reconstructed latent features to provide reconstructed features of a current frame, which can then be processed by a decoder to provide a reconstructed signal.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Speech encoder training method and apparatus, device, medium, and program product

Disclosed are a speech encoder training method performed by a computer device. The method includes: masking a first sub-feature representation at a first feature position in a first text feature representation to obtain a first masked feature representation; performing feature prediction on a masked first feature position in the first masked feature representation based on a first speech feature representation to obtain a first predicted feature representation; and training a first speech encoder based on a difference between the first predicted feature representation and the first sub-feature representation to obtain a second speech encoder. The first speech encoder is trained by combining data in a speech modality with data in a text modality, and information included in the data in the text modality is adopted so that the first speech encoder can learn relatively high-level semantic representations of speech, thereby improving the prediction accuracy of representations.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Method, device and equipment for training voice coding model and readable medium

PendingCN121122255ASpeech recognitionMedicineSpeech code
The embodiment of the invention relates to a method and device for training a voice coding model, equipment and a readable medium. The method comprises: processing a speech feature representation of a speech sample using a first speech coding model to generate a set of discrete features; generating label information based on the set of discrete features, the label information comprising a set of labels, the set of labels indicating clustering centers corresponding to the corresponding discrete features; processing an intermediate feature representation generated based on the voice feature representation by using a second voice coding model to generate probability information corresponding to the tag information; the training loss is determined based on the label information, the probability information and the weight information, and the weight information is determined based on the distance from each discrete feature to the corresponding clustering center; and adjusting parameters of the second speech coding model based on the training loss. Therefore, the data volume required by training can be reduced, the stability of the unsupervised model training process can be improved, the training effect of the model is improved, and the model capability is further improved.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +1

Ultra-low bit rate voice coding and decoding system based on text semantic information fidelity

ActiveCN121214951ASemantic analysisBiological modelsIntelligibility (communication)Communications system
The invention provides an ultra-low bit rate voice coding and decoding system based on text semantic information fidelity, and relates to the technical field of voice coding and decoding. The system comprises the following steps: performing voice feature extraction and text feature extraction on original voice through a multi-modal text-voice combined encoder to obtain voice features and text features, embedding the text features into the voice features to obtain text-voice features of the original voice, and performing voice feature extraction and text feature extraction on the original voice through a multi-modal text-voice combined encoder; sending the text-speech features of the original speech to a receiving end in a semantic communication system; and the receiving end inputs the received text-speech features into a text semantic fidelity decoder based on an attention mechanism to obtain reconstructed speech features and reconstructed text features, decodes the reconstructed speech features layer by layer, and performs attention calculation on the reconstructed speech features and the reconstructed text features to obtain decoded speech. By means of the voice coding and decoding technology, good voice perception quality can still be kept under the ultra-low code rate, and voice intelligibility is guaranteed.
Owner:TSINGHUA UNIVERSITY

Improving speech recognition by a machine learning model

The present disclosure describes techniques for improving speech recognition using a machine learning model. The machine learning model comprises a speech encoder configured to generate acoustic representations based on input speech, an adapter configured to generate adapted representations based on the acoustic representations, and a decoder configured to generate text corresponding to the input speech. A matching loss is applied during training the machine learning model. The matching loss is configured to explicitly force acoustic representations generated by the adapter to align with text embeddings. The machine learning model is fine-tuned by employing parameter-efficient low-rank adaptation. The machine learning model is trained to perform automatic speech recognition with performance improvement and parameter efficiency.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +1

Speech recognition text scoring method and apparatus, electronic device, and storage medium

ActiveCN116403567BSpeech recognitionSpeech codeSpeech sound
The application provides a speech recognition text scoring method and device, electronic equipment and storage medium. The speech recognition text scoring method obtains standard audio features of target speech based on a speech coding process in a speech recognition process of the target speech from the perspective of speech recognition, thereby providing a reference standard for scoring of the recognized text. From the perspective of scoring of the recognized text, audio distribution features corresponding to the first recognized text are obtained based on the first recognized text of the target speech. The first recognized text is scored in combination with the audio distribution features and the standard audio features. The method considers the correlation between the speech recognition text and the audio during scoring of the recognized text, thereby achieving scoring of the speech recognition text.
Owner:IFLYTEK CO LTD

Speech recognition method and device, electronic equipment and storage medium

PendingCN121281502ASpeech recognitionSemantic representationSpeech code
The invention provides a speech recognition method and device, electronic equipment and a storage medium, and relates to the technical field of natural language processing, an adopted target automatic speech recognition model generates context semantic representation through a current text prefix sequence, and combines the context semantic representation with the current text prefix sequence and acoustic features to obtain a speech recognition result. A target text sequence is obtained through step-by-step prediction in an autoregression mode. Context semantic representation is introduced into the target automatic speech recognition model, the language switching moment can be accurately judged when the input speech signal is speech code conversion speech, speech recognition is carried out in time according to a new language during language switching, the recognition precision and robustness of a language switching boundary can be effectively improved, and the speech recognition efficiency is improved. And the speech code conversion speech recognition effect is improved. Moreover, context semantic representation is introduced during prediction, the challenge of ambiguity or ambiguity of acoustic signals can be overcome, the accuracy of the target text sequence is improved, and errors caused by untimely language model switching are reduced.
Owner:IFLYTEK CO LTD

Method and apparatus for encoding bone conduction speech, method and apparatus for decoding bone conduction speech, medium and device

The application discloses a bone conduction speech coding method, a coding and decoding method, a device, a medium and equipment, and belongs to the technical field of audio coding and decoding. The coding method comprises the following steps: updating a bandwidth detection parameter table in advance; updating a time domain noise shaping coding coefficient table in advance; in an LC3 audio coding process, the bone conduction speech is coded by using the updated bandwidth detection parameter table and the updated time domain noise shaping coding coefficient table, and a noise level estimation process of the bone conduction speech is performed on a bone conduction speech frequency band, so that the coding of the bone conduction speech is completed. The bandwidth detection parameter table and the time domain noise shaping coding coefficient table in the LC3 encoder are updated, so that when the LC3 encoder codes the bone conduction speech, the low-frequency part of the bone conduction speech can be recognized and corresponding coding processing can be performed, the coding of the bone conduction speech is realized, and the application range of the LC3 audio coding and decoding device is widened.
Owner:BEIJING BAIRUI INTERNET TECH CO LTD

Semantic perception real-time voice endpoint detection method and device

The semantic perception real-time voice endpoint detection method is applied to a voice interaction system, and comprises the following steps: receiving a voice stream of a user and a voice stream sent by the system from at least two audio channels; performing causal voice coding on the voice stream of the user and the voice stream sent by the system to obtain corresponding double-path voice representation; a causal two-way modeling architecture is constructed, and the causal two-way modeling architecture is realized by adopting an autoregressive prediction network and is used for receiving two-way voice representation and modeling a speech round conversion dynamic state between a user and a system through an embedded causal constraint mechanism and a cross-channel interaction mechanism; based on the output of the autoregressive prediction network, jointly predicting the category of a speech round conversion event and the remaining time from the current moment to the end of the speech round; and a smoothing processing mode is used to correct the predicted speech wheel conversion event type, and a final endpoint detection result is determined. According to the method, the endpoint detection stability in a complex dialogue scene can be improved.
Owner:INST OF ACOUSTICS CHINESE ACAD OF SCI

Joint speech text training for hybrid transducer and attention-based encoder-decoder modeling

A method includes generating speech embeddings corresponding to a received speech input using a speech encoder. The method also includes generating multi-modal embeddings based on at least one of the speech embeddings or a corresponding text embedding using a shared encoder. The method further includes generating conditioned multi-modal embeddings reflected by previous text predictions based on the multi-modal embeddings using a predictor. In addition, the method includes generating a text prediction corresponding to the received speech input based on the multi-modal embeddings and the conditioned multi-modal embeddings using a joint network.
Owner:SAMSUNG ELECTRONICS CO LTD

Real-time speech recognition method, related device and computer program product

PendingCN121214915ASpeech recognitionCode moduleSpeech code
The invention discloses a real-time voice recognition method, related equipment and a computer program product, and the method carries out the audio coding of real-time voice stream data to be recognized, and obtains an audio coding feature as the basis of the subsequent voice recognition. The real-time voice coding module carries out coding processing on the audio coding features, and an acoustic vector dt of the token to be decoded at the current decoding moment can be obtained. The real-time speech coding module adopted by the invention is a coding module in a real-time speech recognition model for token-by-token prediction, so that the acoustic boundary of the current token to be decoded can be determined, and the acoustic vector dt is obtained. According to the method, the acoustic vector dt is further mapped to a large model input space, it is guaranteed that the vector dimension meets the input dimension requirement of the large model, the mapped acoustic vector dt is sent to the large model for forward reasoning calculation, and the real-time speech recognition effect can be improved by means of the powerful context modeling capacity and rich text field knowledge of the large model.
Owner:ANHUI IFLYTEK UNIVERSAL LANGUAGE TECH CO LTD

Voice data processing method and device, storage medium and electronic equipment

The invention discloses a voice data processing method and device, a storage medium and electronic equipment, and the method comprises the steps: collecting to-be-processed user voice data, carrying out the task sharing coding processing of a voice recognition task and a named entity recognition task through a voice data processing model based on the user voice data, and obtaining a voice coding feature; and performing voice recognition processing based on the voice coding features to obtain a target voice recognition result, and performing named entity recognition based on the voice coding features to obtain a target entity recognition result. According to the method, one-time coding can serve two tasks of voice recognition and named entity recognition at the same time, consumption of computing resources is reduced, data processing efficiency is improved, and the method is particularly suitable for a scene of real-time processing of a large amount of voice data.
Owner:MIDEA GRP (SHANGHAI) CO LTD +1

System and method for generating a facial image from a voice sample using a stylegan

System and method for reconstructing a facial image of a speaker from a voice sample of the speaker may include providing the voice sample of the speaker to a trained voice encoder to generate a voice embedding of the speaker, wherein the voice encoder is trained to provide a voice embedding that matches an image embedding of the facial image of the speaker; providing the voice embedding of the speaker to a trained mapping network to generate an intermediate latent vector, wherein the mapping network is trained to generate an intermediate latent vector for a StyleGAN from the voice embedding; and providing the intermediate latent vector to the StyleGAN to generate the facial image of a speaker.
Owner:CORSOUND AI LTD

Audio-visual joint speech enhancement method and apparatus in multi-person environment, device and medium

PendingCN122337229ASpeech soundAudio frequency
This application relates to a method, apparatus, device, and medium for audiovisual joint speech enhancement in a multi-person environment. The method includes: a normalization unit processes facial video and mixed audio to obtain a facial image sequence and a synchronized two-dimensional encoded audio sequence; extracting the lip region of each frame to obtain a lip image sequence; processing this sequence through a network to obtain lip-reading content features; aligning these features with the audio frame; inputting the aligned audio frame to a speech encoder to obtain fixed-length speech encoding features; outputting these features to a speaker feature decoupling network to decouple audio content features from target speaker features; aligning and fusing the lip-reading and audio content features through a cross-modal network to obtain cross-modal fused features; fusing these features with the target speaker features through an attention mechanism to obtain target speaker enhanced features; and finally, using these features, a feature enhancement network enhances the mixed audio and outputs the target audio. This method, by using audiovisual multimodal fusion and speaker feature decoupling, can effectively improve the clarity of target speech in a multi-person environment and reduce background interference.
Owner:NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI

Dynamic VoLTE Allocation (DVA)

A method, computer readable media and a system for determining a Voice over LTE (VOLTE) codec is presented. In one embodiment a method includes discovering a packet header size used in a session using a Silence Insertion Descriptor (SID); inspecting downlink packet sizes coming into a receive buffer; determining when three consecutive packet sizes are greater than zero, then setting a first factor, a second factor, and a third factor equal to zero; calculating a header size; and determining a voice encoding codec by subtracting a header overhead value from the packet size.
Owner:PARALLEL WIRELESS INC

Speech synthesis method, device, equipment and readable medium

The embodiment of the invention provides a voice synthesis method and device, equipment, a storage medium and a program product. The method includes obtaining a speech coding representation corresponding to an initial speech of a first speaker, the speech coding representation indicating speech features unrelated to the speaker, the initial speech including speech content information. A pitch predictive encoded representation related to the initial speech is generated using a predictor model based on the speech encoded representation and a reference speech encoded representation, the reference speech encoded representation indicating a speech encoded representation corresponding to the second speaker, the pitch predictive encoded representation indicating pitch information of the second speaker. And generating a target speech corresponding to a second speaker by using a speech decoder model based on the pitch prediction coded representation and the reference speech coded representation, a speech signal of the second speaker including speech content information. In this way, the quality of the generated target voice is improved.
Owner:JINGDONG CITY BEIJING DIGITS TECH CO LTD +1

Speech synthesis methods, devices, equipment, and media based on gating attention mechanisms

This invention relates to the fields of artificial intelligence and fintech, and discloses a speech synthesis method based on a gated attention mechanism. The method involves acquiring text and speech data. The text data is converted into a sequence of text symbols using a text encoder, while the speech data has its speech features extracted using a speech encoder trained on a self-supervised learning model and quantized into a discrete sequence of speech symbols. The text and speech symbol sequences are then organized into text and speech sequences, respectively. A preliminary alignment process establishes the correspondence between text and speech symbols, and the gated attention mechanism dynamically adjusts the matching degree between them. Finally, a decoder generates the final speech signal. This invention effectively improves the speech feature extraction capability through a speech encoder trained on a self-supervised learning model, especially in scenarios lacking a large amount of labeled data, where it can still learn effective feature representations from unlabeled data.
Owner:PING AN TECH (SHENZHEN) CO LTD

Voice-based image driving method and image-driven data processing method

The embodiment of the specification provides a voice-based image driving method and an image-driven data processing method, wherein the voice-based image driving method comprises the following steps: obtaining a reference voice and a reference facial image of a virtual object, performing voice coding on the reference voice to obtain target voice features, performing image coding on the reference facial image to obtain first image features of a first region and second image features of a second region, performing feature transformation on the first image features based on facial prior features and the target voice features to determine first target image features, wherein the facial prior features comprise facial texture features, and generating a target image after driving according to the first target image features and the second image features. The first image features are transformed based on the target voice features and the facial prior features, and a target image with high fidelity and high definition is obtained through decoding, thereby improving the user experience.
Owner:ALIBABA DAMO (HANGZHOU) TECH CO LTD

Voice interaction method and device and electronic equipment

The invention provides a voice interaction method and device and electronic equipment, and relates to the technical field of voice processing, and the method comprises the steps: obtaining voice information inputted by a user, and obtaining a historical intention text of the user; inputting the voice information into a voice encoder of a spoken language understanding model to obtain acoustic encoding features output by the voice encoder; inputting the historical intention text into a text encoder of the spoken language understanding model to obtain text encoding features output by the text encoder; and inputting the acoustic coding features and the text coding features into an intention recognition module of the spoken language understanding model to obtain an intention recognition result output by the intention recognition module for voice interaction. The accuracy of the intention recognition result can be improved, and the real intention of the user can be accurately obtained.
Owner:HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD

Lightweight semantic preserving coding and decoding and voice signal reconstruction method

PendingCN121905195ASpeech analysisNeural learning methodsNerve networkSpeech reconstruction
The invention discloses a lightweight semantic preserving coding and decoding and voice signal reconstruction method, which comprises the following steps of: acquiring an analog audio signal from an environment, inputting an original audio signal and an audio signal subjected to high-frequency filtering to two ends of a comparator, outputting an ultralow-bit-rate binary bit stream and transmitting the ultralow-bit-rate binary bit stream to a receiving end, and after the receiving end receives the signal, outputting the ultralow-bit-rate binary bit stream to the voice signal reconstruction end. The method comprises the following steps: inputting a voice signal to an end-to-end U-Net convolutional neural network of a bottleneck layer integrated LSTM module, outputting a reconstructed voice signal, training the network by using voice coding sample reconstruction loss and multi-scale short-time Fourier transform, and performing voice signal reconstruction on an ultra-low bit rate binary bit stream of a to-be-reconstructed audio after training. The technical contradiction that in resource-limited remote audio collection and distributed sensing tasks, node hardware resources are extremely limited, and ultra-low bit rate signal compression distortion causes that node microminiaturization deployment and high voice reconstruction quality cannot be considered at the same time is solved.
Owner:ZHEJIANG UNIV

Automatic speech recognition methods, devices, and media based on natural language processing and speech coding.

This application relates to an automatic speech code recognition method, device, and medium based on natural language processing, belonging to the field of voice communication technology. The automatic speech code recognition method based on natural language processing includes the following steps: S1, selecting a speech code; S2, restoring the speech information to a standard speech code speech data stream according to the selected speech code format; S3, obtaining the language matching information of the speech data stream; if the speech data stream does not meet the requirements for all languages ​​under the current speech code format, repeating steps S1-S3; if the speech data stream meets the requirements for any language under the current speech code format, proceeding to the next step; S4, obtaining the current speech code format, and outputting the speech information as a speech data stream in the current speech code format. This application can quickly identify speech codes and ensure smooth communication during communication failures.
Owner:WUHAN HUABO COMM CO LTD

Multimodal abnormal emotion recognition method based on multitask, mixed data enhancement and contrast feature decomposition

The invention discloses a multi-modal abnormal emotion recognition method based on multitask, mixed data enhancement and contrast feature decomposition, and the method comprises the steps: transcribing voice data collected in psychological health census into text data, carrying out the feature extraction of the voice data and the text data, obtaining a word embedding sequence and frame-level acoustic features, and forming a sample; performing data enhancement on the word embedding sequence and the frame-level acoustic features to generate a diversified training sample; respectively constructing a text coding model and a voice coding model, and respectively training the corresponding coding models through word embedding sequences and frame-level acoustic features in the diversified training samples; carrying out dimension expansion on the obtained word embedding sequence and the frame-level acoustic features in a copying mode, and forming an abnormal emotion detection training sample by the expanded word embedding sequence and the frame-level acoustic features; and establishing a multi-modal abnormal emotion recognition model, and carrying out migration training on the multi-modal abnormal emotion recognition model through the abnormal emotion detection training sample.
Owner:NORTHEASTERN UNIV CHINA

A decoupled voice self-supervised pre-training method

The application discloses a kind of decoupling voice self-supervision pre-training methods, including pre-training and fine-tuning two stages.Construct with convolution, Transformer, pitch variation processor and speaker information processor as the core self-supervision pre-training model.Input voice, and convolution module is encoded as frame-level feature;Pitch variation processor extracts pitch variation representation, and is excluded from main branch, and is replaced with masking vector and input transformer encoder.In the middle layer of encoder, speaker processor module is added to extract speaker representation, and is excluded from main branch representation.Continue to encode processing, finally map to target speech representation dimension.After first round pre-training, extract middle layer representation, train second K-Means model to generate new pseudo-label target, and carry out second round pre-training.Using weighted summation mechanism obtains task-specific representation, suitable for various downstream tasks.
Owner:TIANJIN UNIV

A neural network-based lightweight streaming speech coding system and method

The application provides a neural network-based lightweight streaming speech coding system and method, which comprises an encoding compression end and a decoding reconstruction end; the encoding compression end comprises a speech encoder and a quantizer; the speech encoder comprises a time-frequency conversion module, a feature extraction module, an encoding end channel conversion module and a long-range time domain correlation extraction module; the feature extraction module comprises a time-frequency feature extraction submodule and a time-frequency size downsampling submodule; the quantizer comprises two or more layers of quantization modules, each layer of quantization module comprising an encoding end first domain conversion module, an encoding end vector quantization module, an encoding end inverse vector quantization module and an encoding end second domain conversion module; and the decoding reconstruction end comprises an inverse quantizer and a speech decoder. The application can effectively extract multi-scale feature information in a speech signal, so that high-quality speech compression coding and reconstruction can be realized at an extremely low code rate, and the parameter quantity is small, the complexity is low, and streaming coding and decoding can be realized.
Owner:NANJING UNIV

Voice conversion method, training method, device, equipment and medium

ActiveCN115565520BSpeech synthesisSpeech reconstructionVoice transformation
The voice conversion method, the training method, the device, the equipment and the medium provided by the embodiments of the present application obtain the source linear spectrum of the source speaker according to the source voice data of the source speaker; input the source linear spectrum into a pre-trained voice coding model to output corresponding spectral feature prediction data; input the spectral feature prediction data and target speaker feature data of a target speaker into a voice reconstruction model to output corresponding target voice data; in the above manner, the content information of the source voice data and the speaker feature are decoupled, and in the training stage and the voice conversion stage of the voice coding model, the input and the output of the voice coding model respectively only contain content information, the voice coding model is used to reconstruct voice data together with the content information and the speaker feature through the voice reconstruction model, which is beneficial to improve the training speed and the training effect of the voice coding model, and further improves the voice conversion effect.
Owner:PING AN TECH (SHENZHEN) CO LTD

Using machine learning speech synthesizer using synthetic analytic speech coding

PendingCN122374817ASpeech codeSpeech synthesis
An apparatus includes a memory configured to store data associated with a machine learning (ML) based speech synthesis model. The apparatus also includes a speech encoder including the ML based speech synthesis model. The speech encoder is configured to perform a synthetic analysis operation of an input speech signal including generating a synthetic version of the input speech signal by the ML based speech synthesis model.
Owner:QUALCOMM INC

Adaptive speech codec adjustment method, apparatus, device and medium

The application discloses an adaptive speech coding adjustment method and device, equipment and medium, and the method is applied to a talking stage of a voice communication system. The application detects the voice quality under the current channel condition by using a real-time intelligent voice measurement algorithm, and adaptively adjusts the speech coding type based on the voice quality detection result, so that the voice quality of the conversation is improved, and the user experience is improved.
Owner:NO 30 INST OF CHINA ELECTRONIC TECH GRP CORP

Voice interaction method, device and electronic equipment

The application provides a speech interaction method and device and electronic equipment, and relates to the technical field of speech processing. The method comprises the following steps: acquiring speech information input by a user, and acquiring historical intention text of the user; inputting the speech information into a speech encoder of a spoken language understanding model to obtain acoustic coding features output by the speech encoder; inputting the historical intention text into a text encoder of the spoken language understanding model to obtain text coding features output by the text encoder; inputting the acoustic coding features and the text coding features into an intention recognition module of the spoken language understanding model to obtain an intention recognition result output by the intention recognition module, so as to be used for speech interaction. The application can improve the accuracy of the intention recognition result and accurately acquire the real intention of the user.
Owner:HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD

Lightweight streaming speech coding system and method based on neural network

The invention provides a lightweight streaming speech coding system and method based on a neural network. The system comprises a coding compression end and a decoding reconstruction end, the coding compression end comprises a voice coder and a quantizer; the voice encoder comprises a time-frequency transformation module, a feature extraction module, an encoding end channel transformation module and a long-range time domain correlation extraction module. The feature extraction module comprises a time-frequency feature extraction sub-module and a time-frequency size sampling sub-module; the quantizer comprises more than two layers of quantization modules, and each layer of quantization module comprises a coding end first domain transformation module, a coding end vector quantization module, a coding end reverse quantity quantization module and a coding end second domain transformation module; and the decoding reconstruction end comprises an inverse quantizer and a voice decoder. According to the method, the multi-scale feature information in the voice signal can be effectively extracted, so that high-quality voice compression coding and reconstruction can be realized at an extremely low code rate, the parameter quantity is small, the complexity is low, and streaming coding and decoding can be realized.
Owner:NANJING UNIV