Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

22 results about "Speech code" patented technology

Speech synthesis method, device, equipment and readable medium

The embodiment of the invention provides a voice synthesis method and device, equipment, a storage medium and a program product. The method includes obtaining a speech coding representation corresponding to an initial speech of a first speaker, the speech coding representation indicating speech features unrelated to the speaker, the initial speech including speech content information. A pitch predictive encoded representation related to the initial speech is generated using a predictor model based on the speech encoded representation and a reference speech encoded representation, the reference speech encoded representation indicating a speech encoded representation corresponding to the second speaker, the pitch predictive encoded representation indicating pitch information of the second speaker. And generating a target speech corresponding to a second speaker by using a speech decoder model based on the pitch prediction coded representation and the reference speech coded representation, a speech signal of the second speaker including speech content information. In this way, the quality of the generated target voice is improved.
Owner:JINGDONG CITY BEIJING DIGITS TECH CO LTD +1

Lightweight semantic preserving coding and decoding and voice signal reconstruction method

PendingCN121905195ASpeech analysisNeural learning methodsNerve networkSpeech reconstruction
The invention discloses a lightweight semantic preserving coding and decoding and voice signal reconstruction method, which comprises the following steps of: acquiring an analog audio signal from an environment, inputting an original audio signal and an audio signal subjected to high-frequency filtering to two ends of a comparator, outputting an ultralow-bit-rate binary bit stream and transmitting the ultralow-bit-rate binary bit stream to a receiving end, and after the receiving end receives the signal, outputting the ultralow-bit-rate binary bit stream to the voice signal reconstruction end. The method comprises the following steps: inputting a voice signal to an end-to-end U-Net convolutional neural network of a bottleneck layer integrated LSTM module, outputting a reconstructed voice signal, training the network by using voice coding sample reconstruction loss and multi-scale short-time Fourier transform, and performing voice signal reconstruction on an ultra-low bit rate binary bit stream of a to-be-reconstructed audio after training. The technical contradiction that in resource-limited remote audio collection and distributed sensing tasks, node hardware resources are extremely limited, and ultra-low bit rate signal compression distortion causes that node microminiaturization deployment and high voice reconstruction quality cannot be considered at the same time is solved.
Owner:ZHEJIANG UNIV

Using machine learning speech synthesizer using synthetic analytic speech coding

PendingCN122374817ASpeech codeSpeech synthesis
An apparatus includes a memory configured to store data associated with a machine learning (ML) based speech synthesis model. The apparatus also includes a speech encoder including the ML based speech synthesis model. The speech encoder is configured to perform a synthetic analysis operation of an input speech signal including generating a synthetic version of the input speech signal by the ML based speech synthesis model.
Owner:QUALCOMM INC

Adaptive speech codec adjustment method, apparatus, device and medium

The application discloses an adaptive speech coding adjustment method and device, equipment and medium, and the method is applied to a talking stage of a voice communication system. The application detects the voice quality under the current channel condition by using a real-time intelligent voice measurement algorithm, and adaptively adjusts the speech coding type based on the voice quality detection result, so that the voice quality of the conversation is improved, and the user experience is improved.
Owner:NO 30 INST OF CHINA ELECTRONIC TECH GRP CORP

Voice coding method and device, electronic equipment and computer storage medium

The invention provides a voice coding method, a voice coding device, electronic equipment and a computer storage medium. Fusing the encoded semantic features output by the pre-training feature encoder, the video features output by the video encoder and the encoding features output by the existing voice encoder to generate a voice encoding result; and the sound quality of the coded voice is effectively improved.
Owner:UNIV OF SCI & TECH OF CHINA

Training of a voice conversion model and voice conversion method and device and related equipment

ActiveCN114882896BSpeech synthesisSpeech codeAcoustics
The application discloses a speech conversion model training method applied to the field of artificial intelligence. The speech conversion model provided by the application comprises an encoder, an instantiation normalization layer, a voiceprint extractor and a decoder. The method provided by the application comprises the following steps: inputting a speech sample of an original speaker into the encoder to obtain first speech coding data; the instantiation normalization layer removes the speech attribute of the original speaker in the first speech coding data to obtain a first speech hidden vector; a first voiceprint vector of a target speaker is obtained; the decoder synthesizes the first speech hidden vector and the first voiceprint vector to obtain reconstructed speech data; a first loss of the reconstructed speech data and the speech sample of the original speaker is calculated; whether the first loss reaches a maximum is judged; if not, the parameters of the instantiation normalization layer are optimized, and the foregoing steps are circularly performed until the first loss reaches the maximum, so that a trained speech conversion model is obtained.
Owner:PING AN TECH (SHENZHEN) CO LTD

Deep learning-based technology consultation intention accurate identification and response method

The invention relates to the cross technical field of deep learning and technology consultation, in particular to a technology consultation intention accurate identification and response method based on deep learning. Comprising five core processes of multi-modal consultation access and preprocessing, multi-dimensional technical intention deep recognition, dynamic knowledge graph retrieval and adaptation, personalized response generation and optimization, and multi-round interactive feedback and iteration, and all the processes are cooperatively linked with a structured knowledge system through a deep learning model, so that technical consultation whole-process intelligent processing is realized. Through multi-modal information fusion, a technical field pre-training model and implicit intention mining, the technical consultation intention recognition accuracy is greater than or equal to 95%, which is significantly superior to that of a traditional method, response deviation caused by intention misjudgment is effectively avoided, multi-modal input of texts, pictures, voices, codes and the like is supported, technical consultation information is completely received, and the technical consultation intention recognition accuracy is improved. The problem of information missing caused by single text input is solved, and diversified expression scenes of technology consultation are adapted.
Owner:SHANGHAI QUWANG INFORMATION TECHNOLOGY CO LTD

Method and apparatus for adjusting speech coding, electronic device, and storage medium

The application provides a speech coding adjustment method and device in dynamic spectrum sharing, electronic equipment and storage medium, and relates to the technical field of wireless communication. The method comprises the following steps: receiving a notification of whether a next first preset time length occupies a shared frequency band sent by a long term evolution (LTE) network at a current time; in response to the notification including that the LTE network needs to occupy the shared frequency band in the next first preset time length, correcting channel quality information, and adjusting a speech coding rate based on the corrected channel quality information. At least the problem that the LTE system has high priority and occupies the shared spectrum for a long time when the traffic volume is large, resulting in poor voice service quality of the UMTS system, is solved. The application is suitable for spectrum sharing optimization, voice service optimization and the like.
Owner:CHINA UNITED NETWORK COMM GRP CO LTD

Deep learning-based adaptive speech recognition system

PendingCN122337201AFeature extractionSpeech code
This invention discloses a deep learning-based adaptive speech recognition system, relating to the field of speech recognition technology. The system processes raw speech signal data acquired by a speech acquisition module and dynamically adapts it to user identification information to obtain speech coding feature data. Furthermore, it extracts features from environmental metadata to obtain environmental embedding feature data. Based on the environmental embedding feature data and the speech coding feature data, feature recognition processing is performed to calculate recognition feature coefficients. The speech recognition module compares these coefficients with preset recognition feature thresholds and determines the speech quality based on the comparison results. This enables recognition under dynamically changing user and environmental conditions, improving the accuracy of speech recognition.
Owner:IANGSU COLLEGE OF ENG & TECH

Ultra low bit rate speech coding and decoding system based on text semantic information fidelity

ActiveCN121214951BSemantic analysisBiological modelsIntelligibility (communication)Communications system
The application provides a kind of based on text semantic information fidelity super low code rate speech coding system, it is related to speech coding technical field.The system includes: through multimodal text-speech joint encoder, speech feature extraction and text feature extraction are carried out to original speech, speech feature and text feature are obtained, and text feature is embedded into speech feature, text-speech feature of original speech is obtained, and text-speech feature of original speech is sent to receiving end in semantic communication system;The text-speech feature received by receiving end is input into text semantic fidelity decoder based on attention mechanism, to obtain reconstructed speech feature and reconstructed text feature, and the reconstructed speech feature is decoded layer by layer and attention calculation is carried out with the reconstructed text feature, to obtain decoded speech.Through the speech coding technology of the application, good speech perceptual quality can still be maintained under super low code rate, and the speech intelligibility is guaranteed.
Owner:TSINGHUA UNIVERSITY

Audio decoding method and device, electronic equipment and storage medium

The invention provides an audio decoding method and device, electronic equipment and a storage medium, and relates to the field of audio decoding, and the method comprises the steps: receiving L2HC bit stream data, extracting a packet header segment in the L2HC bit stream data, and analyzing a hybrid coding mark in the packet header segment; extracting a side information segment and an effective load segment in the L2HC bit stream data, extracting a frequency division control mark and a voice coding control field from the side information segment when it is determined that the mixed coding mark represents voice coding, and extracting a voice data segment of a specified frequency band from the effective load segment according to the frequency division control mark; decoding the voice data segment according to the voice coding control field to obtain a voice audio; according to the technical scheme, the voice decoding exclusive field can be added in the side information segment of the L2HC bit stream data, the voice data segment is borne by the effective load segment of the L2HC bit stream data, and the voice data segment is independently decoded, so that high-quality voice audio decoding can be realized under the L2HC framework.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Voice coding data processing method, device, equipment and chip

PendingCN121884833ASpeech analysisSpeech codeSpeech sound
The invention relates to a voice coding data processing method and device, equipment and a chip. The method comprises the following steps: performing pre-decoding processing on voice coding data, and determining a pre-decoded disguise frame; adjusting frame header information of a corresponding camouflage frame in the voice coding data into first frame header information, wherein the first frame header information is used for indicating that the corresponding frame is abnormal; wherein the adjusted voice coding data is used for being input into a decoder, so that voice decoding is carried out on the adjusted voice coding data. By adopting the method, noise of the decoded voice data can be avoided, and the use experience is improved.
Owner:UNISOC CHONGQING TECH CO LTD

Voice processing model training method, voice processing method, device and apparatus

ActiveCN119763548BSpeech recognitionSpeech synthesisSpeech codeAcoustics
The present disclosure provides a speech processing model training method, a speech processing method, an apparatus and a device, and belongs to the technical field of computers. The method comprises: performing speech coding on a sample speech signal to obtain semantic embedding representation and acoustic embedding representation of the sample speech signal; performing phoneme extraction and phoneme coding on a reference speech text of the sample speech signal to obtain phoneme embedding representation of the reference speech text; training a speech processing model based on the semantic embedding representation, the acoustic embedding representation and the phoneme embedding representation, the speech processing model being used for real-time speech synthesis on an input speech. The method introduces semantic information and acoustic information in the model training process, so that the model can learn cleaner semantic information, more language information can be retained when processing the speech, and the naturalness of the synthesized speech is improved.
Owner:BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

Weakly supervised target speaker extraction method and system based on pseudo-label signal generation

ActiveCN121483262BSpeech analysisFeature extractionSpeech code
A weakly supervised target speaker extraction method and system based on pseudo-label signal generation are disclosed, including obtaining a to-be-processed far-field multi-channel mixed audio and a corresponding target speaker timestamp; the amplitude spectrum of the far-field multi-channel mixed audio is spliced along the channel dimension and mapped into speech coding hidden features; according to the target speaker timestamp, a speech segment in which the target speaker is active is cut out from a reference channel of the far-field multi-channel mixed audio; a speaker embedding vector of the speech segment is calculated and dimensionally expanded to obtain target speaker embedding features; the speech coding hidden features and the target speaker embedding features are fused and input into a target speaker extraction model for feature extraction to obtain target speaker speech features; the target speaker speech features are mapped into an amplitude spectrum mask, the mask is applied to the amplitude spectrum of the reference channel, and inverse transformation is performed in combination with the phase of the reference channel to obtain clean target speaker speech.
Owner:XIAMEN UNIV

Voice coding and decoding method and system

The invention discloses a voice coding and decoding method, which comprises a receiving end and a transmitting end, and is characterized in that the transmitting end acquires voice data which is in a pulse code modulation (PCM) format, has a single channel, 16 bits and a sampling rate of 8kHz, and acquires 320 bytes per 20ms; compressing each 320 bytes of voice data into 5 bytes of compressed data; multiple groups of compressed data are assembled into data frames, every 18 groups of compressed data form one data frame, and protocol fields are added to the heads of the data frames; performing encryption and scrambling processing on the data frame; the processed data frame is sent to a satellite communication channel through a digital signal processor (DSP); a receiving end receives data frames from a satellite communication channel, and one frame is received every 60ms; splicing the plurality of data frames into a complete voice data frame; performing descrambling and decryption processing on the voice data frame; decompressing the decrypted data into 320-byte voice data by taking every 5 bytes as a group; playing the decompressed voice data; according to the invention, the voice quality is greatly improved.
Owner:NANJING 6902 TECH

A voice-driven three-dimensional virtual figure complex emotional facial motion generation method

The application discloses a speech-driven three-dimensional virtual image complex emotion facial action generation method, relates to the technical field of virtual image animation, and comprises the following steps: receiving a speech segment and an emotion guide image as input, extracting speech features and emotion coding features through a speech coding module and an emotion coding module respectively; randomly sampling a fixed-length noise sequence, splicing the fixed-length noise sequence with time step embedded and processed default shape parameters to form an initial noise input; based on a conditional diffusion model, sequentially guiding a diffusion process with the speech features and the emotion coding features as conditions, generating a three-dimensional facial expression animation with synchronized lip shapes and speech and consistent emotions and input pictures; the method can effectively handle complex emotion scenes, effectively express mixed emotions and hidden emotion characteristics, and quantitative evaluation results show that the method has excellent performance.
Owner:BEIJING INST OF TECH

Speech recognition methods, speech recognition systems, computer equipment and storage media

ActiveCN116959424Bavoid collectingaccurate identificationSpeech recognitionSpeech codeSpeech classification
This application provides a speech recognition method, a speech recognition system, a computer device, and a storage medium, belonging to the field of financial technology. The method includes: inputting target speech with a preset emotion category into a pre-trained multi-task speech recognition model; encoding the target speech using a first speech coding sub-model to obtain initial speech features; performing speech attention processing on the initial speech features using a first attention sub-model to obtain first target attention features; encoding the initial speech features using a second speech coding sub-model to obtain hidden speech features; performing hidden attention processing on the first target attention features and the hidden speech features using a second attention sub-model to obtain second target attention features; and performing speech classification on the second target attention features using a multi-task classification sub-model to obtain a target speech label. This application embodiment can improve the recognition accuracy of multi-task speech recognition.
Owner:PING AN TECH (SHENZHEN) CO LTD

Voice processing method and device, equipment and storage medium

The embodiment of the invention provides a voice processing method and device, equipment and a storage medium. The method comprises the following steps: acquiring a voice feature sequence corresponding to a voice sample, wherein voice features in the voice feature sequence correspond to voice frames in the voice sample; and for a target voice feature in the voice feature sequence, generating one or more post-order voice marks corresponding to the one or more post-order voice features based on the target voice feature and the one or more pre-order voice features by using a voice coding model. A speech encoding model is trained based on the one or more post-order speech markers. The training mode can enable the speech coding model to learn high-quality speech representation.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +1

Speech recognition method and device, storage medium and program product

PendingCN121747541ASpeech recognitionSpeech codeSpeech sound
The invention discloses a voice recognition method and device, a storage medium and a program product, and relates to the technical field of artificial intelligence, and the method comprises the steps: coding a to-be-recognized voice, and obtaining a first voice coding feature of the to-be-recognized voice; encoding the text of each hot word in the target hot word set to obtain a text encoding feature of each hot word; obtaining an audio coding feature of each hot word obtained by coding the audio of each hot word in the target hot word set; fusing the text coding feature and the audio coding feature of the same hot word to obtain a target coding feature of each hot word; and performing hot word enhanced speech recognition based on the target coding feature of each hot word and the first speech coding feature to obtain a speech recognition result of the to-be-recognized speech. According to the hot word enhanced speech recognition method provided by the invention, the text features and the audio features of the hot words are fused, so that the recognition accuracy of the hot words is improved.
Owner:HEFEI IFLY DIGITAL TECH CO LTD

Weak supervision target speaker extraction method and system based on pseudo label signal generation

ActiveCN121483262ASpeech analysisFeature extractionSpeech code
The invention discloses a weak supervision target speaker extraction method and system based on pseudo tag signal generation. The method comprises the following steps: acquiring a far-field multi-channel mixed audio to be processed and a corresponding target speaker timestamp; splicing the amplitude spectrums of the far-field multi-channel mixed audio along channel dimensions, and mapping the amplitude spectrums into voice coding hidden features; according to the timestamp of the target speaker, segmenting an active voice segment of the target speaker from a reference channel of the far-field multi-channel mixed audio; calculating a speaker embedding vector of the voice segment, and performing dimension expansion to obtain a target speaker embedding feature; fusing the voice coding hidden features with the target speaker embedded features, and inputting the fused features into a target speaker extraction model for feature extraction to obtain target speaker voice features; and mapping the speech features of the target speaker into a magnitude spectrum mask, acting the mask on the magnitude spectrum of the reference channel, and performing inverse transformation in combination with the phase of the reference channel to obtain the clean speech of the target speaker.
Owner:XIAMEN UNIV

A system and method for processing audio data

PendingCN122313998AData processing systemNoise
This invention relates to the field of speech coding technology, specifically to an audio data processing system and method. The system includes an audio signal framing module, a spectrum threshold routing module, a threshold balance calibration module, a peak link recognition module, and a curve audio correction module. In this invention, after the audio signal is framed and transformed, the amplitude is compared with adjacent differences to trigger dynamic adjustments, enabling fine correction of spectral abrupt change regions. Local threshold adaptive reconstruction maintains frequency band energy coordination, half-value iterative compensation suppresses single-point deviations and stabilizes the energy structure, and inverse-range weighted calculation of peak positions makes frequency band transitions smoother, reducing energy abrupt changes. Offset correction adjusts the amplitude distribution based on the predicted curve difference, optimizing spectral continuity and residual smoothness, and overall improving transient fidelity and low-noise balance, allowing compressed audio to remain clear and natural in complex scenes.
Owner:NANJING CODE NOTE NETWORK TECH CO LTD

Extending multilingual speech synthesis with zero supervision of discovered data

PendingCN121753095ABiological modelsSpeech recognitionSpeech codeAcoustics
A method (600) includes receiving training data (301) including a plurality of sets of training utterances (310) each associated with a respective language. Each training utterance includes a corresponding reference speech representation (304) paired with a corresponding input text sequence (302). For each training utterance, the method includes generating a corresponding encoded text representation for a corresponding input text sequence (312, 313), generating a corresponding speech code for a corresponding reference speech representation (314), generating a shared encoder output (332, 334), and outputting the shared encoder output (332, 334). And determine a text-to-speech (TTS) loss based on the corresponding encoded text representation, the corresponding speech coding, and the shared encoder output (305). The method further comprises training the TTS model (501) based on the TTS loss determined for the training utterances in each set of training utterances to teach the TTS model how to learn to synthesize speech in each of the respective languages.
Owner:GOOGLE LLC