Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

55 results about "Speech code" patented technology

Method, device and equipment for training voice coding model and readable medium

The embodiment of the invention relates to a method and device for training a voice coding model, equipment and a readable medium. The method comprises: processing a speech feature representation of a speech sample using a first speech coding model to generate a set of discrete features; generating label information based on the set of discrete features, the label information comprising a set of labels, the set of labels indicating clustering centers corresponding to the corresponding discrete features; processing an intermediate feature representation generated based on the voice feature representation by using a second voice coding model to generate probability information corresponding to the tag information; the training loss is determined based on the label information, the probability information and the weight information, and the weight information is determined based on the distance from each discrete feature to the corresponding clustering center; and adjusting parameters of the second speech coding model based on the training loss. Therefore, the data volume required by training can be reduced, the stability of the unsupervised model training process can be improved, the training effect of the model is improved, and the model capability is further improved.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +1

Ultra-low bit rate voice coding and decoding system based on text semantic information fidelity

ActiveCN121214951ASemantic analysisBiological modelsIntelligibility (communication)Communications system
The invention provides an ultra-low bit rate voice coding and decoding system based on text semantic information fidelity, and relates to the technical field of voice coding and decoding. The system comprises the following steps: performing voice feature extraction and text feature extraction on original voice through a multi-modal text-voice combined encoder to obtain voice features and text features, embedding the text features into the voice features to obtain text-voice features of the original voice, and performing voice feature extraction and text feature extraction on the original voice through a multi-modal text-voice combined encoder; sending the text-speech features of the original speech to a receiving end in a semantic communication system; and the receiving end inputs the received text-speech features into a text semantic fidelity decoder based on an attention mechanism to obtain reconstructed speech features and reconstructed text features, decodes the reconstructed speech features layer by layer, and performs attention calculation on the reconstructed speech features and the reconstructed text features to obtain decoded speech. By means of the voice coding and decoding technology, good voice perception quality can still be kept under the ultra-low code rate, and voice intelligibility is guaranteed.
Owner:TSINGHUA UNIVERSITY

Speech recognition text scoring method and apparatus, electronic device, and storage medium

The application provides a speech recognition text scoring method and device, electronic equipment and storage medium. The speech recognition text scoring method obtains standard audio features of target speech based on a speech coding process in a speech recognition process of the target speech from the perspective of speech recognition, thereby providing a reference standard for scoring of the recognized text. From the perspective of scoring of the recognized text, audio distribution features corresponding to the first recognized text are obtained based on the first recognized text of the target speech. The first recognized text is scored in combination with the audio distribution features and the standard audio features. The method considers the correlation between the speech recognition text and the audio during scoring of the recognized text, thereby achieving scoring of the speech recognition text.
Owner:IFLYTEK CO LTD

Speech recognition method and device, electronic equipment and storage medium

The invention provides a speech recognition method and device, electronic equipment and a storage medium, and relates to the technical field of natural language processing, an adopted target automatic speech recognition model generates context semantic representation through a current text prefix sequence, and combines the context semantic representation with the current text prefix sequence and acoustic features to obtain a speech recognition result. A target text sequence is obtained through step-by-step prediction in an autoregression mode. Context semantic representation is introduced into the target automatic speech recognition model, the language switching moment can be accurately judged when the input speech signal is speech code conversion speech, speech recognition is carried out in time according to a new language during language switching, the recognition precision and robustness of a language switching boundary can be effectively improved, and the speech recognition efficiency is improved. And the speech code conversion speech recognition effect is improved. Moreover, context semantic representation is introduced during prediction, the challenge of ambiguity or ambiguity of acoustic signals can be overcome, the accuracy of the target text sequence is improved, and errors caused by untimely language model switching are reduced.
Owner:IFLYTEK CO LTD

Real-time speech recognition method, related device and computer program product

The invention discloses a real-time voice recognition method, related equipment and a computer program product, and the method carries out the audio coding of real-time voice stream data to be recognized, and obtains an audio coding feature as the basis of the subsequent voice recognition. The real-time voice coding module carries out coding processing on the audio coding features, and an acoustic vector dt of the token to be decoded at the current decoding moment can be obtained. The real-time speech coding module adopted by the invention is a coding module in a real-time speech recognition model for token-by-token prediction, so that the acoustic boundary of the current token to be decoded can be determined, and the acoustic vector dt is obtained. According to the method, the acoustic vector dt is further mapped to a large model input space, it is guaranteed that the vector dimension meets the input dimension requirement of the large model, the mapped acoustic vector dt is sent to the large model for forward reasoning calculation, and the real-time speech recognition effect can be improved by means of the powerful context modeling capacity and rich text field knowledge of the large model.
Owner:ANHUI IFLYTEK UNIVERSAL LANGUAGE TECH CO LTD

Speech synthesis method, device, equipment and readable medium

The embodiment of the invention provides a voice synthesis method and device, equipment, a storage medium and a program product. The method includes obtaining a speech coding representation corresponding to an initial speech of a first speaker, the speech coding representation indicating speech features unrelated to the speaker, the initial speech including speech content information. A pitch predictive encoded representation related to the initial speech is generated using a predictor model based on the speech encoded representation and a reference speech encoded representation, the reference speech encoded representation indicating a speech encoded representation corresponding to the second speaker, the pitch predictive encoded representation indicating pitch information of the second speaker. And generating a target speech corresponding to a second speaker by using a speech decoder model based on the pitch prediction coded representation and the reference speech coded representation, a speech signal of the second speaker including speech content information. In this way, the quality of the generated target voice is improved.
Owner:JINGDONG CITY BEIJING DIGITS TECH CO LTD +1

Lightweight semantic preserving coding and decoding and voice signal reconstruction method

PendingCN121905195ASpeech analysisNeural learning methodsNerve networkSpeech reconstruction
The invention discloses a lightweight semantic preserving coding and decoding and voice signal reconstruction method, which comprises the following steps of: acquiring an analog audio signal from an environment, inputting an original audio signal and an audio signal subjected to high-frequency filtering to two ends of a comparator, outputting an ultralow-bit-rate binary bit stream and transmitting the ultralow-bit-rate binary bit stream to a receiving end, and after the receiving end receives the signal, outputting the ultralow-bit-rate binary bit stream to the voice signal reconstruction end. The method comprises the following steps: inputting a voice signal to an end-to-end U-Net convolutional neural network of a bottleneck layer integrated LSTM module, outputting a reconstructed voice signal, training the network by using voice coding sample reconstruction loss and multi-scale short-time Fourier transform, and performing voice signal reconstruction on an ultra-low bit rate binary bit stream of a to-be-reconstructed audio after training. The technical contradiction that in resource-limited remote audio collection and distributed sensing tasks, node hardware resources are extremely limited, and ultra-low bit rate signal compression distortion causes that node microminiaturization deployment and high voice reconstruction quality cannot be considered at the same time is solved.
Owner:ZHEJIANG UNIV

Automatic speech recognition methods, devices, and media based on natural language processing and speech coding.

This application relates to an automatic speech code recognition method, device, and medium based on natural language processing, belonging to the field of voice communication technology. The automatic speech code recognition method based on natural language processing includes the following steps: S1, selecting a speech code; S2, restoring the speech information to a standard speech code speech data stream according to the selected speech code format; S3, obtaining the language matching information of the speech data stream; if the speech data stream does not meet the requirements for all languages ​​under the current speech code format, repeating steps S1-S3; if the speech data stream meets the requirements for any language under the current speech code format, proceeding to the next step; S4, obtaining the current speech code format, and outputting the speech information as a speech data stream in the current speech code format. This application can quickly identify speech codes and ensure smooth communication during communication failures.
Owner:WUHAN HUABO COMM CO LTD

AI edge server (B033)

ActiveCN309535918SCode generationEdge server
1. The name of the design product: AI edge server (B033). 2. The use of the design product: used to carry intelligent visual deep learning processor, through the combination of various open source large models, realize general large model, text to image, text to search image, text to search video, speech to text, text to speech, code generation and other applications. 3. The design points of the design product: in shape. 4. The picture or photo that best indicates the design points: perspective view 1.
Owner:深圳市元智奇点科技有限公司

A neural network-based lightweight streaming speech coding system and method

The application provides a neural network-based lightweight streaming speech coding system and method, which comprises an encoding compression end and a decoding reconstruction end; the encoding compression end comprises a speech encoder and a quantizer; the speech encoder comprises a time-frequency conversion module, a feature extraction module, an encoding end channel conversion module and a long-range time domain correlation extraction module; the feature extraction module comprises a time-frequency feature extraction submodule and a time-frequency size downsampling submodule; the quantizer comprises two or more layers of quantization modules, each layer of quantization module comprising an encoding end first domain conversion module, an encoding end vector quantization module, an encoding end inverse vector quantization module and an encoding end second domain conversion module; and the decoding reconstruction end comprises an inverse quantizer and a speech decoder. The application can effectively extract multi-scale feature information in a speech signal, so that high-quality speech compression coding and reconstruction can be realized at an extremely low code rate, and the parameter quantity is small, the complexity is low, and streaming coding and decoding can be realized.
Owner:NANJING UNIV

Using machine learning speech synthesizer using synthetic analytic speech coding

PendingCN122374817ASpeech codeSpeech synthesis
An apparatus includes a memory configured to store data associated with a machine learning (ML) based speech synthesis model. The apparatus also includes a speech encoder including the ML based speech synthesis model. The speech encoder is configured to perform a synthetic analysis operation of an input speech signal including generating a synthetic version of the input speech signal by the ML based speech synthesis model.
Owner:QUALCOMM INC

Adaptive speech codec adjustment method, apparatus, device and medium

The application discloses an adaptive speech coding adjustment method and device, equipment and medium, and the method is applied to a talking stage of a voice communication system. The application detects the voice quality under the current channel condition by using a real-time intelligent voice measurement algorithm, and adaptively adjusts the speech coding type based on the voice quality detection result, so that the voice quality of the conversation is improved, and the user experience is improved.
Owner:NO 30 INST OF CHINA ELECTRONIC TECH GRP CORP

Lightweight streaming speech coding system and method based on neural network

The invention provides a lightweight streaming speech coding system and method based on a neural network. The system comprises a coding compression end and a decoding reconstruction end, the coding compression end comprises a voice coder and a quantizer; the voice encoder comprises a time-frequency transformation module, a feature extraction module, an encoding end channel transformation module and a long-range time domain correlation extraction module. The feature extraction module comprises a time-frequency feature extraction sub-module and a time-frequency size sampling sub-module; the quantizer comprises more than two layers of quantization modules, and each layer of quantization module comprises a coding end first domain transformation module, a coding end vector quantization module, a coding end reverse quantity quantization module and a coding end second domain transformation module; and the decoding reconstruction end comprises an inverse quantizer and a voice decoder. According to the method, the multi-scale feature information in the voice signal can be effectively extracted, so that high-quality voice compression coding and reconstruction can be realized at an extremely low code rate, the parameter quantity is small, the complexity is low, and streaming coding and decoding can be realized.
Owner:NANJING UNIV

Voice coding method and device, electronic equipment and computer storage medium

The invention provides a voice coding method, a voice coding device, electronic equipment and a computer storage medium. Fusing the encoded semantic features output by the pre-training feature encoder, the video features output by the video encoder and the encoding features output by the existing voice encoder to generate a voice encoding result; and the sound quality of the coded voice is effectively improved.
Owner:UNIV OF SCI & TECH OF CHINA

Training of a voice conversion model and voice conversion method and device and related equipment

The application discloses a speech conversion model training method applied to the field of artificial intelligence. The speech conversion model provided by the application comprises an encoder, an instantiation normalization layer, a voiceprint extractor and a decoder. The method provided by the application comprises the following steps: inputting a speech sample of an original speaker into the encoder to obtain first speech coding data; the instantiation normalization layer removes the speech attribute of the original speaker in the first speech coding data to obtain a first speech hidden vector; a first voiceprint vector of a target speaker is obtained; the decoder synthesizes the first speech hidden vector and the first voiceprint vector to obtain reconstructed speech data; a first loss of the reconstructed speech data and the speech sample of the original speaker is calculated; whether the first loss reaches a maximum is judged; if not, the parameters of the instantiation normalization layer are optimized, and the foregoing steps are circularly performed until the first loss reaches the maximum, so that a trained speech conversion model is obtained.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speech synthesis methods, devices, electronic devices and storage media

This invention provides a speech synthesis method, apparatus, electronic device, and storage medium. The method includes: extracting text features of the text to be synthesized; searching for a target speech code matching the target speaker in a speaker codebook based on the identifier of the target speaker; and synthesizing speech from the text based on the text features and the target speech code to obtain synthesized speech. The speaker codebook stores speech codes of different speakers, where each speaker's speech code is a partial code selected from multiple basic speech codes of each speaker, and the total number of partial codes corresponding to all speakers is greater than or equal to the total number of speakers. The speech synthesis method, apparatus, electronic device, and storage medium provided by this invention significantly reduce model redundancy and improve the utilization rate of model parameters while ensuring the accuracy of speech synthesis.
Owner:ANHUI IFLYTEK UNIVERSAL LANGUAGE TECH CO LTD

Progressive language code conversion method for cross-language migration

The invention provides a progressive speech code conversion method for cross-language migration, which comprises the following steps: designing a difficulty measurer to measure the influence of each word replacement on a sentence, and then generating speech code conversion data with gradually increased difficulty based on a controllable temperature variable, finally, determining when to sample more difficult speech code conversion data for model training through a training scheduler; the method overcomes the defects that in the prior art, speech code conversion is prone to causing loss of original context information, insufficient representation cross-language alignment and limited data change, and the ability of model cross-language learning and knowledge migration is limited, and effective utilization of speech code conversion data and model generalization are achieved.
Owner:BEIHANG UNIV

A speech synthesis method, apparatus, device, and medium

This application provides a speech synthesis method, apparatus, device, and medium, belonging to the field of computer technology. The method includes: encoding style prompt speech into discrete speech code using a pre-trained speech quantizer to obtain prompt speech code; using a pre-trained text-to-discrete speech coding module, based on a penalty-before-sampling strategy, performing multiple inferences according to the target text and the prompt text and prompt speech code corresponding to the style prompt speech to obtain the target discrete speech code corresponding to the target text; and using a pre-trained discrete speech code to waveform module, converting the target discrete speech code into target speech based on the timbre prompt speech.
Owner:CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD

Deep learning-based technology consultation intention accurate identification and response method

The invention relates to the cross technical field of deep learning and technology consultation, in particular to a technology consultation intention accurate identification and response method based on deep learning. Comprising five core processes of multi-modal consultation access and preprocessing, multi-dimensional technical intention deep recognition, dynamic knowledge graph retrieval and adaptation, personalized response generation and optimization, and multi-round interactive feedback and iteration, and all the processes are cooperatively linked with a structured knowledge system through a deep learning model, so that technical consultation whole-process intelligent processing is realized. Through multi-modal information fusion, a technical field pre-training model and implicit intention mining, the technical consultation intention recognition accuracy is greater than or equal to 95%, which is significantly superior to that of a traditional method, response deviation caused by intention misjudgment is effectively avoided, multi-modal input of texts, pictures, voices, codes and the like is supported, technical consultation information is completely received, and the technical consultation intention recognition accuracy is improved. The problem of information missing caused by single text input is solved, and diversified expression scenes of technology consultation are adapted.
Owner:SHANGHAI QUWANG INFORMATION TECHNOLOGY CO LTD

Speech synthesis method and system based on shared subspace speech codec

The invention discloses a voice synthesis method and system based on a shared subsurface space voice codec. The method comprises the steps of shared subsurface space voice codec training, voice synthesis model training and voice synthesis. According to the invention, by constructing a unified coding framework sharing the submerged space, two different types of speech coding sequences of discrete speech coding and continuous speech coding can be obtained by using the same language codec, and a speech synthesis system based on the codec has the advantages of being convenient to train, high in speech synthesis quality and the like.
Owner:IOL WUHAN INFORMATION TECH CO LTD

Multi-mode-based intention recognition method and device and readable medium

The invention discloses an intention recognition method and device based on multiple modes and a readable medium. The intention recognition method comprises the steps that to-be-recognized texts, videos and voices of the same time sequence are obtained and input into a trained multi-mode intention recognition model; the text coding layer, the video coding layer and the voice coding layer are used for coding the text, the video and the voice to obtain text coding features, video coding features and voice coding features; inputting the text coding feature, the video coding feature and the voice coding feature into a shallow fusion module to obtain a final video pseudo tag and a final voice pseudo tag; and the text coding feature, the final video pseudo tag and the final voice pseudo tag are spliced and then input into a deep fusion module and a classifier to obtain an intention recognition prediction result and an output probability thereof. According to the method, the LLM architecture is fully utilized, and the cross attention and the LoRA module are utilized to perform modal alignment, so that the feature fusion effect and the intention recognition accuracy are effectively improved.
Owner:XIAMEN KUAISHANGTONG TECH CORP LTD

Method and apparatus for adjusting speech coding, electronic device, and storage medium

The application provides a speech coding adjustment method and device in dynamic spectrum sharing, electronic equipment and storage medium, and relates to the technical field of wireless communication. The method comprises the following steps: receiving a notification of whether a next first preset time length occupies a shared frequency band sent by a long term evolution (LTE) network at a current time; in response to the notification including that the LTE network needs to occupy the shared frequency band in the next first preset time length, correcting channel quality information, and adjusting a speech coding rate based on the corrected channel quality information. At least the problem that the LTE system has high priority and occupies the shared spectrum for a long time when the traffic volume is large, resulting in poor voice service quality of the UMTS system, is solved. The application is suitable for spectrum sharing optimization, voice service optimization and the like.
Owner:CHINA UNITED NETWORK COMM GRP CO LTD

Deep learning-based adaptive speech recognition system

This invention discloses a deep learning-based adaptive speech recognition system, relating to the field of speech recognition technology. The system processes raw speech signal data acquired by a speech acquisition module and dynamically adapts it to user identification information to obtain speech coding feature data. Furthermore, it extracts features from environmental metadata to obtain environmental embedding feature data. Based on the environmental embedding feature data and the speech coding feature data, feature recognition processing is performed to calculate recognition feature coefficients. The speech recognition module compares these coefficients with preset recognition feature thresholds and determines the speech quality based on the comparison results. This enables recognition under dynamically changing user and environmental conditions, improving the accuracy of speech recognition.
Owner:IANGSU COLLEGE OF ENG & TECH

Hearing aid system comprising haptic unit

The invention discloses a hearing aid system comprising a haptic unit. The hearing aid system comprises a hearing aid and a wrist-worn haptic device. The hearing aid comprises: an input unit comprising at least one microphone configured to pick up audio from a sound environment and configured to provide an audio input signal based on the audio; a signal processor including a speech enhancer configured to attenuate noise in the audio input signal and provide a speech enhancement signal based on the audio input signal; the signal processor is further configured to provide a voice coding signal based on the voice enhancement signal, and the voice coding signal comprises one or more of the following: a voice envelope signal, a voice timing signal and a voice intensity signal; an output unit configured to provide an auditory output sound indicative of the sound environment based on the speech enhancement signal; the wrist-worn haptic device includes an actuator configured to provide a haptic stimulus based on a speech-encoded signal, the haptic stimulus indicating a speech-enhanced signal; wherein the hearing aid and the haptic device are configured for wireless communication.
Owner:OTICON

Speech coding and decoding method, apparatus, computer device and storage medium

The application relates to a speech coding method, a speech coding device, a computer device, a storage medium and a computer program product, and a speech decoding method, a speech decoding device, a computer device, a storage medium and a computer program product. The speech coding method comprises the following steps: performing signal sub-band decomposition based on an initial speech signal to be coded to obtain a first sub-band excitation signal and a second sub-band excitation signal; the frequency value corresponding to the second sub-band excitation signal is higher than the frequency value corresponding to the first sub-band excitation signal; performing quantization processing on the first sub-band excitation signal according to a first quantization precision to obtain a first excitation quantization signal, and performing quantization processing on the second sub-band excitation signal according to a second quantization precision to obtain a second excitation quantization signal; the first quantization precision is greater than the second quantization precision; and target coding data is obtained based on the first excitation quantization signal and the second excitation quantization signal. The method can improve coding efficiency.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Speech recognition method, apparatus, device, and storage medium

ActiveCN115512695BSpeech recognitionSpeech codeSpeech classification
The application discloses a speech recognition method, device and equipment and a storage medium. A speech recognition model configured by the application predicts an initial predicted text based on speech coding features output by a speech encoder through a first speech classification layer. A text encoder encodes the initial predicted text, fuses the text coding features and the speech coding features, inputs the fused coding features into a shared encoder for secondary coding, and obtains a final predicted text based on the secondary coding features by a second speech classification layer. Since the speech recognition model can extract more rich fused coding features as a whole, the recognition accuracy can be further improved. In addition, since the speech recognition model includes a text encoder and a shared encoder, the text encoder and the shared encoder can be additionally trained using pure text data in the training process. Pure text data is easier to obtain in large quantities than annotated text of speech, greatly reducing the cost of manual annotation.
Owner:IFLYTEK CO LTD

Ultra low bit rate speech coding and decoding system based on text semantic information fidelity

ActiveCN121214951BSemantic analysisBiological modelsIntelligibility (communication)Communications system
The application provides a kind of based on text semantic information fidelity super low code rate speech coding system, it is related to speech coding technical field.The system includes: through multimodal text-speech joint encoder, speech feature extraction and text feature extraction are carried out to original speech, speech feature and text feature are obtained, and text feature is embedded into speech feature, text-speech feature of original speech is obtained, and text-speech feature of original speech is sent to receiving end in semantic communication system;The text-speech feature received by receiving end is input into text semantic fidelity decoder based on attention mechanism, to obtain reconstructed speech feature and reconstructed text feature, and the reconstructed speech feature is decoded layer by layer and attention calculation is carried out with the reconstructed text feature, to obtain decoded speech.Through the speech coding technology of the application, good speech perceptual quality can still be maintained under super low code rate, and the speech intelligibility is guaranteed.
Owner:TSINGHUA UNIVERSITY

Speech recognition methods, speech recognition devices, electronic devices and storage media

This application provides a speech recognition method, speech recognition device, electronic device, and storage medium, belonging to the field of artificial intelligence technology. The method includes: acquiring target speech data to be processed; inputting the target speech data into a speech recognition model including a convolutional network, a first encoding network, a second encoding network, and a decoding network; performing variable-dimensional processing on the target speech data through the convolutional network to obtain target speech convolutional features; extracting features from the target speech convolutional features through the first and second encoding networks respectively to obtain first speech coding features and second speech coding features; concatenating the first and second speech coding features to obtain target speech coding features; decoding the target speech coding features through the decoding network to obtain target word embedding vectors; and performing recognition processing on the target word embedding vectors through a preset function to obtain target text data, thereby improving the accuracy of speech recognition.
Owner:PING AN TECH (SHENZHEN) CO LTD

Audio decoding method and device, electronic equipment and storage medium

The invention provides an audio decoding method and device, electronic equipment and a storage medium, and relates to the field of audio decoding, and the method comprises the steps: receiving L2HC bit stream data, extracting a packet header segment in the L2HC bit stream data, and analyzing a hybrid coding mark in the packet header segment; extracting a side information segment and an effective load segment in the L2HC bit stream data, extracting a frequency division control mark and a voice coding control field from the side information segment when it is determined that the mixed coding mark represents voice coding, and extracting a voice data segment of a specified frequency band from the effective load segment according to the frequency division control mark; decoding the voice data segment according to the voice coding control field to obtain a voice audio; according to the technical scheme, the voice decoding exclusive field can be added in the side information segment of the L2HC bit stream data, the voice data segment is borne by the effective load segment of the L2HC bit stream data, and the voice data segment is independently decoded, so that high-quality voice audio decoding can be realized under the L2HC framework.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY