Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

77 results about "Speech code" patented technology

Voice generation method and device, equipment and medium

The invention relates to the technical field of speech synthesis, can be applied to business scenes such as financial science and technology, medical health and the like, and discloses a speech generation method, device, equipment and medium. And inputting a text speech language model to generate an intermediate code in combination with a speech code extracted based on a codebook generation mode, extracting a speaker feature vector in the prompt speech, decoding the intermediate code and the speaker feature vector by a generative adversarial decoder, and outputting a target speech. According to the method, fine-grained text representation is established by fusing character semantics and pinyin pronunciation information, and voice cloning is completed through unified voice language modeling and an adversarial generation mechanism in combination with codebook-driven acoustic coding and speaker personality characteristics, so that the naturalness, similarity and pronunciation accuracy of generated voices are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Fine emotion control TTS method and device based on large model, equipment and medium

The invention discloses a fine emotion control TTS method and device based on a large model, equipment and a medium, relates to the technical field of artificial intelligence, and can generate a voice service with fine emotion expression and improve the user experience in the high-sensitivity emotion interaction fields of banks, financial customer service, insurance consultation, medical hospital guide and the like. Acquiring an emotional feature vector of the input text based on a preset large language model; encoding the emotion feature vector to obtain a voice encoding vector; determining an emotion curve vector of the input text based on a pre-trained neural network model; and generating target voice corresponding to the input text based on the voice coding vector and the emotion curve vector. According to the method, emotion feature vectors are extracted through a large language model, a dynamic emotion curve is generated in combination with a neural network, speech synthesis is cooperatively driven, emotion fineness and continuity are remarkably improved, and the problems of traditional TTS emotion expression roughening and fragmentation are solved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Method, device and equipment for training voice coding model and readable medium

The embodiment of the invention relates to a method and device for training a voice coding model, equipment and a readable medium. The method comprises: processing a speech feature representation of a speech sample using a first speech coding model to generate a set of discrete features; generating label information based on the set of discrete features, the label information comprising a set of labels, the set of labels indicating clustering centers corresponding to the corresponding discrete features; processing an intermediate feature representation generated based on the voice feature representation by using a second voice coding model to generate probability information corresponding to the tag information; the training loss is determined based on the label information, the probability information and the weight information, and the weight information is determined based on the distance from each discrete feature to the corresponding clustering center; and adjusting parameters of the second speech coding model based on the training loss. Therefore, the data volume required by training can be reduced, the stability of the unsupervised model training process can be improved, the training effect of the model is improved, and the model capability is further improved.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +1

Ultra-low bit rate voice coding and decoding system based on text semantic information fidelity

ActiveCN121214951ASemantic analysisBiological modelsIntelligibility (communication)Communications system
The invention provides an ultra-low bit rate voice coding and decoding system based on text semantic information fidelity, and relates to the technical field of voice coding and decoding. The system comprises the following steps: performing voice feature extraction and text feature extraction on original voice through a multi-modal text-voice combined encoder to obtain voice features and text features, embedding the text features into the voice features to obtain text-voice features of the original voice, and performing voice feature extraction and text feature extraction on the original voice through a multi-modal text-voice combined encoder; sending the text-speech features of the original speech to a receiving end in a semantic communication system; and the receiving end inputs the received text-speech features into a text semantic fidelity decoder based on an attention mechanism to obtain reconstructed speech features and reconstructed text features, decodes the reconstructed speech features layer by layer, and performs attention calculation on the reconstructed speech features and the reconstructed text features to obtain decoded speech. By means of the voice coding and decoding technology, good voice perception quality can still be kept under the ultra-low code rate, and voice intelligibility is guaranteed.
Owner:TSINGHUA UNIVERSITY

Speech recognition text scoring method and apparatus, electronic device, and storage medium

The application provides a speech recognition text scoring method and device, electronic equipment and storage medium. The speech recognition text scoring method obtains standard audio features of target speech based on a speech coding process in a speech recognition process of the target speech from the perspective of speech recognition, thereby providing a reference standard for scoring of the recognized text. From the perspective of scoring of the recognized text, audio distribution features corresponding to the first recognized text are obtained based on the first recognized text of the target speech. The first recognized text is scored in combination with the audio distribution features and the standard audio features. The method considers the correlation between the speech recognition text and the audio during scoring of the recognized text, thereby achieving scoring of the speech recognition text.
Owner:IFLYTEK CO LTD

Speech recognition method and device, electronic equipment and storage medium

The invention provides a speech recognition method and device, electronic equipment and a storage medium, and relates to the technical field of natural language processing, an adopted target automatic speech recognition model generates context semantic representation through a current text prefix sequence, and combines the context semantic representation with the current text prefix sequence and acoustic features to obtain a speech recognition result. A target text sequence is obtained through step-by-step prediction in an autoregression mode. Context semantic representation is introduced into the target automatic speech recognition model, the language switching moment can be accurately judged when the input speech signal is speech code conversion speech, speech recognition is carried out in time according to a new language during language switching, the recognition precision and robustness of a language switching boundary can be effectively improved, and the speech recognition efficiency is improved. And the speech code conversion speech recognition effect is improved. Moreover, context semantic representation is introduced during prediction, the challenge of ambiguity or ambiguity of acoustic signals can be overcome, the accuracy of the target text sequence is improved, and errors caused by untimely language model switching are reduced.
Owner:IFLYTEK CO LTD

Real-time speech recognition method, related device and computer program product

The invention discloses a real-time voice recognition method, related equipment and a computer program product, and the method carries out the audio coding of real-time voice stream data to be recognized, and obtains an audio coding feature as the basis of the subsequent voice recognition. The real-time voice coding module carries out coding processing on the audio coding features, and an acoustic vector dt of the token to be decoded at the current decoding moment can be obtained. The real-time speech coding module adopted by the invention is a coding module in a real-time speech recognition model for token-by-token prediction, so that the acoustic boundary of the current token to be decoded can be determined, and the acoustic vector dt is obtained. According to the method, the acoustic vector dt is further mapped to a large model input space, it is guaranteed that the vector dimension meets the input dimension requirement of the large model, the mapped acoustic vector dt is sent to the large model for forward reasoning calculation, and the real-time speech recognition effect can be improved by means of the powerful context modeling capacity and rich text field knowledge of the large model.
Owner:ANHUI IFLYTEK UNIVERSAL LANGUAGE TECH CO LTD

Hatred recognition method combining multi-layer hatred generation and progressive comparative learning

The invention belongs to the field of natural language processing, and relates to a hatred recognition method combining multi-layer hatred generation and progressive comparative learning, which comprises the following steps: acquiring a private hatred recognition model trained by speech input to obtain speech characteristics, and obtaining a hatred recognition result according to the speech characteristics; the training process of the hatred recognition model comprises the following steps: acquiring a speech set, and inputting the speech set into a feature coding module to obtain a speech coding feature fi; the speech coding characteristics comprise hatred speech coding characteristics fih and normal speech coding characteristics fin; inputting the characteristic fih into a multi-layer hatred generation module to obtain a multi-layer hatred speech Xih; inputting the speech Xih into a feature extraction module to obtain a multi-layer hatred speech feature Fih; calculating a total loss function value according to the features Fih and fi, and updating model parameters according to the total loss function value until a trained model is obtained; according to the method, different layers of hatred expression modes are fully covered through the multi-layer hatred generation module, so that various expressions can be accurately recognized, and the recognition robustness is improved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Speech synthesis method, device, equipment and readable medium

The embodiment of the invention provides a voice synthesis method and device, equipment, a storage medium and a program product. The method includes obtaining a speech coding representation corresponding to an initial speech of a first speaker, the speech coding representation indicating speech features unrelated to the speaker, the initial speech including speech content information. A pitch predictive encoded representation related to the initial speech is generated using a predictor model based on the speech encoded representation and a reference speech encoded representation, the reference speech encoded representation indicating a speech encoded representation corresponding to the second speaker, the pitch predictive encoded representation indicating pitch information of the second speaker. And generating a target speech corresponding to a second speaker by using a speech decoder model based on the pitch prediction coded representation and the reference speech coded representation, a speech signal of the second speaker including speech content information. In this way, the quality of the generated target voice is improved.
Owner:JINGDONG CITY BEIJING DIGITS TECH CO LTD +1

Robust speech recognition method and system based on multi-stage feature fusion

The invention discloses a robust speech recognition method and system based on multi-stage feature fusion, and relates to the technical field of speech recognition. The invention provides a speech recognition model for robust speech recognition, which comprises the following steps: firstly, extracting a magnitude spectrum | Y | from noisy speech Y through a speech coding part, coding the magnitude spectrum | Y | into a preliminary feature Ybasic, then, processing the Ybasic through a speech enhancement part to obtain a hidden feature Yhidden, a masking feature Ymask and a mapping feature Ymap, and finally, carrying out speech recognition on the hidden feature Yhidden, the masking feature Ymask and the mapping feature Ymap. Then, a fusion feature Ffuse is obtained through three-stage feature fusion, and finally, the Ffuse is decoded through a voice decoder to obtain Result. According to the model, feature complementation, feature interaction and high-low layer semantic alignment in the process are enhanced, the problem of voice distortion introduced by voice enhancement is systematically relieved, information loss in different stages is solved, and therefore the final voice recognition effect is guaranteed.
Owner:ANHUI UNIV

Lightweight semantic preserving coding and decoding and voice signal reconstruction method

PendingCN121905195ASpeech analysisNeural learning methodsNerve networkSpeech reconstruction
The invention discloses a lightweight semantic preserving coding and decoding and voice signal reconstruction method, which comprises the following steps of: acquiring an analog audio signal from an environment, inputting an original audio signal and an audio signal subjected to high-frequency filtering to two ends of a comparator, outputting an ultralow-bit-rate binary bit stream and transmitting the ultralow-bit-rate binary bit stream to a receiving end, and after the receiving end receives the signal, outputting the ultralow-bit-rate binary bit stream to the voice signal reconstruction end. The method comprises the following steps: inputting a voice signal to an end-to-end U-Net convolutional neural network of a bottleneck layer integrated LSTM module, outputting a reconstructed voice signal, training the network by using voice coding sample reconstruction loss and multi-scale short-time Fourier transform, and performing voice signal reconstruction on an ultra-low bit rate binary bit stream of a to-be-reconstructed audio after training. The technical contradiction that in resource-limited remote audio collection and distributed sensing tasks, node hardware resources are extremely limited, and ultra-low bit rate signal compression distortion causes that node microminiaturization deployment and high voice reconstruction quality cannot be considered at the same time is solved.
Owner:ZHEJIANG UNIV

Automatic speech recognition methods, devices, and media based on natural language processing and speech coding.

This application relates to an automatic speech code recognition method, device, and medium based on natural language processing, belonging to the field of voice communication technology. The automatic speech code recognition method based on natural language processing includes the following steps: S1, selecting a speech code; S2, restoring the speech information to a standard speech code speech data stream according to the selected speech code format; S3, obtaining the language matching information of the speech data stream; if the speech data stream does not meet the requirements for all languages ​​under the current speech code format, repeating steps S1-S3; if the speech data stream meets the requirements for any language under the current speech code format, proceeding to the next step; S4, obtaining the current speech code format, and outputting the speech information as a speech data stream in the current speech code format. This application can quickly identify speech codes and ensure smooth communication during communication failures.
Owner:WUHAN HUABO COMM CO LTD

A speech filtering method, device, storage medium and equipment

The present application discloses a speech filtering method, device, storage medium and equipment, which belongs to the field of speech coding and decoding technology. The method mainly includes: encoding the speech signal according to a standard Bluetooth encoder without a post-filtering module, and decoding the encoded speech signal to a transform domain noise shaping decoding module according to a standard decoder without a post-filtering module to obtain speech spectrum coefficients; inputting the speech spectrum coefficients into a pre-trained neural network model to obtain target spectrum coefficients corresponding to the speech spectrum coefficients; and according to the remaining decoding steps of the standard decoder without a post-filtering module, inputting the target spectrum coefficients into the low-latency improved inverse discrete cosine transform module of the standard decoder to obtain the target speech signal corresponding to the target spectrum coefficients. The present application omits the complex post-filtering operation in the Bluetooth encoding process, and only uses the pre-trained neural network model for filtering in the Bluetooth decoding process, so that it achieves a sound quality close to that of standard decoding.
Owner:BEIJING BAIRUI INTERNET TECH CO LTD

Speech code automatic identification method and device based on natural language processing, and medium

The invention relates to a voice code automatic identification method and device based on natural language processing and a medium, and relates to the technical field of voice communication, and the voice code automatic identification method based on natural language processing comprises the following steps: S1, selecting a voice code; s2, restoring the voice information into a voice data stream of standard voice coding according to the selected voice coding format; s3, obtaining a language condition that the voice data stream meets a requirement; if the voice data streams in all languages do not meet the requirements in the current voice coding format, repeating the steps S1 to S3; if the voice data stream in any language meets the requirement in the current voice coding format, executing the next step; and S4, acquiring a current voice coding format, and outputting a voice data stream by the voice information in the current voice coding format. The voice codes can be quickly recognized when communication is abnormal, and smooth communication is guaranteed.
Owner:WUHAN HUABO COMM CO LTD

AI edge server (B033)

ActiveCN309535918SCode generationEdge server
1. The name of the design product: AI edge server (B033). 2. The use of the design product: used to carry intelligent visual deep learning processor, through the combination of various open source large models, realize general large model, text to image, text to search image, text to search video, speech to text, text to speech, code generation and other applications. 3. The design points of the design product: in shape. 4. The picture or photo that best indicates the design points: perspective view 1.
Owner:深圳市元智奇点科技有限公司

Speech emotion recognition method and device based on target scene, equipment and medium

The invention relates to the technical field of speech semantics, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a speech emotion recognition method, device, equipment and medium based on a target scene, and the method comprises the steps: obtaining an initial speech signal, carrying out the signal coding of the initial speech signal, and obtaining a speech coding feature; performing acoustic feature extraction on the voice coding features to obtain initial acoustic features; performing feature hierarchy division on the initial acoustic feature to obtain a plurality of voice hierarchy features, and performing hierarchical attention processing on the plurality of voice hierarchy features to obtain a target voice feature; performing emotion feature extraction on the target voice feature to obtain a plurality of voice emotion features; and obtaining a target voice scene, mapping the plurality of voice emotion features to the target voice scene, and identifying a target voice emotion track in the mapped target voice scene. According to the invention, the accuracy and efficiency of speech emotion recognition are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

A neural network-based lightweight streaming speech coding system and method

The application provides a neural network-based lightweight streaming speech coding system and method, which comprises an encoding compression end and a decoding reconstruction end; the encoding compression end comprises a speech encoder and a quantizer; the speech encoder comprises a time-frequency conversion module, a feature extraction module, an encoding end channel conversion module and a long-range time domain correlation extraction module; the feature extraction module comprises a time-frequency feature extraction submodule and a time-frequency size downsampling submodule; the quantizer comprises two or more layers of quantization modules, each layer of quantization module comprising an encoding end first domain conversion module, an encoding end vector quantization module, an encoding end inverse vector quantization module and an encoding end second domain conversion module; and the decoding reconstruction end comprises an inverse quantizer and a speech decoder. The application can effectively extract multi-scale feature information in a speech signal, so that high-quality speech compression coding and reconstruction can be realized at an extremely low code rate, and the parameter quantity is small, the complexity is low, and streaming coding and decoding can be realized.
Owner:NANJING UNIV

Using machine learning speech synthesizer using synthetic analytic speech coding

PendingCN122374817ASpeech codeSpeech synthesis
An apparatus includes a memory configured to store data associated with a machine learning (ML) based speech synthesis model. The apparatus also includes a speech encoder including the ML based speech synthesis model. The speech encoder is configured to perform a synthetic analysis operation of an input speech signal including generating a synthetic version of the input speech signal by the ML based speech synthesis model.
Owner:QUALCOMM INC

Massive multilingual speech-text joint semi-supervised learning for text-to-speech

A method (600) includes receiving training data (301) including a plurality of sets of text-to-speech (TTS) spoken utterances (510), each set of TTS spoken utterances being associated with a respective language and including TTS utterances of synthetic speech including a corresponding reference speech representation (504) paired with a corresponding input text sequence (502). For each TTS utterance, the method includes generating a corresponding TTS encoded text representation for the corresponding input text sequence (512); generating a corresponding speech code for the TTS utterance of the corresponding synthetic speech (514); generating a shared encoder output (532, 534); generating a predicted speech representation for the TTS utterance of the corresponding synthetic speech (522); and determining a reconstruction loss (545). The method further includes training a TTS model based on the reconstruction loss determined for the TTS utterance in each set of TTS spoken training utterances (501).
Owner:GOOGLE LLC

Adaptive speech codec adjustment method, apparatus, device and medium

The application discloses an adaptive speech coding adjustment method and device, equipment and medium, and the method is applied to a talking stage of a voice communication system. The application detects the voice quality under the current channel condition by using a real-time intelligent voice measurement algorithm, and adaptively adjusts the speech coding type based on the voice quality detection result, so that the voice quality of the conversation is improved, and the user experience is improved.
Owner:NO 30 INST OF CHINA ELECTRONIC TECH GRP CORP

Bone conduction speech enhancement method and device based on cepstrum and deep learning, medium and equipment

The invention discloses a bone conduction speech enhancement method and device based on cepstrum and deep learning, a storage medium and equipment, and belongs to the technical field of Bluetooth audio coding and decoding, and the method comprises the steps: inputting PCM audio data, executing low-delay improved discrete cosine transform, and outputting a bone conduction speech spectrum coefficient; obtaining an enhanced bone conduction speech spectrum coefficient by combining cepstrum and deep learning according to the bone conduction speech spectrum coefficient; and according to the enhanced bone conduction speech spectrum coefficient, continuing to execute a standard LC3 coding process, and outputting an enhanced speech code stream. According to the invention, through combination of cepstrum and deep learning, LC3 coding is carried out on the bone conduction voice, the tone quality is enhanced, the user experience is improved, the feature extraction steps of coding and deep learning are reused, and the storage space and the computing power demand are saved.
Owner:CHONGQING BAIRUI INTERNET ELECTRONICS TECH CO LTD

Lightweight streaming speech coding system and method based on neural network

The invention provides a lightweight streaming speech coding system and method based on a neural network. The system comprises a coding compression end and a decoding reconstruction end, the coding compression end comprises a voice coder and a quantizer; the voice encoder comprises a time-frequency transformation module, a feature extraction module, an encoding end channel transformation module and a long-range time domain correlation extraction module. The feature extraction module comprises a time-frequency feature extraction sub-module and a time-frequency size sampling sub-module; the quantizer comprises more than two layers of quantization modules, and each layer of quantization module comprises a coding end first domain transformation module, a coding end vector quantization module, a coding end reverse quantity quantization module and a coding end second domain transformation module; and the decoding reconstruction end comprises an inverse quantizer and a voice decoder. According to the method, the multi-scale feature information in the voice signal can be effectively extracted, so that high-quality voice compression coding and reconstruction can be realized at an extremely low code rate, the parameter quantity is small, the complexity is low, and streaming coding and decoding can be realized.
Owner:NANJING UNIV

Voice coding method and device, electronic equipment and computer storage medium

The invention provides a voice coding method, a voice coding device, electronic equipment and a computer storage medium. Fusing the encoded semantic features output by the pre-training feature encoder, the video features output by the video encoder and the encoding features output by the existing voice encoder to generate a voice encoding result; and the sound quality of the coded voice is effectively improved.
Owner:UNIV OF SCI & TECH OF CHINA

A method, device and medium for adjusting speech coding and / or bit rate

The present invention provides a method, device and medium for adjusting speech coding and / or bit rate, the method comprising: determining whether the geographic clustering area corresponding to the terminal has changed based on the geographic location of the terminal; in response to the geographic clustering area corresponding to the terminal changing, obtaining the target coding and / or bit rate scheme corresponding to the changed current geographic clustering area based on a pre-constructed mapping table of geographic clustering areas and speech coding and / or bit rate schemes, and triggering the change of speech coding and / or bit rate based on the target coding and / or bit rate scheme; in response to the geographic clustering area corresponding to the terminal not changing, maintaining the current speech coding and / or bit rate unchanged. The method, device and medium can solve the problem that the existing speech coding and / or bit rate adjustment method does not take into account the obvious fluctuations in user experience caused by frequent adjustments, and that there is a risk of signaling storms and waste of air interface resources after the network starts adaptive adjustment on a large scale.
Owner:CHINA UNITED NETWORK COMM GRP CO LTD

Training of a voice conversion model and voice conversion method and device and related equipment

The application discloses a speech conversion model training method applied to the field of artificial intelligence. The speech conversion model provided by the application comprises an encoder, an instantiation normalization layer, a voiceprint extractor and a decoder. The method provided by the application comprises the following steps: inputting a speech sample of an original speaker into the encoder to obtain first speech coding data; the instantiation normalization layer removes the speech attribute of the original speaker in the first speech coding data to obtain a first speech hidden vector; a first voiceprint vector of a target speaker is obtained; the decoder synthesizes the first speech hidden vector and the first voiceprint vector to obtain reconstructed speech data; a first loss of the reconstructed speech data and the speech sample of the original speaker is calculated; whether the first loss reaches a maximum is judged; if not, the parameters of the instantiation normalization layer are optimized, and the foregoing steps are circularly performed until the first loss reaches the maximum, so that a trained speech conversion model is obtained.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speech synthesis methods, devices, electronic devices and storage media

This invention provides a speech synthesis method, apparatus, electronic device, and storage medium. The method includes: extracting text features of the text to be synthesized; searching for a target speech code matching the target speaker in a speaker codebook based on the identifier of the target speaker; and synthesizing speech from the text based on the text features and the target speech code to obtain synthesized speech. The speaker codebook stores speech codes of different speakers, where each speaker's speech code is a partial code selected from multiple basic speech codes of each speaker, and the total number of partial codes corresponding to all speakers is greater than or equal to the total number of speakers. The speech synthesis method, apparatus, electronic device, and storage medium provided by this invention significantly reduce model redundancy and improve the utilization rate of model parameters while ensuring the accuracy of speech synthesis.
Owner:ANHUI IFLYTEK UNIVERSAL LANGUAGE TECH CO LTD

Progressive language code conversion method for cross-language migration

The invention provides a progressive speech code conversion method for cross-language migration, which comprises the following steps: designing a difficulty measurer to measure the influence of each word replacement on a sentence, and then generating speech code conversion data with gradually increased difficulty based on a controllable temperature variable, finally, determining when to sample more difficult speech code conversion data for model training through a training scheduler; the method overcomes the defects that in the prior art, speech code conversion is prone to causing loss of original context information, insufficient representation cross-language alignment and limited data change, and the ability of model cross-language learning and knowledge migration is limited, and effective utilization of speech code conversion data and model generalization are achieved.
Owner:BEIHANG UNIV

Intelligent anti-packet-loss voice coding method and terminal for VOIP (Voice Over Internet Protocol)

The invention relates to the technical field of voice coding, in particular to an intelligent anti-packet-loss voice coding method and terminal for a VOIP (Voice Over Internet Protocol), and the method comprises the following steps: carrying out the parametric coding of an original voice frame based on a real-time network state parameter and a voice signal time-frequency feature, and generating a coding feature vector; inputting the coding feature vector into a hierarchical protection encoder, determining the bit allocation of a basic coding layer according to the fundamental frequency track parameter, generating an enhanced coding layer differential parameter based on the formant bandwidth parameter, and configuring a high-frequency compensation coding layer according to unvoiced and voiced sound identifiers; and in combination with the energy correlation characteristics of the adjacent voice frames, adaptive interleaving packaging of coding parameters is executed, and the packaging interval of the high-frequency compensation coding layer is in negative correlation with the current network packet loss rate. According to the method, the problem of sudden reduction of speech intelligibility caused by loss of unvoiced sound is effectively relieved, and the recovery capability and subjective hearing stability of a coding structure in a severe network are improved.
Owner:LVSWITCHES INC

A speech synthesis method, apparatus, device, and medium

This application provides a speech synthesis method, apparatus, device, and medium, belonging to the field of computer technology. The method includes: encoding style prompt speech into discrete speech code using a pre-trained speech quantizer to obtain prompt speech code; using a pre-trained text-to-discrete speech coding module, based on a penalty-before-sampling strategy, performing multiple inferences according to the target text and the prompt text and prompt speech code corresponding to the style prompt speech to obtain the target discrete speech code corresponding to the target text; and using a pre-trained discrete speech code to waveform module, converting the target discrete speech code into target speech based on the timbre prompt speech.
Owner:CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD

Deep learning-based technology consultation intention accurate identification and response method

The invention relates to the cross technical field of deep learning and technology consultation, in particular to a technology consultation intention accurate identification and response method based on deep learning. Comprising five core processes of multi-modal consultation access and preprocessing, multi-dimensional technical intention deep recognition, dynamic knowledge graph retrieval and adaptation, personalized response generation and optimization, and multi-round interactive feedback and iteration, and all the processes are cooperatively linked with a structured knowledge system through a deep learning model, so that technical consultation whole-process intelligent processing is realized. Through multi-modal information fusion, a technical field pre-training model and implicit intention mining, the technical consultation intention recognition accuracy is greater than or equal to 95%, which is significantly superior to that of a traditional method, response deviation caused by intention misjudgment is effectively avoided, multi-modal input of texts, pictures, voices, codes and the like is supported, technical consultation information is completely received, and the technical consultation intention recognition accuracy is improved. The problem of information missing caused by single text input is solved, and diversified expression scenes of technology consultation are adapted.
Owner:SHANGHAI QUWANG INFORMATION TECHNOLOGY CO LTD