Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

229 results about "Speech coding" patented technology

Speech coding is an application of data compression of digital audio signals containing speech. Speech coding uses speech-specific parameter estimation using audio signal processing techniques to model the speech signal, combined with generic data compression algorithms to represent the resulting modeled parameters in a compact bitstream.

Dynamic interaction method based on multi-modal dynamic fusion large model and intelligent agent collaboration

The invention discloses a dynamic interaction method based on cooperation of a multi-modal dynamic fusion large model and an intelligent agent. The method comprises the following steps: performing feature extraction on user voice information to obtain a voice coding vector, a text semantic vector and an emotion feature vector; performing dynamic weight feature fusion on the voice coding vector, the text semantic vector and the emotion feature vector through a multi-modal dynamic fusion large model to obtain a fusion feature vector; inputting the fusion feature vector into an intention-scene coupling network, and identifying to obtain a user intention label; and identifying according to the user behavior log to obtain a user portrait tag, inputting the user intention tag and the user portrait tag into an autonomous decision-making agent, generating a target decision-making action through a lightweight policy network, and then interacting with the user according to the target decision-making action. The intelligent interaction efficiency and accuracy of the customer service system are improved, the interaction experience of the user is also improved, and the method can be widely applied to the technical field of artificial intelligence.
Owner:E SURFING IOT CO LTD

Voice generation method and device, equipment and medium

The invention relates to the technical field of speech synthesis, can be applied to business scenes such as financial science and technology, medical health and the like, and discloses a speech generation method, device, equipment and medium. And inputting a text speech language model to generate an intermediate code in combination with a speech code extracted based on a codebook generation mode, extracting a speaker feature vector in the prompt speech, decoding the intermediate code and the speaker feature vector by a generative adversarial decoder, and outputting a target speech. According to the method, fine-grained text representation is established by fusing character semantics and pinyin pronunciation information, and voice cloning is completed through unified voice language modeling and an adversarial generation mechanism in combination with codebook-driven acoustic coding and speaker personality characteristics, so that the naturalness, similarity and pronunciation accuracy of generated voices are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speech translation model training method, speech translation method and device based on cross-modal attention, global memory and dynamic convolution

The invention discloses a cross-modal attention, global memory and dynamic convolution-based speech translation model training method and device, and a speech translation method and device, and relates to the technical field of speech processing and machine translation. A speech translation model is designed to comprise a speech encoder, a text embedding layer, a cross-modal attention adapter, a large language model decoder, a global memory network, a dynamic convolution decoder and an output layer. The cross-modal attention adapter projects audio features and performs multi-head cross attention fusion with text embedding; the global memory network updates and enhances historical memory on the basis of a gating mechanism and a Transform Encoder; and the dynamic convolution decoder performs multi-scale convolution extraction on the decoded hidden representation and fuses with the memory, so that the translation quality is improved. According to the method, deep fusion of voice and text, continuous memory with contextual coherence and high-quality translation generation can be realized, the end-to-end voice translation performance is remarkably improved, and the actual requirements of real-time and high-quality end-to-end voice translation in a complex scene are met.
Owner:BEIJING YUNSHANG TECH CO LTD

Fine emotion control TTS method and device based on large model, equipment and medium

The invention discloses a fine emotion control TTS method and device based on a large model, equipment and a medium, relates to the technical field of artificial intelligence, and can generate a voice service with fine emotion expression and improve the user experience in the high-sensitivity emotion interaction fields of banks, financial customer service, insurance consultation, medical hospital guide and the like. Acquiring an emotional feature vector of the input text based on a preset large language model; encoding the emotion feature vector to obtain a voice encoding vector; determining an emotion curve vector of the input text based on a pre-trained neural network model; and generating target voice corresponding to the input text based on the voice coding vector and the emotion curve vector. According to the method, emotion feature vectors are extracted through a large language model, a dynamic emotion curve is generated in combination with a neural network, speech synthesis is cooperatively driven, emotion fineness and continuity are remarkably improved, and the problems of traditional TTS emotion expression roughening and fragmentation are solved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Speech coding method and decoding method based on time sequence modeling and related devices

The embodiment of the invention provides a voice coding method and decoding method based on time sequence modeling and a related device. The voice coding method comprises the following steps: acquiring target voice; extracting voice features corresponding to the voice waveform of the target voice through an input cavity convolution extraction module; performing down-sampling processing on the voice features through a down-sampling module to obtain a down-sampling sequence; performing long time sequence modeling on the down-sampling sequence through a long and short time memory neural network added with exponential activation and matrix operation to obtain a coding sequence; and converting the feature dimension of the coding sequence into 128 dimensions through the first convolutional layer to obtain a target code stream. According to the embodiment of the invention, the feature extraction module, the down-sampling module, the long and short term memory neural network and the output cavity convolution layer code the target voice, so that the compression of the voice with the extremely low bit rate can be realized, and the reconstruction of the voice with the extremely low bit rate and high quality under 150bps is realized.
Owner:SUN YAT SEN UNIV

Signal encoding using latent feature prediction

Techniques and solutions are described for encoding and decoding signals, such as audio data. Disclosed innovations can find particular use in speech coding applications, such as for real time communications. Using a neural network, contextual coding can be used to encode latent features for a current frame using a prediction from reconstructed latent features of past frames as a context. An extractor learns a residual-like feature based on such prediction and latent features of the current frame obtained using an encoder. The residual-like feature is then quantized. At a decoder portion of a coding framework, the quantized feature is dequantized and then combined with a prediction from prior reconstructed latent features to provide reconstructed features of a current frame, which can then be processed by a decoder to provide a reconstructed signal.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Voice processing method and device, equipment, medium and product

The invention provides a voice processing method and device, equipment, a medium and a product. The method comprises the following steps: acquiring a target voice and a translation mode of the target voice; based on a voice encoder module and a translation mode in the voice processing model, encoding the target voice to obtain voice content features and acoustic features of the target voice; translating the target voice based on a large language model, a translation mode and voice content features in the voice processing model to obtain a target translation text of the target voice; and based on a voice decoder module, the translation mode and the acoustic features in the voice processing model, performing voice synthesis on the target translation text to obtain a target translation voice of the target voice. According to the invention, adaptive processing is carried out on voice coding, translation and voice decoding through a voice processing model of an integrated framework integrating voice translation and voice simultaneous transmission in combination with a translation mode, so that the deployment cost is better reduced, and the real-time performance and quality of voice processing are optimized.
Owner:XIAN XUNFEI SUPER BRAIN INFORMATION TECH CO LTD

Dialect recognition model training method and device and dialect recognition method

The invention discloses a training method and device of a dialect recognition model and a dialect recognition method. The model training method comprises the following steps: training a first initial model comprising a feature extraction module, a voice encoder and a natural language large model by using a first dialect sample set with a text label to obtain a first dialect recognition model; training a second initial model containing a feature extraction module, a voice encoder and a voice decoder in the first dialect recognition model by using a text result of a first dialect sample set predicted by the first dialect recognition model to obtain a second dialect recognition model; and finally, training a second dialect recognition model by using the second dialect sample set without text annotation and a text result of the second dialect sample set predicted by the first dialect recognition model to obtain a target dialect recognition model. The technical problems that a large number of samples are needed in traditional dialect recognition model training, the data quality is difficult to guarantee, and the labeling cost is high are solved.
Owner:CHINA TELECOM CORP LTD

Speech encoder training method and apparatus, device, medium, and program product

Disclosed are a speech encoder training method performed by a computer device. The method includes: masking a first sub-feature representation at a first feature position in a first text feature representation to obtain a first masked feature representation; performing feature prediction on a masked first feature position in the first masked feature representation based on a first speech feature representation to obtain a first predicted feature representation; and training a first speech encoder based on a difference between the first predicted feature representation and the first sub-feature representation to obtain a second speech encoder. The first speech encoder is trained by combining data in a speech modality with data in a text modality, and information included in the data in the text modality is adopted so that the first speech encoder can learn relatively high-level semantic representations of speech, thereby improving the prediction accuracy of representations.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Large model real-time voice interaction method and device based on autoregression voice synthesis

The invention provides a large-model real-time voice interaction method and device based on autoregressive voice synthesis, and the method comprises the steps: obtaining a voice instruction marked with a target text response and a target voice response, enabling a voice encoder to code the voice instruction into voice representation, and enabling a voice adapter to carry out the dimension reduction and feature conversion of the original voice representation; the large language model generates a hidden state according to the converted voice representation and samples the hidden state to obtain a text sequence; and processing the text sequence by adopting a text-voice language model based on an autoregression Transform structure, generating a voice marking sequence in a streaming manner, and converting the voice marking sequence into a voice signal through a vocoder. According to the method provided by the invention, the naturalness and fluency of speech synthesis are greatly improved while high real-time performance is ensured. The optimized voice decoding architecture effectively reduces the voice generation delay and improves the response speed of the voice interaction system.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Speech synthesis method and system for controllable latent variable modeling based on semantic distillation

The invention relates to the technical field of speech synthesis, and particularly discloses a speech synthesis method and system for controllable latent variable modeling based on semantic distillation, and the method comprises the steps: converting a Mel spectrum into continuous latent variable distribution through a speech coding module, generating continuous latent variables through re-parameterization sampling, introducing a self-supervised model for semantic distillation, and carrying out the semantic distillation. According to the method, alignment of latent variables and semantic features is constrained through marginal cosine similarity and distance matrix structure loss, a text encoder maps a phoneme sequence into latent variable distribution, time sequence alignment of a text and the latent variables is achieved in combination with monotonic alignment search, and a decoder reconstructs the latent variables into a Mel spectrum. According to the method, waveform synthesis through a vocoder and total loss function joint optimization reconstruction, KL divergence, distillation, text alignment and confrontation loss are carried out, discrete information loss is avoided through continuous latent variable modeling, semantic consistency and text alignment efficiency are enhanced, the naturalness, coherence and real-time performance of synthesized voice are improved, and the method is suitable for scenes such as voice assistants and virtual anchors.
Owner:BEIJING TIMES RUILANG TECH CO LTD

Method, device and equipment for training voice coding model and readable medium

The embodiment of the invention relates to a method and device for training a voice coding model, equipment and a readable medium. The method comprises: processing a speech feature representation of a speech sample using a first speech coding model to generate a set of discrete features; generating label information based on the set of discrete features, the label information comprising a set of labels, the set of labels indicating clustering centers corresponding to the corresponding discrete features; processing an intermediate feature representation generated based on the voice feature representation by using a second voice coding model to generate probability information corresponding to the tag information; the training loss is determined based on the label information, the probability information and the weight information, and the weight information is determined based on the distance from each discrete feature to the corresponding clustering center; and adjusting parameters of the second speech coding model based on the training loss. Therefore, the data volume required by training can be reduced, the stability of the unsupervised model training process can be improved, the training effect of the model is improved, and the model capability is further improved.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +1

Efficient simultaneous interpretation method based on expert routing threshold

The invention discloses an efficient simultaneous interpretation method based on an expert routing threshold, and relates to the field of voice processing, the method is based on a classical Transform architecture model, an expert routing strategy model based on the expert routing threshold is constructed, and multi-language streaming translation is realized, the model comprises a streaming voice encoder, and the streaming voice encoder adopts a hybrid design and is connected with the expert routing threshold. The block-by-block autoregression block is composed of an autoregression block and a non-autoregression block; the text decoder is used for simultaneously processing the complete offline voice and the randomly truncated voice prefix to generate a hidden state; the routing threshold module is realized by a feedforward network and projects the final hidden state into a scalar value to determine an expert weight; and the hybrid expert post-processing module shares a language model head with the text decoder, and predicts a target translation sequence in combination with prefix information and global information. According to the method, a mixed expert threshold scheme is adopted to learn the strategy, the self-learning ability of the neural network is fully played, good effects are achieved in streaming translation and streaming TTS, and the method can be used for generating more streaming sequences.
Owner:SHANGHAI JIAOTONG UNIV

Speech recognition method and device, electronic equipment, storage medium and program product

The invention relates to a voice recognition method and device, electronic equipment, a storage medium and a program product. The method may comprise the following steps: acquiring a to-be-recognized voice, the to-be-recognized voice comprising a wake-up voice and a target voice; the to-be-recognized voice is input into a pre-trained voice recognition model to obtain a prediction text, the prediction text is a text used for representing voice content of a target user in the target voice, and the target user is a speaker waking up the voice; wherein the speech recognition model comprises a speech encoder, and the speech encoder is used for obtaining a first feature vector comprising semantic features and speaker features according to the speech to be recognized so as to determine a prediction text. Therefore, in a noisy scene such as multiple speakers, the voice content of the target user in the target voice can be accurately recognized. In addition, the voice encoder is subjected to self-supervised training by using single-person voice, so that the dependence on annotated data can be effectively reduced, and the generalization ability and adaptability of the model are enhanced.
Owner:XIAOMI EV TECH CO LTD +3

Ultra-low bit rate voice coding and decoding system based on text semantic information fidelity

ActiveCN121214951ASemantic analysisBiological modelsIntelligibility (communication)Communications system
The invention provides an ultra-low bit rate voice coding and decoding system based on text semantic information fidelity, and relates to the technical field of voice coding and decoding. The system comprises the following steps: performing voice feature extraction and text feature extraction on original voice through a multi-modal text-voice combined encoder to obtain voice features and text features, embedding the text features into the voice features to obtain text-voice features of the original voice, and performing voice feature extraction and text feature extraction on the original voice through a multi-modal text-voice combined encoder; sending the text-speech features of the original speech to a receiving end in a semantic communication system; and the receiving end inputs the received text-speech features into a text semantic fidelity decoder based on an attention mechanism to obtain reconstructed speech features and reconstructed text features, decodes the reconstructed speech features layer by layer, and performs attention calculation on the reconstructed speech features and the reconstructed text features to obtain decoded speech. By means of the voice coding and decoding technology, good voice perception quality can still be kept under the ultra-low code rate, and voice intelligibility is guaranteed.
Owner:TSINGHUA UNIVERSITY

Improving speech recognition by a machine learning model

The present disclosure describes techniques for improving speech recognition using a machine learning model. The machine learning model comprises a speech encoder configured to generate acoustic representations based on input speech, an adapter configured to generate adapted representations based on the acoustic representations, and a decoder configured to generate text corresponding to the input speech. A matching loss is applied during training the machine learning model. The matching loss is configured to explicitly force acoustic representations generated by the adapter to align with text embeddings. The machine learning model is fine-tuned by employing parameter-efficient low-rank adaptation. The machine learning model is trained to perform automatic speech recognition with performance improvement and parameter efficiency.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +1

Voiceprint comparison method and device based on large language model, and readable medium

The invention discloses a voiceprint comparison method and device based on a large language model, and a readable medium. The voiceprint comparison method comprises the following steps: acquiring a first voice to be compared and a second voice to be compared which are acquired respectively, and splicing the first voice and the second voice into a combined voice; the combined voice and the prompt words are input into a trained voiceprint comparison model, and the combined voice passes through an audio encoder to obtain voice encoding features; the prompt words pass through a text encoder to obtain text encoding features; inputting the voice coding features into an adapter, and performing dimension conversion on the voice coding features to obtain the voice coding features after dimension conversion; the text coding features and the voice coding features after dimension conversion are spliced and then input into an ontology structure of the improved large language model, output features of the ontology structure of the pre-trained large language model and output features of the LoRA module are added to obtain an output token sequence, the output token sequence passes through a text decoder to obtain a corresponding output text, and the output text is input into the LoRA module. And the robustness and accuracy of the model algorithm on voiceprint discrimination are improved.
Owner:XIAMEN KUAISHANGTONG TECH CORP LTD

Business complaint report generation method and device, equipment and medium

The invention relates to the technical field of voice semantics, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a business complaint report generation method, device, equipment and medium, and the method comprises the following steps: obtaining business data and customer complaint voice, carrying out text conversion on the customer complaint voice to obtain a customer complaint text, and sending the customer complaint text to a server; performing character recognition on the business data and the customer complaint text to obtain business data fields and text data fields; performing template filling on the customer complaint report template according to the business data field and the text data field to obtain an initial customer complaint report; performing voice coding on the customer complaint voice to obtain coding features, and performing emotion node recognition on the coding features to obtain emotion fluctuation nodes; performing emotion intensity analysis on the emotion fluctuation node to obtain an emotion intensity score; and performing content optimization on the initial customer complaint report according to the emotion fluctuation node and the emotion intensity score to obtain a target customer complaint report. The report generation accuracy and the report generation efficiency can be improved.
Owner:CHINA PING AN LIFE INSURANCE CO LTD

Speech recognition text scoring method and apparatus, electronic device, and storage medium

The application provides a speech recognition text scoring method and device, electronic equipment and storage medium. The speech recognition text scoring method obtains standard audio features of target speech based on a speech coding process in a speech recognition process of the target speech from the perspective of speech recognition, thereby providing a reference standard for scoring of the recognized text. From the perspective of scoring of the recognized text, audio distribution features corresponding to the first recognized text are obtained based on the first recognized text of the target speech. The first recognized text is scored in combination with the audio distribution features and the standard audio features. The method considers the correlation between the speech recognition text and the audio during scoring of the recognized text, thereby achieving scoring of the speech recognition text.
Owner:IFLYTEK CO LTD

Speech recognition method and device, electronic equipment and storage medium

The invention provides a speech recognition method and device, electronic equipment and a storage medium, and relates to the technical field of natural language processing, an adopted target automatic speech recognition model generates context semantic representation through a current text prefix sequence, and combines the context semantic representation with the current text prefix sequence and acoustic features to obtain a speech recognition result. A target text sequence is obtained through step-by-step prediction in an autoregression mode. Context semantic representation is introduced into the target automatic speech recognition model, the language switching moment can be accurately judged when the input speech signal is speech code conversion speech, speech recognition is carried out in time according to a new language during language switching, the recognition precision and robustness of a language switching boundary can be effectively improved, and the speech recognition efficiency is improved. And the speech code conversion speech recognition effect is improved. Moreover, context semantic representation is introduced during prediction, the challenge of ambiguity or ambiguity of acoustic signals can be overcome, the accuracy of the target text sequence is improved, and errors caused by untimely language model switching are reduced.
Owner:IFLYTEK CO LTD

Multi-modal speech recognition method based on large model, storage medium, electronic equipment and product

The invention relates to the technical field of voice recognition, and particularly provides a multi-modal voice recognition method based on a large model, a storage medium, electronic equipment and a product, and the method can comprise the steps: carrying out the preprocessing of an original voice signal of a user, and obtaining a processed voice signal; inputting the voice coding data and the historical dialogue data corresponding to the processed voice signal into a large language model to obtain a text vector corresponding to the processed voice signal; performing feature extraction on the processed voice signal to obtain a voice feature vector; processing a target vector sequence formed by splicing the voice feature vector and the text vector by using a pre-trained voice recognition module to obtain a text sequence; wherein the voice recognition module comprises a plurality of encoder layers and a plurality of decoder layers which are trained in advance; and cleaning and formatting the text sequence to obtain text data corresponding to the original voice signal. Some embodiments of the present application can improve the accuracy of speech recognition.
Owner:LINGXI TECHNOLOGY CO LTD

Dialogue voice generation method and system based on emotion perception adapter and large model reasoning

The invention discloses a dialogue voice generation method and system based on an emotion perception adapter and large model reasoning. The dialogue voice generation method adopted by the invention comprises the following steps: extracting voice emotion characteristics in original dialogue voice data by using a voice encoder and a time and hierarchical attention network; aligning the speech emotion features with text features of the large language model through an emotion perception coding module based on a query converter network, and generating emotion embedding compatible with the large language model; generating text embedding from statements of the input dialogue text by using a word segmentation device of a large language model; performing emotion embedding and text embedding by adopting an emotion adapter and a text adapter based on a partial low-rank adaptive network, and reasoning a text reply and a reply emotion state of the dialogue; and in combination with the text reply and the reply emotional state, generating a target voice conforming to an emotional context by using a voice generation model. According to the invention, the difference between the emotion and the text is effectively reduced, and the target voice with emotion consistency is generated.
Owner:STATE GRID ZHEJIANG ELECTRIC POWER CO MARKETING SERVICE CENT

Speech coding, speech decoding methods, devices, computer equipment, and storage media

The present application relates to a voice encoding, voice decoding method, device, computer device, and storage medium. The method includes: obtaining target feature information corresponding to a first frequency band based on initial feature information corresponding to the first frequency band in the initial frequency band feature information corresponding to the voice signal to be processed, performing feature compression on the initial feature information corresponding to the second frequency band in the initial frequency band feature information to obtain target feature information corresponding to the compressed frequency band, obtaining a compressed voice signal corresponding to the voice signal to be processed based on the target feature information corresponding to the first frequency band and the compressed frequency band, and performing encoding processing on the compressed voice signal through a voice encoding module to obtain encoded voice data. The sampling rate of the compressed voice signal is less than or equal to the supported sampling rate corresponding to the voice encoding module and less than the sampling rate corresponding to the voice signal to be processed, and the acquisition of the voice signal is not restricted by the sampling rate supported by the voice encoder.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Method and apparatus for encoding bone conduction speech, method and apparatus for decoding bone conduction speech, medium and device

The application discloses a bone conduction speech coding method, a coding and decoding method, a device, a medium and equipment, and belongs to the technical field of audio coding and decoding. The coding method comprises the following steps: updating a bandwidth detection parameter table in advance; updating a time domain noise shaping coding coefficient table in advance; in an LC3 audio coding process, the bone conduction speech is coded by using the updated bandwidth detection parameter table and the updated time domain noise shaping coding coefficient table, and a noise level estimation process of the bone conduction speech is performed on a bone conduction speech frequency band, so that the coding of the bone conduction speech is completed. The bandwidth detection parameter table and the time domain noise shaping coding coefficient table in the LC3 encoder are updated, so that when the LC3 encoder codes the bone conduction speech, the low-frequency part of the bone conduction speech can be recognized and corresponding coding processing can be performed, the coding of the bone conduction speech is realized, and the application range of the LC3 audio coding and decoding device is widened.
Owner:BEIJING BAIRUI INTERNET TECH CO LTD

Semantic perception real-time voice endpoint detection method and device

The semantic perception real-time voice endpoint detection method is applied to a voice interaction system, and comprises the following steps: receiving a voice stream of a user and a voice stream sent by the system from at least two audio channels; performing causal voice coding on the voice stream of the user and the voice stream sent by the system to obtain corresponding double-path voice representation; a causal two-way modeling architecture is constructed, and the causal two-way modeling architecture is realized by adopting an autoregressive prediction network and is used for receiving two-way voice representation and modeling a speech round conversion dynamic state between a user and a system through an embedded causal constraint mechanism and a cross-channel interaction mechanism; based on the output of the autoregressive prediction network, jointly predicting the category of a speech round conversion event and the remaining time from the current moment to the end of the speech round; and a smoothing processing mode is used to correct the predicted speech wheel conversion event type, and a final endpoint detection result is determined. According to the method, the endpoint detection stability in a complex dialogue scene can be improved.
Owner:INST OF ACOUSTICS CHINESE ACAD OF SCI

Joint speech text training for hybrid transducer and attention-based encoder-decoder modeling

A method includes generating speech embeddings corresponding to a received speech input using a speech encoder. The method also includes generating multi-modal embeddings based on at least one of the speech embeddings or a corresponding text embedding using a shared encoder. The method further includes generating conditioned multi-modal embeddings reflected by previous text predictions based on the multi-modal embeddings using a predictor. In addition, the method includes generating a text prediction corresponding to the received speech input based on the multi-modal embeddings and the conditioned multi-modal embeddings using a joint network.
Owner:SAMSUNG ELECTRONICS CO LTD

Real-time speech recognition method, related device and computer program product

The invention discloses a real-time voice recognition method, related equipment and a computer program product, and the method carries out the audio coding of real-time voice stream data to be recognized, and obtains an audio coding feature as the basis of the subsequent voice recognition. The real-time voice coding module carries out coding processing on the audio coding features, and an acoustic vector dt of the token to be decoded at the current decoding moment can be obtained. The real-time speech coding module adopted by the invention is a coding module in a real-time speech recognition model for token-by-token prediction, so that the acoustic boundary of the current token to be decoded can be determined, and the acoustic vector dt is obtained. According to the method, the acoustic vector dt is further mapped to a large model input space, it is guaranteed that the vector dimension meets the input dimension requirement of the large model, the mapped acoustic vector dt is sent to the large model for forward reasoning calculation, and the real-time speech recognition effect can be improved by means of the powerful context modeling capacity and rich text field knowledge of the large model.
Owner:ANHUI IFLYTEK UNIVERSAL LANGUAGE TECH CO LTD

Voice data processing method and device, storage medium and electronic equipment

The invention discloses a voice data processing method and device, a storage medium and electronic equipment, and the method comprises the steps: collecting to-be-processed user voice data, carrying out the task sharing coding processing of a voice recognition task and a named entity recognition task through a voice data processing model based on the user voice data, and obtaining a voice coding feature; and performing voice recognition processing based on the voice coding features to obtain a target voice recognition result, and performing named entity recognition based on the voice coding features to obtain a target entity recognition result. According to the method, one-time coding can serve two tasks of voice recognition and named entity recognition at the same time, consumption of computing resources is reduced, data processing efficiency is improved, and the method is particularly suitable for a scene of real-time processing of a large amount of voice data.
Owner:MIDEA GRP (SHANGHAI) CO LTD +1

System and method for performing identification based on scene representation and information related to the scene

A method for performing identification comprises: receiving a representation of a scene, receiving first information related to the scene for performing identification on the representation of the scene, and processing at least the representation of the scene and the first information to obtain a first output associated with an identification result. The method further includes receiving second information related to the scene for performing identification on the representation of the scene, and processing at least the first output and the second information to obtain a second output associated with an updated identification result. Processing the scene representation and first information may comprise feature extraction from the scene representation and first information, and fusing features so extracted to obtain the first output. Extraction of features from the scene may use a first encoder, e.g. unimodal image encoder; feature extraction from the information may use a second encoder e.g. unimodal text or speech encoder, and fusing of extracted features may use a third, e.g. multimodal, fusion encoder. Further processing, of the first output, may use a heatmap decoder. An artificial intelligence (AI), machine learning (ML) system may be used.
Owner:KK TOSHIBA

System and method for generating a facial image from a voice sample using a stylegan

System and method for reconstructing a facial image of a speaker from a voice sample of the speaker may include providing the voice sample of the speaker to a trained voice encoder to generate a voice embedding of the speaker, wherein the voice encoder is trained to provide a voice embedding that matches an image embedding of the facial image of the speaker; providing the voice embedding of the speaker to a trained mapping network to generate an intermediate latent vector, wherein the mapping network is trained to generate an intermediate latent vector for a StyleGAN from the voice embedding; and providing the intermediate latent vector to the StyleGAN to generate the facial image of a speaker.
Owner:CORSOUND AI LTD