Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

83 results about "Language speech" patented technology

Speech is the verbal expression of language and includes articulation (the way sounds and words are formed). Language is the entire system of giving and getting information in a meaningful way. It's understanding and being understood through communication — verbal, nonverbal, and written.

Cross-language voice migration synthesis method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes such as financial science and technology, medical health and the like, and discloses a cross-language voice migration synthesis method, device, equipment and medium. And training an acoustic model by adopting a hierarchical self-adaptive fine tuning strategy, fusing the phoneme sequence for reasoning and the tone mark to generate a representation sequence, and finally synthesizing a target language speech signal. According to the method, a sharing and separation parallel cross-language modeling structure is constructed, the accuracy of phoneme and tone modeling in a low-resource language is effectively improved, the target language acoustic model has higher generalization ability and migration efficiency in combination with a multi-stage self-adaptive fine tuning and training data enhancement strategy, and the method is suitable for being applied to the field of multi-language modeling. And finally, synchronous improvement of the voice naturalness and the tone fidelity is realized.
Owner:PING AN TECH (SHENZHEN) CO LTD

Simultaneous interpretation data processing method and system based on POE microphone array

The invention relates to the technical field of simultaneous interpretation, and discloses a simultaneous interpretation data processing method and system based on a POE microphone array. The method comprises the following steps: synchronously acquiring multi-language original audio streams and meeting place environment noise spectrum features through a distributed microphone array powered by the Ethernet; after time domain framing is carried out on the audio stream, adaptive filtering is carried out by using a dynamic noise reduction weight coefficient to obtain a primary pure voice segment; dividing the multi-language speech endpoint detection model into independent speech units with language labels through a pre-trained multi-language speech endpoint detection model, and matching a corresponding acoustic model to generate a phoneme-level time alignment sequence; comparing and outputting a term replacement instruction stream in real time in combination with a simultaneous transfer term library, and generating an intermediate semantic representation vector after fusion; and the low-delay encoder converts the voice parameter sequence into a target language voice parameter sequence, and drives the waveform synthesizer to generate final simultaneous transmission audio. The method optimizes the whole process processing, gives consideration to the simultaneous transmission accuracy and real-time performance, and is suitable for a multilingual meeting place scene.
Owner:SUZHOU FUCHUAN TECH

Speech recognition method, device and system, electronic equipment and storage medium

The invention provides a voice recognition method, device and system, electronic equipment and a storage medium, and the method is applied to terminal equipment, and comprises the steps: carrying out the acoustic feature extraction of a voice signal based on the language information of the voice signal, obtaining an acoustic feature, decoding the acoustic feature, and obtaining a plurality of initial recognition results of the voice signal; determining a voice recognition result of the voice signal; the speech recognition result is obtained by applying a speech recognition model to perform semantic error correction on the basis of each initial recognition result, the acoustic feature and the language information; a speech recognition model is constructed on the basis of a large-scale language model, the defects that current multi-language speech recognition is low in accuracy and prone to misjudgment are overcome, through a two-step recognition process, a plurality of initial recognition results are generated locally and rapidly, and then the speech recognition accuracy is improved through the strong semantic understanding ability of the large-scale model. And the acoustic features and the language information are fused to carry out multi-modal deep error correction, so that the accuracy of multi-language speech recognition is greatly improved.
Owner:ANHUI IFLYTEK UNIVERSAL LANGUAGE TECH CO LTD

Personalized sound cloning method and device based on dynamic adaptation, equipment and medium

The invention discloses a personalized sound cloning method and device based on dynamic adaptation, equipment and a medium, and relates to the technical field of speech synthesis, and the method comprises the steps: carrying out the feature extraction of an original speech sample based on a preset deep neural network, so as to obtain a multi-modal speech feature, carrying out the fusion coding, and generating an acoustic representation model; determining to-be-synthesized text content and a target scene template, and extracting target acoustic parameters from a preset scene template library, so as to carry out acoustic parameter adjustment on the to-be-synthesized text content based on the target acoustic parameters and an acoustic representation model, and generating target cloned voice; extracting acoustic features irrelevant to languages by using a multi-language speech encoder, determining a to-be-migrated target language, and generating a corresponding text corpus; and carrying out phoneme mapping and alignment operation on the target clone voice based on the text corpus, and carrying out cross-language migration operation on the target clone voice through a preset generative adversarial network and the acoustic features to generate target language voice.
Owner:HUNAN AUTOMOTIVE ENG VOCATIONAL COLLEGE

Cross-language voice interaction method and system based on multi-modal semantic understanding, and storage medium

The invention discloses a cross-language voice interaction method and system based on multi-modal semantic understanding and a storage medium, and relates to the technical field of artificial intelligence and natural language processing. The method comprises the following steps: acquiring a source language speech stream, performing parallel double-flow feature extraction, extracting text semantic features through a mixed language code recognition model based on a unified phoneme space, and extracting acoustic features containing rhythm information at the same time; performing alignment fusion on the text and the acoustic features by using a cross-modal attention mechanism to generate multi-modal semantic representation; analyzing the explicit intention and the implicit emotion based on the representation, and generating a reply strategy and an emotion control label of the target language; and finally, synthesizing a target voice with a corresponding emotion color. According to the invention, the bottleneck that the traditional cascade architecture loses side language information is broken through, the precise understanding and strategic feedback of complex contexts such as Chinese-English mixed language codes (Code-mixing), anti-quiescence, hesitation and the like are realized, and the method is particularly suitable for transnational business negotiation and international customer service scenes.
Owner:杭州智慧沟通智能科技有限公司

Method, device, storage medium and electronic device for generating virtual voice

The present invention discloses a method, device, storage medium, and electronic device for generating virtual speech. The method comprises: obtaining multiple different speech text samples and speech attribute information, wherein each speech text sample in multiple different language speech text samples corresponds to a language and an object; inputting each speech text sample into a multi-stream encoder to obtain text features corresponding to each speech text sample; and training a preset speech acoustic model based on generative adversarial network modeling using the text features and speech features to obtain a target acoustic model for generating virtual speech. The present invention can support cross-language data training and the generation of cross-language speakers. The multi-stream encoder can better capture text features in different languages, improve the flexibility and reliability of virtual preset generation, and thus solve the technical problem of low flexibility and reliability in generating virtual speech in the prior art.
Owner:BEIJING UNISOUND INFORMATION TECH CO LTD

Yi language speech recognition method based on self-supervision and attention feature fusion

The invention relates to the technical field of natural language processing, and discloses a Yi language speech recognition method based on self-supervision and attention feature fusion. The method comprises a feature encoder module, a comparative learning module, a mask language modeling module, a joint optimization and feature fusion module and a decoder module, the feature encoder module adopts a convolutional neural network structure and converts continuous waveform signals into feature representation suitable for subsequent modeling, and the comparative learning module performs feature fusion on the continuous waveform signals through a Gumbel-Softmax technology. The method comprises the following steps that: a mask language modeling module and a feature fusion module are integrated, deviation caused by manual definition or clustering is avoided, the mask language modeling module obviously enhances semantic understanding and tone modeling capabilities of a model in a low-resource scene, a self-attention feature fusion mechanism is introduced into the feature fusion module, continuous features, discrete unit representation and semantic context representation from an acoustic level are spliced, and a self-attention feature fusion mechanism is introduced into the self-attention feature fusion mechanism. The decoder module adopts a decoder structure based on connection time sequence classification, and the corresponding relation between the voice and the text can be achieved without strict alignment labeling.
Owner:KUNMING UNIVERSITY

Speech translation method and device

The invention provides a speech translation method and device, and the method comprises the steps: translating source language speech data based on a speech translation model, and obtaining a target language text; a training target of the speech translation model comprises minimizing a difference between a first target language prediction text generated based on source language sample speech data and a translation label corresponding to the source language sample speech data, and minimizing a difference between speech features of the source language sample speech data and text features of a first source language sample text. And minimizing the difference between a second target language prediction text generated based on the pseudo source language speech features and a translation label corresponding to the second source language sample text. According to the method and the device, the problem of scarcity of annotated voice data is solved by efficiently utilizing relatively rich text data in a low-resource language scene, so that the performance of a voice translation model is improved.
Owner:IFLYTEK CO LTD

Cross-language speech recognition method, system and device and storage medium

The invention discloses a cross-language speech recognition method, system and device and a storage medium, and the method comprises the steps: carrying out the preprocessing of training speech, and obtaining a training speech frame sequence; extracting content characterization, speaker characterization and pitch characterization from the training voice frame sequence; performing voice reconstruction according to the content representation, the speaker representation and the pitch representation to obtain a target language voice; based on the training voice, constructing a target language voice recognition model according to the target language voice; in response to the target language recognition instruction, obtaining a target voice; and inputting the target voice into the target language voice recognition model to obtain a recognition result output by the target language voice recognition model, so that accurate modeling of dialect acoustic characteristics can be realized under the condition of extremely low annotation data through cross-language feature decoupling and a self-supervised migration mechanism. The recognition robustness of the complex tones and the characteristic vocabularies of the Guiwili is remarkably improved, and efficient generalization application in a dialect scene is achieved.
Owner:GUANGXI COMM IND SERVICE CO LTD +1

Methods and devices for task performance

Techniques for performing tasks are provided. An example method includes receiving, via the one or more input devices, a natural-language speech input including a request to perform a task; providing, at a language model, a plan corresponding to the task; determining whether the plan satisfies a set of resolution criteria; in accordance with a determination that the plan satisfies the set of resolution criteria, initiating performance of the task according to the selected plan; and in accordance with a determination that the plan does not satisfy the set of resolution criteria: providing a query to an information retrieval service requesting a set of resolution data; receiving, from the information retrieval service, the set of resolution data; resolving the plan based on the set of resolution data; and initiating performance of the task according to the resolved plan.
Owner:APPLE INC

A method and system for constructing a small language voice recognition based on a whisper token

The application provides a speech recognition method and system for constructing a small language based on a whisper token, and relates to the technical field of natural language processing and speech recognition. The method comprises the following steps: extracting all tokens related to a target small language in a whisper tokenizer to form an initial candidate set; matching and analyzing the tokens in the initial candidate set with collected training text corpus of the target small language, and counting the frequency of the tokens in the corpus; and screening high-frequency tokens and supplementing low-frequency tokens according to the frequency counting result to construct a dynamic vocabulary. The application improves the vocabulary quality, optimizes the model training efficiency, enhances the speech recognition accuracy, improves the model generalization ability, and simplifies the model construction process, thereby providing an efficient, accurate and easy-to-implement solution for the field of small language speech recognition.
Owner:BEIJING RUI KELUN INTELLIGENT TECH CO LTD

Performance optimization for real-time large language speech to text systems

Methods and systems for transcribing communications are provided. Methods may include receiving a communication. Methods may include splitting the communication into a plurality of communication segments. Each communication segment may include two or more words. Methods may include transcribing each segment included in the plurality of communication segments, in parallel. The transcribing may include using a transformer neural network to transcribe each segment included in the plurality of communication segments. Methods may include generating a transcription from the transcribing. The transcription may be generated by combining the transcription of each of the communication segments into a combined transcription. Methods may include correcting the combined transcription.
Owner:BANK OF AMERICA CORP

Speech recognition method and server

The application relates to a speech recognition method and a server. The method comprises the following steps: obtaining a to-be-recognized speech signal; recognizing each frame of the to-be-recognized speech signal according to an acoustic model of each language, and respectively outputting corresponding language phonemes and prediction probabilities; wherein the acoustic model of each language is respectively constructed according to shared hidden layer training; sequentially traversing a sentence decoding graph and a multi-language slot decoding graph connected with each other to obtain a corresponding path; wherein the sentence decoding graph is used for decoding phonemes entering a non-slot, and the slot decoding graph is used for decoding phonemes entering a slot; when it is determined that the path passes through the multi-language slot decoding graph in the speech decoding graph, screening the path according to the prediction probabilities of the language phonemes corresponding to each language, and determining the text information corresponding to the target path as a speech recognition result. The scheme provided by the application can accurately recognize mixed multi-language speech information.
Owner:GUANGZHOU XIAOPENG MOTORS TECH CO LTD

Speech synthesis method and apparatus, electronic device, and computer readable medium

The application discloses a speech synthesis method and device, electronic equipment and a computer readable medium, and relates to the technical field of speech synthesis. The method comprises the following steps: based on input text, obtaining a first synthesized speech according to a pre-acquired basic language speech synthesis model, and obtaining a second synthesized speech according to a pre-acquired target language speech synthesis model, wherein the similarity of the training speech of the target language speech synthesis model to the training speech of the basic language speech synthesis model is higher than a preset value; performing speech conversion on the second synthesized speech based on pre-acquired basic language training speech to obtain third synthesized speech; and obtaining target synthesized speech based on the first synthesized speech and the third synthesized speech. Therefore, the similarity of synthesized speech of different languages is further improved, and the target synthesized speech including bilingual or even multilingual speech has high timbre consistency, thereby improving the hearing effect.
Owner:VOICEAI TECH CO LTD

System and method for translating and transcribing

A computer-implemented method, computer program product and computing system for: receiving speech in a source language to define source language speech; performing a first token-based transcription of the source language speech into text of the source language using a first look-ahead encoder to define source language text; and performing a first token-based translation of the source language speech into text of a target language using a second look-ahead encoder to define target language text, wherein the first look-ahead encoder is smaller than the second look-ahead encoder.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Vehicle-mounted multi-language speech recognition method and system, electronic equipment and storage medium

The invention discloses a vehicle-mounted multi-language speech recognition method and system, electronic equipment and a storage medium, and relates to the field of cockpit speech control, and the method comprises the steps: collecting multi-language audio data of different acoustic environments in a cockpit; preprocessing the multi-voice audio data, extracting voice features, and outputting feature vectors; acquiring a multi-language speech recognition engine, and training the recognition engine according to the feature vector; obtaining target audio data, performing recognition through the trained multi-language speech recognition engine, and generating a recognition result; obtaining a verification mechanism, verifying the identification result according to the verification mechanism, and generating a verification result; generating a correction result set according to the verification result; and correcting the multi-language speech recognition engine according to the correction result set, and reducing the noise interference and increasing the robustness of speech recognition through a mode of firstly performing speech enhancement and then performing feature extraction.
Owner:CHINA FAW CO LTD +1

A Text Printing Method and System Based on Speech Recognition

This invention belongs to the field of text printing technology and discloses a text printing method and system based on speech recognition. The method includes the following steps: constructing a text printing template database; constructing a mixed-language speech recognition model; acquiring speech audio data in real time and performing speech recognition; matching several corresponding text printing templates; selecting a text printing template; fusing the real-time speech text data with the selected text printing template; and printing the real-time text printing data. The system includes a database construction unit, a model construction unit, a storage unit, a speech audio acquisition unit, a speech recognition application unit, a template matching unit, a template selection unit, a data fusion unit, and a printer. This invention solves the problems of low intelligence, poor speech recognition effect, low recognition efficiency, and lack of organic integration in existing technologies.
Owner:魏鹏飞

Speech translation method, server, storage medium and program product

The invention provides a speech translation method, a server, a storage medium and a program product. The method relates to the field of artificial intelligence. The method comprises the following steps: acquiring original voice of a to-be-translated original language and a to-be-translated target language; and searching translation knowledge matched with the original voice in a translation knowledge base, inputting the original voice and the translation knowledge matched with the original voice into a voice translation model, translating the original language voice in the translation knowledge appearing in the original voice into a corresponding target language text through the voice translation model, and generating a translation text of the target language. On the basis of a cross-modal translation knowledge base, external knowledge intervention is carried out on the speech translation model, translation errors of customized vocabularies are effectively avoided on the premise that parameters of the speech translation model do not need to be retrained, the translation accuracy of the speech translation model to customized vocabularies such as proper nouns, terminologies and new words is improved, and the translation efficiency is improved. Therefore, the accuracy of the speech translation result is improved.
Owner:ALIBABA (CHINA) CO LTD

Chinese language pronunciation intonation evaluation method

The invention discloses a Chinese language pronunciation intonation evaluation method, and belongs to the technical field of speech processing. The method comprises the following steps: firstly, dividing a Chinese language speech signal by adopting an autocorrelation method to obtain a plurality of sound state sets; eMD decomposition is carried out on short-time frames in the sound state set, and four characteristics of low-frequency gravity center, low-frequency energy, high-frequency gravity center and high-frequency energy are extracted; for a single sound state set, obtaining a sound state instability coefficient and constructing a sequence based on the four EMD features of each short-time frame, and meanwhile, respectively averaging each EMD feature; obtaining a change rate according to the EMD feature mean value difference of the adjacent sound state sets, constructing four change rate sequences, and subtracting each change rate sequence from the corresponding change rate storage sequence to obtain a change rate difference sequence; and finally, processing the four change rate difference sequences by using a pronunciation intonation evaluation neural network, and obtaining a pronunciation intonation score by combining the pronunciation intonation evaluation neural network with the sound state instability coefficient sequence. The precision of Chinese language pronunciation intonation evaluation is improved.
Owner:SICHUAN VOCATIONAL & TECHN COLLEGE OF COMM

Voice conversion method and device, electronic equipment and storage medium

PendingCN120727020ASpeech analysisSpeaking styleLanguage speech
The invention provides a voice conversion method and device, electronic equipment and a storage medium, and the method comprises the steps: carrying out the extraction of content features and fundamental frequency features of a source audio, and carrying out the extraction of speaker features of a target audio; inputting the content features, the fundamental frequency features and the speaker features into a voice conversion model for joint modeling processing, linear mapping processing and waveform reconstruction processing to generate a voice waveform; performing one-dimensional deep separation convolution processing and multi-receptive-field fusion processing on the voice waveform based on a vocoder in the voice conversion model to generate converted voice audio; wherein the converted voice audio shows that the speaking style of the source audio is replaced by the speaker style of the target audio. The voice conversion model is used for voice conversion, so that the timbre consistency and semantic retention capability of cross-language voice conversion are improved, and the naturalness and definition of reconstructed voice are improved.
Owner:BEIJING YUANJIAN INFORMATION TECH CO LTD

Speech synthesis methods, devices, media and electronic equipment

This disclosure relates to a speech synthesis method, apparatus, medium, and electronic device. The method includes: acquiring a phoneme sequence of a text to be synthesized and target prosodic information corresponding to the phoneme sequence, wherein the target prosodic information is prosodic information under a first language to which the text to be synthesized belongs; and synthesizing target speech based on the phoneme sequence, the target prosodic information, a speech vector of a first speaker under the first language, and a timbre vector of a second speaker, wherein the target speech represents the speech of the second speaker speaking the text to be synthesized under the first language. The speech synthesis method of this disclosure can effectively solve the problem of cross-language speech synthesis, and the synthesized speech has high pronunciation accuracy and natural prosody.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Pre-training method, system and electronic device of multilingual self-supervised model

Embodiments of the present application provide a multilingual self-supervised model pre-training method, system and electronic equipment. The method comprises: inputting unpaired unsupervised speech data selected from a multilingual data set to a language identification network to build a language classifier; extracting target language speech embedding and speech sentence embedding from a target language speech and speech sentence set; determining a training difficulty standard of extended dynamic curriculum learning of the multilingual self-supervised model based on an initial running loss of the multilingual self-supervised model, the target language speech embedding and the speech sentence embedding; and the language classifier performing pre-training of the multilingual self-supervised model in extended dynamic curriculum learning based on the training difficulty standard and a training set of different data amounts dynamically determined from the multilingual data set. The method of the embodiments of the present application makes the multilingual data set more efficient, eliminates potential harmful data on low-resource target language speech, and improves the performance of multilingual self-supervised learning downstream tasks.
Owner:AISPEECH CO LTD

Multi-language speech translation method

The invention discloses a multi-language speech translation method, which relates to the technical field of speech recognition and machine translation, and comprises the following steps of: receiving a speech signal containing intra-word code switching; performing phoneme-level identification on the voice signal through a multi-language acoustic model, and allocating a language tag to each phoneme; grouping the phonemes according to the language tags, and performing transcription according to a normal character rule of a corresponding language; segmenting and translating the transcribed text into a target language according to the recognized source language; the translated texts are combined, final output is generated, and the multi-language acoustic model adopts an end-to-end neural network architecture. According to the multi-language speech translation method, a multi-language acoustic model is adopted, language tags are distributed at a phoneme level instead of a conventional word or sentence level, a normal character rule that the language tags dynamically switch different languages is adopted, correct transcription of mixed words is ensured, a transcribed text is translated in a segmented mode according to the languages, and then the translated text is combined into final output. And translation accuracy is improved.
Owner:ZHEJIANG GONGSHANG UNIVERSITY

Cross-language speech recognition conversion method

The invention relates to a cross-language speech recognition conversion method, and relates to the technical field of language speech recognition processing, the cross-language speech recognition conversion method comprises a multi-language corpus composed of a matching model containing at least two languages, and the matching model comprises speech data of the languages and corresponding text data; the method comprises the following steps: S01, carding input voice data to obtain a de-noised audio set; s02, extracting and classifying features in the audio set to obtain text data and semantic data; and S03, inputting the text data into the obtained corresponding matching model based on the semantic data so as to obtain language specific text features. Through the steps of multi-language corpus building, fine feature processing, efficient acoustic model, flexible matching model and the like, accuracy and high efficiency of cross-language speech recognition are realized, the complexity of an algorithm is reduced, and the accuracy of a recognition result is remarkably improved.
Owner:宋婧婧 +1

Multi-speaker multi-lingual speech synthesis system based on self-learning text representation

The application discloses a multi-speaker multi-language speech synthesis system based on self-learning text representation, self-learning multi-language text representation, and is embodied in two modules, namely a text-to-SMTR prediction module and an SMTR-to-multi-language acoustic spectrum prediction module. Specifically, the application comprises the following steps: constructing an SMTR extraction method based on a self-learning system; constructing a multi-language text-to-SMTR prediction method; constructing an SMTR-to-multi-language acoustic spectrum prediction method; and constructing an end-to-end multi-language speech synthesis method based on SMTR fusion. The application can improve the accuracy of multi-language speech synthesis.
Owner:TIANJIN UNIV

Low-resource language speech recognition and model training method, device and program product

ActiveCN121506111BData setEngineering
The application discloses a text processing method and device, related equipment and computer program product. A data set composed of three different data of a target language is used to train a model in a LoRA fine-tuning manner to obtain three low-rank adaptive models corresponding to the three data sets respectively. The first data set uses real voice-text pair data, the second data set has the same text as the first data set and the voice is a synthesized voice, and the third data set includes collected high-resource text corpus and the voice is a synthesized voice. A task arithmetic merging strategy is used to calculate the sum of the first and third low-rank adaptive models and the difference of the second low-rank adaptive model to obtain a merged low-rank adaptive model. The voice recognition model of the target language is composed of the merged low-rank adaptive model and a pre-trained voice recognition model. The model is optimized by fully utilizing the synthesized data without affecting the real data effect, and the voice recognition effect of the low-resource language is improved.
Owner:ANHUI IFLYTEK UNIVERSAL LANGUAGE TECH CO LTD

Voice conversion related methods, systems and devices

The present application discloses methods, systems, devices and equipment related to speech conversion. Among them, the speech conversion method constructs a speech posterior probability graph PPG feature extractor and a speech synthesis model for each user; through the PPG feature extractor, based on the first acoustic feature data of the first speech data of the first user, the PPG feature data of the first speech data is determined; through the speech synthesis model of the second user, based on the PPG feature data and the second acoustic feature data of the first speech data, the second speech data of the second user corresponding to the first speech data is generated. This processing method uses the PPG feature as the input of the speech synthesis module. The PPG feature retains acoustic information (such as rhythm and pronunciation information). The PPG feature of each frame can be universal between different speakers and different languages, thereby realizing cross-language, cross-dialect, and same-language speech conversion, and can obtain the same speech conversion quality in different languages.
Owner:ALIBABA GROUP HOLDING LTD

Simultaneous interpretation delay elimination method, device, equipment, medium and program product

The invention provides a simultaneous interpretation delay elimination method, device and equipment, a medium and a program product, and relates to the technical field of simultaneous interpretation. The method comprises the following steps: acquiring a mixed audio comprising original language speech of at least one speaker; the original language voice of the target speaker is separated from the mixed audio, and background audio shielding the original language voice of the target speaker is obtained; performing simultaneous interpretation on the original language voice of the target speaker to obtain a target language voice of the target speaker; and fusing the target language voice with the background audio to obtain an output audio. According to the method, the original language voice serving as a delay perception anchor point is shielded, the possibility of time comparison of a user is directly eliminated, delay pseudo elimination is finally realized, and the user experience of simultaneous interpretation is improved; and the text content corresponding to the original language voice of the at least one speaker is displayed, so that the accuracy and convenience when the user selects the target speaker are improved, and the user experience is further improved.
Owner:IFLYTEK CO LTD

A method and apparatus for modeling an end-to-end speech translation model based on cross-language CTC

The application relates to an end-to-end speech translation model modeling method and device based on cross-language CTC, and belongs to the technical field of natural language processing; the method solves the problem that the speech translation method in the prior art ignores the guidance of target language text to an encoder and the monotone hypothesis and conditional independent hypothesis problems of CTC; the modeling method comprises the following steps: constructing an initial speech translation model; the initial speech translation model comprises an acoustic encoder, a text encoder and a decoder; obtaining a speech data set; the speech data set comprises source language speech data, source language labeled text corresponding to the speech data and target language labeled text; the initial speech translation model is trained by using the speech data set, is iteratively updated by using a loss function, and the speech translation model is obtained.
Owner:XIAONIU FANYI

Automated multi-speaker and multi-lingual speech analysis

PCT designated stageWO2026142921A1Semantic vectorSystems analysis
Exemplary system and methods use a combination of application modules and neural network architecture for multi-speaker and multi-language speech analysis. The exemplary system can receive a natural language input, which it decomposes into plural segments. A sub-group of the plural segments are accumulated in a buffer where each segment representing a period during which voice activity is detected. The sub-groups are analyzed for voice activity of multiple speakers and one or more text segments are generated based on the speakers. A semantic vector for each text segment is generated and stored in vector memory. Relevant data associated with each semantic vector is retrieved from the vector memory based on a similarity measure; and a response including specified information extracted from the one or more text segments is generated based on at least the relevant data.
Owner:ERESTECH