Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

57 results about "Language speech" patented technology

Speech is the verbal expression of language and includes articulation (the way sounds and words are formed). Language is the entire system of giving and getting information in a meaningful way. It's understanding and being understood through communication — verbal, nonverbal, and written.

Simultaneous interpretation data processing method and system based on POE microphone array

The invention relates to the technical field of simultaneous interpretation, and discloses a simultaneous interpretation data processing method and system based on a POE microphone array. The method comprises the following steps: synchronously acquiring multi-language original audio streams and meeting place environment noise spectrum features through a distributed microphone array powered by the Ethernet; after time domain framing is carried out on the audio stream, adaptive filtering is carried out by using a dynamic noise reduction weight coefficient to obtain a primary pure voice segment; dividing the multi-language speech endpoint detection model into independent speech units with language labels through a pre-trained multi-language speech endpoint detection model, and matching a corresponding acoustic model to generate a phoneme-level time alignment sequence; comparing and outputting a term replacement instruction stream in real time in combination with a simultaneous transfer term library, and generating an intermediate semantic representation vector after fusion; and the low-delay encoder converts the voice parameter sequence into a target language voice parameter sequence, and drives the waveform synthesizer to generate final simultaneous transmission audio. The method optimizes the whole process processing, gives consideration to the simultaneous transmission accuracy and real-time performance, and is suitable for a multilingual meeting place scene.
Owner:SUZHOU FUCHUAN TECH

Cross-language voice interaction method and system based on multi-modal semantic understanding, and storage medium

The invention discloses a cross-language voice interaction method and system based on multi-modal semantic understanding and a storage medium, and relates to the technical field of artificial intelligence and natural language processing. The method comprises the following steps: acquiring a source language speech stream, performing parallel double-flow feature extraction, extracting text semantic features through a mixed language code recognition model based on a unified phoneme space, and extracting acoustic features containing rhythm information at the same time; performing alignment fusion on the text and the acoustic features by using a cross-modal attention mechanism to generate multi-modal semantic representation; analyzing the explicit intention and the implicit emotion based on the representation, and generating a reply strategy and an emotion control label of the target language; and finally, synthesizing a target voice with a corresponding emotion color. According to the invention, the bottleneck that the traditional cascade architecture loses side language information is broken through, the precise understanding and strategic feedback of complex contexts such as Chinese-English mixed language codes (Code-mixing), anti-quiescence, hesitation and the like are realized, and the method is particularly suitable for transnational business negotiation and international customer service scenes.
Owner:杭州智慧沟通智能科技有限公司

Yi language speech recognition method based on self-supervision and attention feature fusion

The invention relates to the technical field of natural language processing, and discloses a Yi language speech recognition method based on self-supervision and attention feature fusion. The method comprises a feature encoder module, a comparative learning module, a mask language modeling module, a joint optimization and feature fusion module and a decoder module, the feature encoder module adopts a convolutional neural network structure and converts continuous waveform signals into feature representation suitable for subsequent modeling, and the comparative learning module performs feature fusion on the continuous waveform signals through a Gumbel-Softmax technology. The method comprises the following steps that: a mask language modeling module and a feature fusion module are integrated, deviation caused by manual definition or clustering is avoided, the mask language modeling module obviously enhances semantic understanding and tone modeling capabilities of a model in a low-resource scene, a self-attention feature fusion mechanism is introduced into the feature fusion module, continuous features, discrete unit representation and semantic context representation from an acoustic level are spliced, and a self-attention feature fusion mechanism is introduced into the self-attention feature fusion mechanism. The decoder module adopts a decoder structure based on connection time sequence classification, and the corresponding relation between the voice and the text can be achieved without strict alignment labeling.
Owner:KUNMING UNIVERSITY

Speech translation method and device

The invention provides a speech translation method and device, and the method comprises the steps: translating source language speech data based on a speech translation model, and obtaining a target language text; a training target of the speech translation model comprises minimizing a difference between a first target language prediction text generated based on source language sample speech data and a translation label corresponding to the source language sample speech data, and minimizing a difference between speech features of the source language sample speech data and text features of a first source language sample text. And minimizing the difference between a second target language prediction text generated based on the pseudo source language speech features and a translation label corresponding to the second source language sample text. According to the method and the device, the problem of scarcity of annotated voice data is solved by efficiently utilizing relatively rich text data in a low-resource language scene, so that the performance of a voice translation model is improved.
Owner:IFLYTEK CO LTD

Methods and devices for task performance

Techniques for performing tasks are provided. An example method includes receiving, via the one or more input devices, a natural-language speech input including a request to perform a task; providing, at a language model, a plan corresponding to the task; determining whether the plan satisfies a set of resolution criteria; in accordance with a determination that the plan satisfies the set of resolution criteria, initiating performance of the task according to the selected plan; and in accordance with a determination that the plan does not satisfy the set of resolution criteria: providing a query to an information retrieval service requesting a set of resolution data; receiving, from the information retrieval service, the set of resolution data; resolving the plan based on the set of resolution data; and initiating performance of the task according to the resolved plan.
Owner:APPLE INC

A method and system for constructing a small language voice recognition based on a whisper token

The application provides a speech recognition method and system for constructing a small language based on a whisper token, and relates to the technical field of natural language processing and speech recognition. The method comprises the following steps: extracting all tokens related to a target small language in a whisper tokenizer to form an initial candidate set; matching and analyzing the tokens in the initial candidate set with collected training text corpus of the target small language, and counting the frequency of the tokens in the corpus; and screening high-frequency tokens and supplementing low-frequency tokens according to the frequency counting result to construct a dynamic vocabulary. The application improves the vocabulary quality, optimizes the model training efficiency, enhances the speech recognition accuracy, improves the model generalization ability, and simplifies the model construction process, thereby providing an efficient, accurate and easy-to-implement solution for the field of small language speech recognition.
Owner:BEIJING RUI KELUN INTELLIGENT TECH CO LTD

Performance optimization for real-time large language speech to text systems

Methods and systems for transcribing communications are provided. Methods may include receiving a communication. Methods may include splitting the communication into a plurality of communication segments. Each communication segment may include two or more words. Methods may include transcribing each segment included in the plurality of communication segments, in parallel. The transcribing may include using a transformer neural network to transcribe each segment included in the plurality of communication segments. Methods may include generating a transcription from the transcribing. The transcription may be generated by combining the transcription of each of the communication segments into a combined transcription. Methods may include correcting the combined transcription.
Owner:BANK OF AMERICA CORP

Speech recognition method and server

The application relates to a speech recognition method and a server. The method comprises the following steps: obtaining a to-be-recognized speech signal; recognizing each frame of the to-be-recognized speech signal according to an acoustic model of each language, and respectively outputting corresponding language phonemes and prediction probabilities; wherein the acoustic model of each language is respectively constructed according to shared hidden layer training; sequentially traversing a sentence decoding graph and a multi-language slot decoding graph connected with each other to obtain a corresponding path; wherein the sentence decoding graph is used for decoding phonemes entering a non-slot, and the slot decoding graph is used for decoding phonemes entering a slot; when it is determined that the path passes through the multi-language slot decoding graph in the speech decoding graph, screening the path according to the prediction probabilities of the language phonemes corresponding to each language, and determining the text information corresponding to the target path as a speech recognition result. The scheme provided by the application can accurately recognize mixed multi-language speech information.
Owner:GUANGZHOU XIAOPENG MOTORS TECH CO LTD

Speech synthesis method and apparatus, electronic device, and computer readable medium

The application discloses a speech synthesis method and device, electronic equipment and a computer readable medium, and relates to the technical field of speech synthesis. The method comprises the following steps: based on input text, obtaining a first synthesized speech according to a pre-acquired basic language speech synthesis model, and obtaining a second synthesized speech according to a pre-acquired target language speech synthesis model, wherein the similarity of the training speech of the target language speech synthesis model to the training speech of the basic language speech synthesis model is higher than a preset value; performing speech conversion on the second synthesized speech based on pre-acquired basic language training speech to obtain third synthesized speech; and obtaining target synthesized speech based on the first synthesized speech and the third synthesized speech. Therefore, the similarity of synthesized speech of different languages is further improved, and the target synthesized speech including bilingual or even multilingual speech has high timbre consistency, thereby improving the hearing effect.
Owner:VOICEAI TECH CO LTD

System and method for translating and transcribing

A computer-implemented method, computer program product and computing system for: receiving speech in a source language to define source language speech; performing a first token-based transcription of the source language speech into text of the source language using a first look-ahead encoder to define source language text; and performing a first token-based translation of the source language speech into text of a target language using a second look-ahead encoder to define target language text, wherein the first look-ahead encoder is smaller than the second look-ahead encoder.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Vehicle-mounted multi-language speech recognition method and system, electronic equipment and storage medium

The invention discloses a vehicle-mounted multi-language speech recognition method and system, electronic equipment and a storage medium, and relates to the field of cockpit speech control, and the method comprises the steps: collecting multi-language audio data of different acoustic environments in a cockpit; preprocessing the multi-voice audio data, extracting voice features, and outputting feature vectors; acquiring a multi-language speech recognition engine, and training the recognition engine according to the feature vector; obtaining target audio data, performing recognition through the trained multi-language speech recognition engine, and generating a recognition result; obtaining a verification mechanism, verifying the identification result according to the verification mechanism, and generating a verification result; generating a correction result set according to the verification result; and correcting the multi-language speech recognition engine according to the correction result set, and reducing the noise interference and increasing the robustness of speech recognition through a mode of firstly performing speech enhancement and then performing feature extraction.
Owner:CHINA FAW CO LTD +1

A Text Printing Method and System Based on Speech Recognition

ActiveCN116339655BImprove user experiencereal-time identificationSpeech recognitionNeural architecturesComputer printingTimed text
This invention belongs to the field of text printing technology and discloses a text printing method and system based on speech recognition. The method includes the following steps: constructing a text printing template database; constructing a mixed-language speech recognition model; acquiring speech audio data in real time and performing speech recognition; matching several corresponding text printing templates; selecting a text printing template; fusing the real-time speech text data with the selected text printing template; and printing the real-time text printing data. The system includes a database construction unit, a model construction unit, a storage unit, a speech audio acquisition unit, a speech recognition application unit, a template matching unit, a template selection unit, a data fusion unit, and a printer. This invention solves the problems of low intelligence, poor speech recognition effect, low recognition efficiency, and lack of organic integration in existing technologies.
Owner:魏鹏飞

Speech translation method, server, storage medium and program product

The invention provides a speech translation method, a server, a storage medium and a program product. The method relates to the field of artificial intelligence. The method comprises the following steps: acquiring original voice of a to-be-translated original language and a to-be-translated target language; and searching translation knowledge matched with the original voice in a translation knowledge base, inputting the original voice and the translation knowledge matched with the original voice into a voice translation model, translating the original language voice in the translation knowledge appearing in the original voice into a corresponding target language text through the voice translation model, and generating a translation text of the target language. On the basis of a cross-modal translation knowledge base, external knowledge intervention is carried out on the speech translation model, translation errors of customized vocabularies are effectively avoided on the premise that parameters of the speech translation model do not need to be retrained, the translation accuracy of the speech translation model to customized vocabularies such as proper nouns, terminologies and new words is improved, and the translation efficiency is improved. Therefore, the accuracy of the speech translation result is improved.
Owner:ALIBABA (CHINA) CO LTD

Pre-training method, system and electronic device of multilingual self-supervised model

Embodiments of the present application provide a multilingual self-supervised model pre-training method, system and electronic equipment. The method comprises: inputting unpaired unsupervised speech data selected from a multilingual data set to a language identification network to build a language classifier; extracting target language speech embedding and speech sentence embedding from a target language speech and speech sentence set; determining a training difficulty standard of extended dynamic curriculum learning of the multilingual self-supervised model based on an initial running loss of the multilingual self-supervised model, the target language speech embedding and the speech sentence embedding; and the language classifier performing pre-training of the multilingual self-supervised model in extended dynamic curriculum learning based on the training difficulty standard and a training set of different data amounts dynamically determined from the multilingual data set. The method of the embodiments of the present application makes the multilingual data set more efficient, eliminates potential harmful data on low-resource target language speech, and improves the performance of multilingual self-supervised learning downstream tasks.
Owner:AISPEECH CO LTD

Multi-speaker multi-lingual speech synthesis system based on self-learning text representation

The application discloses a multi-speaker multi-language speech synthesis system based on self-learning text representation, self-learning multi-language text representation, and is embodied in two modules, namely a text-to-SMTR prediction module and an SMTR-to-multi-language acoustic spectrum prediction module. Specifically, the application comprises the following steps: constructing an SMTR extraction method based on a self-learning system; constructing a multi-language text-to-SMTR prediction method; constructing an SMTR-to-multi-language acoustic spectrum prediction method; and constructing an end-to-end multi-language speech synthesis method based on SMTR fusion. The application can improve the accuracy of multi-language speech synthesis.
Owner:TIANJIN UNIV

Low-resource language speech recognition and model training method, device and program product

ActiveCN121506111BData setEngineering
The application discloses a text processing method and device, related equipment and computer program product. A data set composed of three different data of a target language is used to train a model in a LoRA fine-tuning manner to obtain three low-rank adaptive models corresponding to the three data sets respectively. The first data set uses real voice-text pair data, the second data set has the same text as the first data set and the voice is a synthesized voice, and the third data set includes collected high-resource text corpus and the voice is a synthesized voice. A task arithmetic merging strategy is used to calculate the sum of the first and third low-rank adaptive models and the difference of the second low-rank adaptive model to obtain a merged low-rank adaptive model. The voice recognition model of the target language is composed of the merged low-rank adaptive model and a pre-trained voice recognition model. The model is optimized by fully utilizing the synthesized data without affecting the real data effect, and the voice recognition effect of the low-resource language is improved.
Owner:ANHUI IFLYTEK UNIVERSAL LANGUAGE TECH CO LTD

Simultaneous interpretation delay elimination method, device, equipment, medium and program product

The invention provides a simultaneous interpretation delay elimination method, device and equipment, a medium and a program product, and relates to the technical field of simultaneous interpretation. The method comprises the following steps: acquiring a mixed audio comprising original language speech of at least one speaker; the original language voice of the target speaker is separated from the mixed audio, and background audio shielding the original language voice of the target speaker is obtained; performing simultaneous interpretation on the original language voice of the target speaker to obtain a target language voice of the target speaker; and fusing the target language voice with the background audio to obtain an output audio. According to the method, the original language voice serving as a delay perception anchor point is shielded, the possibility of time comparison of a user is directly eliminated, delay pseudo elimination is finally realized, and the user experience of simultaneous interpretation is improved; and the text content corresponding to the original language voice of the at least one speaker is displayed, so that the accuracy and convenience when the user selects the target speaker are improved, and the user experience is further improved.
Owner:IFLYTEK CO LTD

A method and apparatus for modeling an end-to-end speech translation model based on cross-language CTC

The application relates to an end-to-end speech translation model modeling method and device based on cross-language CTC, and belongs to the technical field of natural language processing; the method solves the problem that the speech translation method in the prior art ignores the guidance of target language text to an encoder and the monotone hypothesis and conditional independent hypothesis problems of CTC; the modeling method comprises the following steps: constructing an initial speech translation model; the initial speech translation model comprises an acoustic encoder, a text encoder and a decoder; obtaining a speech data set; the speech data set comprises source language speech data, source language labeled text corresponding to the speech data and target language labeled text; the initial speech translation model is trained by using the speech data set, is iteratively updated by using a loss function, and the speech translation model is obtained.
Owner:XIAONIU FANYI

Automated multi-speaker and multi-lingual speech analysis

PCT designated stageWO2026142921A1Semantic vectorSystems analysis
Exemplary system and methods use a combination of application modules and neural network architecture for multi-speaker and multi-language speech analysis. The exemplary system can receive a natural language input, which it decomposes into plural segments. A sub-group of the plural segments are accumulated in a buffer where each segment representing a period during which voice activity is detected. The sub-groups are analyzed for voice activity of multiple speakers and one or more text segments are generated based on the speakers. A semantic vector for each text segment is generated and stored in vector memory. Relevant data associated with each semantic vector is retrieved from the vector memory based on a similarity measure; and a response including specified information extracted from the one or more text segments is generated based on at least the relevant data.
Owner:ERESTECH

Multi-language voice real-time translation method and system based on large model

The invention discloses a multi-language voice real-time translation method and system based on a large model, and belongs to the technical field of voice real-time translation, and the system comprises a voice collection module which is used for pre-collecting a voice sample of a speaker and collecting to-be-translated voice in real time; the voice recognition module is used for converting voice to be translated into a source language text; the large model translation module is used for sending the source language text to a selected large model API through a network request and receiving a target language text obtained after the selected large model translates the source language text in real time; the timbre feature extraction module is used for carrying out timbre feature extraction on a pre-collected voice sample; and the speech synthesis and output module is used for generating target language speech about the speaker from the extracted timbre features and the target speech text through a speech synthesis model. According to the method, the user experience is enhanced, the end-to-end translation delay is reduced, and application scenes such as real-time dialogues and conferences with extremely high delay requirements are met.
Owner:CHINA TELECOM DIGITAL INTELLIGENCE TECH CO LTD

Simultaneous interpretation data processing method and system based on POE microphone array

The present application relates to simultaneous interpretation technical field, disclose a kind of simultaneous interpretation data processing method and system based on POE microphone array.The method includes the synchronous acquisition of multilingual original audio stream and meeting environment noise spectrum characteristics by power over Ethernet distributed microphone array;After audio stream time domain framing, using dynamic noise reduction weight coefficient executes adaptive filtering and obtains preliminary clean speech segment;It is segmented into independent speech unit with language label by pre-training multilingual speech endpoint detection model, and phoneme level time alignment sequence is generated by matching corresponding acoustic model;Output terminology replacement instruction stream in real time by combining simultaneous terminology library comparison, generate intermediate semantic representation vector after fusion;It is converted into target language speech parameter sequence by low-delay encoder, and final simultaneous interpretation audio is generated by driving waveform synthesizer.The method optimizes whole process processing, and simultaneous interpretation accuracy and real-time performance are considered, and it is suitable for multilingual meeting scene.
Owner:SUZHOU FUCHUAN TECH

Multi-modal digital assistant

Systems and processes for a multi-modal digital assistant are provided. An example method includes, at a computer system, while operating in a hands-free operating mode, receiving a natural-language speech input indicative of a task; in response to receiving the natural-language speech input, initiating performance of the task; and providing an output corresponding to the task, wherein providing the output includes displaying a first user interface associated with an application, the first user interface displayed according to the hands-free operating mode; while displaying the first user interface, receiving, via the one or more input devices, a touch input at a location corresponding to the first user interface; and in response to the touch input: transitioning from the hands-free operating mode to a hands-on operating mode different than the hands-free operating mode; and launching the application, wherein launching the application includes displaying a second user interface according to the hands-on operating mode.
Owner:APPLE INC

Han-Zhuang translation and speech synthesis system based on Tacotron2 and HiFi-GAN

The invention discloses a Chinese-Zhuang translation and speech synthesis system based on Tacotron2 and HiFi-GAN, the system comprises text preprocessing, Chinese-Zhuang mapping and speech synthesis module realization, the text preprocessing comprises text standardization, text legality verification and word segmentation and part-of-speech tagging, the Chinese-Zhuang mapping is used for utilizing international phonetic symbol correspondence, and the Chinese-Zhuang translation and speech synthesis module is used for realizing Chinese-Zhuang translation and speech synthesis. The tone characteristics of Chinese can be reasonably converted into the tone mode of Zhuang language, and the implementation of the speech synthesis module comprises Tacotron2 model processing and a vocoder. Compared with the prior art, the method has the advantages that a seamless conversion system platform from Chinese text to Zhuang language speech synthesis is provided, the minority language is protected and inherited, convenient speech interaction experience is provided for Zhuang language users, the accessibility of information is enhanced, and the user experience is improved. The method is especially suitable for visually impaired people or scenes requiring hand-eye liberation.
Owner:GUANGXI UNIV FOR NATITIES +1

Low-resource language speech recognition and model training method, device and program product

The invention discloses a text processing method and device, related equipment and a computer program product, and the method comprises the steps: training a model according to an LoRA fine tuning mode through employing data sets composed of three different data of a target language, obtaining three low-rank adaptation models corresponding to the three data sets respectively, employing real voice-text pair data in the first data set, and employing real voice-text pair data in the second data set; the second data set is the same as the first data set in text, the voice is synthetic voice, the third data set comprises collected high-resource text corpora, and the voice is synthetic voice. And calculating a difference value between the sum of the first low-rank adaptation model and the third low-rank adaptation model and the second low-rank adaptation model by adopting a task arithmetic merging strategy to obtain a merged low-rank adaptation model, and forming a speech recognition model of the target language by the merged low-rank adaptation model and the pre-trained speech recognition model. The method achieves the optimization of the model through the full utilization of the synthetic data without affecting the real data effect, and improves the voice recognition effect of a low-resource language.
Owner:ANHUI IFLYTEK UNIVERSAL LANGUAGE TECH CO LTD

Prison break audio generation system and method based on IPA cross-language voice conversion

The invention discloses a jailbreak audio generation system and method based on IPA cross-language voice conversion in the technical field of artificial intelligence security. The system comprises an acquisition module used for acquiring original jailbreak information; the keyword extraction module is used for analyzing text segments of the original jailbreak information segment by segment and classifying the text segments; the cross-language similar speech conversion module is used for obtaining a Chinese attack audio similar to the original English pronunciation based on the sensitive segment; the audio synthesis module is used for generating and splicing audios to obtain corresponding candidate cross-language jailbreak audios; and the optimization output module is used for inputting the candidate cross-language jailbreak audio into the target model and optimizing the candidate cross-language jailbreak audio according to the response to obtain the cross-language jailbreak audio. The method does not need to depend on complex long prompts and background bedding, improves the flexibility and universality of attacks through cross-language similar voice conversion, can adapt to the rapid and hidden attack requirements in an actual scene, achieves more accurate pronunciation matching through IPA, and enhances the concealment and effectiveness of attacks.
Owner:SHANDONG UNIV

Cross-language speech synthesis method and system based on dual speaker embedding

The embodiment of the application provides a cross-language speech synthesis method based on double speaker embedding. The method comprises the following steps: inputting text and native language speaker embedding into a txt2vec acoustic model, determining the phoneme sequence code of the text through a text encoder, and determining the vector quantization acoustic feature and the auxiliary feature from the phoneme sequence code and the native language speaker embedding through a decoder; inputting the target language speaker embedding, the vector quantization acoustic feature and the auxiliary feature into a vec2wav vocoder, extracting the X-vector feature of the target language speaker embedding, inputting the X-vector feature, the vector quantization acoustic feature and the auxiliary feature into a feature encoder, and obtaining a cross-language acoustic feature; and determining the cross-language synthesized speech of the cross-language acoustic feature by using a generator. The embodiment of the application constructs a cross-language TTS model based on VQTTS, models the language speaking style and the speaker timbre respectively, and thus realizes cross-language speech synthesis with high nativeness and similar timbre to the target speaker.
Owner:AISPEECH CO LTD

Cross-lingual speech conversion method fusing an enhanced coding module and a codec structure of an LGNet network

The application discloses a cross-language speech conversion method of a coding-decoding structure of a fusion enhanced coding module and an LGNet network, comprising a training stage and a conversion stage, in the training stage, an encoder disentangles acoustic features to obtain speaker information representation and content representation; a U-shaped connection is adopted between the encoder and the decoder, and the speaker information representation is transmitted from the encoder to the decoder; an LGNet network further optimizes the extracted content representation; the decoder reconstructs the obtained speaker information representation and the optimized content representation; the application introduces an enhanced coding module in the encoder, and improves the quality of the converted speech; the LGNet network is used to make the optimized content representation of the source sentence and the speaker information representation of the target sentence fully fused in the adaptive instance normalization layer in the decoder, further improves the naturalness of the converted speech and the speaker similarity, and thus high-quality cross-language speech conversion is realized.
Owner:NANJING UNIV OF POSTS & TELECOMM

Speaker verification method and system based on dual-stream low-rank adaptive and adversarial decoupling

The application discloses a speaker verification method and system based on double-flow low-rank adaptive and anti-decoupling, comprising: based on a pre-trained speech network, a double-flow low-rank adaptive anti-decoupling network is constructed, and the original weight parameters of the pre-trained speech network are frozen; based on a language feature extraction branch and a speaker feature extraction branch, original speech data is subjected to feature extraction respectively to obtain language features and speaker features; the language features are input into the shared discriminator to perform language classification prediction to complete language boundary anchoring; after gradient inversion processing of the speaker features, the speaker features are input into the shared discriminator which has completed language boundary anchoring to perform anti-decoupling; identity recognition is performed based on the speaker features subjected to anti-decoupling constraint; corresponding training losses are calculated respectively, and iterative updating is performed based on the training losses. The application can improve the acceptance rate of cross-language speech of the same person and the rejection rate of the same language speech of different persons, and is suitable for high-precision speaker verification in a multi-language environment.
Owner:NANJING UNIV

Online conference translation method and system

The invention discloses an online conference translation method and system, and belongs to the technical field of voice recognition. According to the method, a target language mark generated by a sliding gesture is introduced in a recording stage, the mark reflects the language of a replied person facing the current speaking, the mark is combined with a native language mark of a speaker, a candidate language set is preferentially converged into a two-language set of native language and target language, and the recognition uncertainty is reduced. Therefore, the search space of speech recognition in the language dimension is remarkably shrunk from a'conference all language set 'to a'two-language set'. Based on this, the probability of misrecognition caused by multilingual competition is reduced, especially the probability of misrecognition of foreign language terms as native language near-pronunciation words is reduced, which languages are involved in each sentence of speech can be quickly determined, which language hybrid model is used for speech recognition can be quickly determined, and the speech recognition efficiency of the mixed language speech is improved.
Owner:QUEEN BEE NETWORK TECH (SHENZHEN) CO LTD

A real-time multi-lingual speech translation system based on single-screen interaction of earphone compartment

The present application belongs to the technical field of artificial intelligence and intelligent wearable devices, and specifically relates to a real-time multi-language speech translation system based on single-screen interaction of earphone compartments, which comprises an overlapping three-stage pipeline architecture combined with dynamic segmented speech processing, a ring-shaped semantic buffer area and a screen refresh synchronization mechanism; the earphone compartment integrates a double-microphone array and a single-color electronic ink screen. Through technical means such as overlapping pipeline processing mechanism, dynamic segmented window adjustment, hardware-level mutual exclusion control, layered reasoning architecture and electronic ink screen local refresh, the real-time multi-language speech translation function with low delay, high fidelity and bidirectional naturalness is realized in the single-screen interaction scene of the earphone compartment, and the synergistic effect of various technical features solves the timing coupling conflict between speech recognition, translation generation and screen rendering, effectively improving the user experience and system energy efficiency.
Owner:深圳卓隆智能电子有限公司