Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

185 results about "Speech translation" patented technology

Speech translation is the process by which conversational spoken phrases are instantly translated and spoken aloud in a second language. This differs from phrase translation, which is where the system only translates a fixed and finite set of phrases that have been manually entered into the system. Speech translation technology enables speakers of different languages to communicate. It thus is of tremendous value for humankind in terms of science, cross-cultural exchange and global business. so when you type something in you can press a button and it will read it out and pronounce it for you so you know how to say it in a different language.

Face-translator: end-to-end system for speech-translated lip-synchronized and voice preserving video generation

A neural end-to-end system is provided for the face and voice preserving translation of videos. The system is a pipeline of multiple models that produces a video of the original speaker speaking in the target language with modified lip movement to match the target speech, while preserving emphases and prosody of the original speech, and voice characteristics of the original speaker. The pipeline starts with automatic speech recognition including emphasis detection, followed by the translation model. The translated text is then synthesized by a Text-to-Speech model that recreates the original emphases in the target sentence. The resulting synthetic speech is then converted back to the original speakers' voice using a voice conversion model. Finally, to synchronize the lips of the speaker with the translated audio, a generative model generates frames of adapted lip movements which are combined with the audio to produce the final output. The disclosure further describes several use-cases and configurations that apply these techniques to video conferencing, dubbing, low-bandwidth transmission, speech enhancement and assistive technology for the hearing impaired.
Owner:WAIBEL ALEXANDER

Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement

The invention relates to the technical field of artificial intelligence, and discloses a Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement, and the method comprises the steps: collecting a multi-channel audio signal through a built-in multi-microphone array of a Bluetooth device, carrying out the dynamic direction self-adaptive beam forming of the multi-channel audio signal, and carrying out the self-adaptive beam forming of the multi-channel audio signal; extracting a Mel spectrogram feature of the direction enhancement signal, identifying lip regions of a plurality of candidate speakers in each frame of real-time speaking video captured by a camera, performing time sequence convolution on the lip regions to obtain a lip movement time sequence embedded vector, calculating a correlation score with the Mel spectrogram feature, separating the direction enhancement signal, and obtaining a lip movement time sequence embedded vector; and performing text transcription and conversion on the high-confidence separation voice to obtain a translation language text, and sending the synthesized target translation voice to a preset mobile terminal through the Bluetooth device to obtain a target translation result. According to the method, the real-time performance and accuracy of speech translation are improved in a multi-person scene, far-field speech, noise interference and accent difference.
Owner:SHENZHEN DIE MICRO SEMICON CO LTD

Speech translation model training method, speech translation method and device based on cross-modal attention, global memory and dynamic convolution

The invention discloses a cross-modal attention, global memory and dynamic convolution-based speech translation model training method and device, and a speech translation method and device, and relates to the technical field of speech processing and machine translation. A speech translation model is designed to comprise a speech encoder, a text embedding layer, a cross-modal attention adapter, a large language model decoder, a global memory network, a dynamic convolution decoder and an output layer. The cross-modal attention adapter projects audio features and performs multi-head cross attention fusion with text embedding; the global memory network updates and enhances historical memory on the basis of a gating mechanism and a Transform Encoder; and the dynamic convolution decoder performs multi-scale convolution extraction on the decoded hidden representation and fuses with the memory, so that the translation quality is improved. According to the method, deep fusion of voice and text, continuous memory with contextual coherence and high-quality translation generation can be realized, the end-to-end voice translation performance is remarkably improved, and the actual requirements of real-time and high-quality end-to-end voice translation in a complex scene are met.
Owner:BEIJING YUNSHANG TECH CO LTD

Speech translation method, device, equipment and product

The invention provides a speech translation method, device, equipment and product, and is applied to the technical field of artificial intelligence. The speech translation method comprises the steps that speech to be translated and associated data of the speech to be translated are acquired, the associated data comprise associated images and / or associated videos, the associated images comprise image information of at least part of objects described by the speech to be translated, and the associated videos display lip actions when a speaking object expresses the speech to be translated; and based on the to-be-translated speech and the associated data, translating the to-be-translated speech through the multi-modal speech translation model to obtain a translation result of the to-be-translated speech. Therefore, in combination with the images and / or videos related to the voice, end-to-end voice translation is carried out through the multi-mode voice translation model, and the voice translation quality is effectively improved.
Owner:XIAN XUNFEI SUPER BRAIN INFORMATION TECH CO LTD

Voice processing method and device, equipment, medium and product

The invention provides a voice processing method and device, equipment, a medium and a product. The method comprises the following steps: acquiring a target voice and a translation mode of the target voice; based on a voice encoder module and a translation mode in the voice processing model, encoding the target voice to obtain voice content features and acoustic features of the target voice; translating the target voice based on a large language model, a translation mode and voice content features in the voice processing model to obtain a target translation text of the target voice; and based on a voice decoder module, the translation mode and the acoustic features in the voice processing model, performing voice synthesis on the target translation text to obtain a target translation voice of the target voice. According to the invention, adaptive processing is carried out on voice coding, translation and voice decoding through a voice processing model of an integrated framework integrating voice translation and voice simultaneous transmission in combination with a translation mode, so that the deployment cost is better reduced, and the real-time performance and quality of voice processing are optimized.
Owner:XIAN XUNFEI SUPER BRAIN INFORMATION TECH CO LTD

Speech translation method and device of end cloud translation system, equipment and medium

The invention relates to a speech translation method, device and equipment of an end cloud translation system and a medium, and relates to the technical field of intelligent speech translations, the end cloud translation system comprises an audio receiving end and a cloud, the method is applied to the audio receiving end, and the method comprises the following steps: obtaining a first language audio of the audio receiving end; sending the first language audio to a cloud; and receiving and playing a second language audio, wherein the second language audio is translated by the cloud according to the first language audio and then is sent out. The method can be widely applied to cross-border conference scenes, and the cross-language communication efficiency and experience are improved.
Owner:SHENZHEN XINZHILIAN SOFTWARE CO LTD

Speech translation using a wearable device

A speech translation system may provide real-time or near real-time translation of speech uttered by a person or emitted from a media device. The speech translation system may include a device that may receive audio representing speech in a source language and output audio representing speech in a target language. The speech translation system may translate the speech in portions representing semantically cohesive speech segments such that the target speech reflects the semantic meaning of words, phrases, and / or clauses as used in the context of the source speech. The speech translation system may condense the speech segments prior to or during translation to reduce verbosity. The speech translation system may selectively translate some speakers and not others, and may determine voice characteristics of source speech and apply identifying characteristics to the target speech that allow a user to differentiate respective target speech from different speakers based on the identifying characteristics.
Owner:AMAZON TECH INC

Methods for non-audible speech detection

Provided herein is a method for non-audible speech detection and output. The method comprises providing a radio frequency (RF) sensing device configured to be coupled to a head of a user. The method further comprises using the RF sensing device to collect RF signal data associated with movement of one or more speech articulators of the user. The method further comprises outputting or facilitating an output comprising a non-audible speech translation using at least in part processed RF signal data, wherein the non-audible speech of the user comprises continuous speech by the user.
Owner:REFLEX TECHNOLOGIES INC

Intelligent speech translation mobile phone and system capable of realizing multi-language inter-translation

The invention belongs to the technical field of mobile terminal communication and language translation, and particularly relates to an intelligent speech translation mobile phone and system capable of realizing multilingual inter-translation, which are characterized in that a two-way speech separation module and unit, a language recognition related module and unit, an end-cloud collaborative translation module and unit and a system are compatible with related module and unit collaboration; two-way voice collection, separation, language recognition and end-cloud collaborative translation are completed in a call scene, the system is compatible with a mainstream mobile operating system and communication application, a resource isolation and process linkage mechanism is adopted to guarantee operation compatibility, multilingual communication obstacles in the call scene are effectively solved, complex environment interference is adapted, the real-time performance and accuracy of translation are considered, and the system is suitable for large-scale popularization and application. And the native function of the terminal does not need to be modified.
Owner:SHENZHEN GUO ELECTRONIC INFORMATION CO LTD

Speech translation method and device

The invention provides a speech translation method and device, and the method comprises the steps: translating source language speech data based on a speech translation model, and obtaining a target language text; a training target of the speech translation model comprises minimizing a difference between a first target language prediction text generated based on source language sample speech data and a translation label corresponding to the source language sample speech data, and minimizing a difference between speech features of the source language sample speech data and text features of a first source language sample text. And minimizing the difference between a second target language prediction text generated based on the pseudo source language speech features and a translation label corresponding to the second source language sample text. According to the method and the device, the problem of scarcity of annotated voice data is solved by efficiently utilizing relatively rich text data in a low-resource language scene, so that the performance of a voice translation model is improved.
Owner:IFLYTEK CO LTD

Language recognition method and device, electronic equipment and product

The invention provides a language recognition method and device, electronic equipment, a storage medium and a product, and the method comprises the steps: carrying out the sliding extraction of a voice segment from a to-be-recognized voice according to a sliding window of a preset size, and detecting a voiceprint turning point from the to-be-recognized voice; under the condition that the voiceprint turning point is detected from the voice segment in the first sliding window, carrying out displacement adjustment on the first sliding window according to the voiceprint turning point so as to enable the voiceprint turning point to be located at the end point position of the first sliding window; and performing language recognition on the shifted voice segment in the first sliding window to obtain a language recognition result. According to the scheme, the voiceprint turning point in the sliding window can be detected, the jumping moment of the speaker in the sliding window can be recognized, the sliding window is moved according to the jumping moment to divide the voice segments, language recognition of the voice segments of multiple languages and a single language in the sliding window can be avoided, and the user experience is improved. Therefore, the accuracy of language recognition of the to-be-recognized speech is higher, and the accuracy of speech translation is further improved.
Owner:IFLYTEK CO LTD

Time-length-controllable end-to-end speech translation method and translation system

The invention discloses an end-to-end speech translation method and translation system capable of controlling duration, which effectively reduces the translation delay by introducing a speech end-to-end scheme. By constructing a brand new token and introducing a token alignment scheme, the length of a translation result is effectively controlled; by controlling the token duration and adding a length control variable, cross-language timbre and rhythm cloning is guided, so that relatively excellent cross-language timbre and rhythm cloning is achieved.
Owner:ZHILING WORKSHOP (HUZHOU) INTELLIGENT TECHNOLOGY CO LTD

Speech translation method, system, and storage medium

The application discloses a speech translation method and system and a storage medium, relates to the technical field of speech processing, and comprises the following steps: acquiring a real-time audio stream of source audio, and incrementally generating source language text corresponding to the real-time audio stream; detecting the semantic integrity of the generated source language text in real time; in the case where the semantic integrity is greater than or equal to the integrity threshold value corresponding to a minimum semantic unit, taking the generated source language text as a source language text block; determining target translation corresponding to the source language text block, and outputting the target translation. After the source language text is incrementally generated, the translation of the source language text block and the output are driven based on the semantic integrity, online segmentation translation of long sentences is realized, the translation waiting time is shortened, and the target translation and the original sound are synchronized.
Owner:ZHUHAI MOJIE TECH CO LTD

Artificial intelligence-based voice translation method and device, computer device and medium

The application is suitable for the field of digital medical technology, and particularly relates to a voice translation method and device based on artificial intelligence, computer equipment and medium. The application uses an acoustic encoder to perform feature coding on target voice to obtain an acoustic feature sequence, uses a boundary predictor to perform boundary prediction on the acoustic feature sequence to obtain a probability value of each feature value being predicted as a boundary and use the probability value as a weight of the corresponding feature value, performs weighted summation on all feature values to obtain acoustic contraction features, uses a semantic encoder to extract semantic features in the acoustic contraction features, uses a decoder to decode the semantic features to obtain target translation text in a preset target language, contracts the acoustic features to eliminate the length gap problem between voice features and text features, effectively inherits the knowledge of a pre-trained model when performing voice translation, improves the accuracy of the target translation text, and improves the work efficiency and work quality of doctors in the field of digital medical technology.
Owner:PING AN TECH (SHENZHEN) CO LTD

A crosswise intermodal transport obfuscation method and apparatus based on adaptive optimal transport

This invention relates to the field of speech translation technology, and particularly to a cross-transport obfuscation method and apparatus based on adaptive optimal transmission. The method includes: constructing an optimal transmission model within a multi-task general framework; performing attention-enhanced optimal transmission alignment on speech and text sequences; optimizing the attention-enhanced optimal transmission alignment based on a dynamic window strategy to obtain the alignment relationship between the speech and text sequences; fusing speech and text features of the speech and text sequences through a similarity-based adaptive fusion strategy; and combining a contrastive learning loss function with a multi-task learning framework, constructing positive and negative sample pairs, and combining the multi-task loss functions to obtain a unified optimization objective. This invention effectively reduces the representational differences between speech and text through dynamic window strategies, optimal transmission, and contrastive learning, achieving significant improvements in translation tasks for low-resource languages.
Owner:MINZU UNIVERSITY OF CHINA

Speech-to-speech translation

PendingUS20260154515A1Natural language translationSound input/outputSpeech to speech translationSpeech translation
A speech-to-speech translation method comprises transcribing speech spoken in a source language into transcribed text data in the source language using an on-premises speech recognition model. The transcribed text data is translated into translated text data in a target language using a first on-premises machine translation model. The translated text data is reverse translated into retranslated text data in the source language using a second, different on-premises machine translation model. The transcribed text and the retranslated text are displayed on a screen. The method also involves synthesizing, using an on-premises speech synthesis model, translated speech data in the target language based on the translated text data and play back, in response to a user confirmation, translated speech in the target language based on the translated speech data in the target language.
Owner:MABEL AI AB

Method and system for providing video conference service including artificial intelligence-based speech interpretation function

PCT designated stageWO2026146907A1Speech translationSpeech sound
Disclosed are a method and system for providing a speech conference service including an artificial intelligence-based speech interpretation function. According to one embodiment, the method for providing a video conference service may comprise the steps of: setting two or more interpretation bots participating as virtual participants in a video conference service; interpreting a speech of a first language input by participants of the video conference service into a speech of at least one other language different from the first language through processes of speech recognition, translation, and synthetic speech generation between the two or more interpretation bots and artificial intelligence; and providing the interpreted speech.
Owner:LINE PLUS

Voice real-time translation and proofreading system for cross-border e-commerce

The invention relates to the technical field of voice real-time translation, in particular to a voice real-time translation and proofreading system oriented to cross-border e-commerce, which comprises the following modules: a voice semantic processing module used for receiving voice input, converting the voice input into text data, performing semantic analysis on the text data, extracting commodity names, attribute information and logic relationships, and sending the extracted commodity names, attribute information and logic relationships to a server; semantic structure data is generated and stored in a context cache; and the anaphora analysis module is used for reading the semantic structure data from the context cache, detecting anaphora words in the semantic structure data, determining anaphora objects and updating the semantic data subjected to anaphora resolution to the context cache. According to the method, word segmentation, part-of-speech tagging, syntactic analysis and dependency analysis are carried out on a text, commodity names, the number, colors, sizes and attribute relations are extracted to generate structured semantic data, the structured semantic data are stored in a context cache, and when anaphora words are detected, anaphora objects are determined through semantic similarity and context logic analysis, and the semantic data are updated.
Owner:WUHAN CITY VOCATIONAL COLLEGE

system

We provide a system that enables more natural and accurate speech translation. [Solution] The system includes means for receiving the user's voice and acquiring audio data, means for providing speech recognition technology that analyzes the audio data and converts it into text data, means for motion recognition technology that captures the user's mouth movements and generates motion data, means for transmitting the text data and motion data to a server, means for the server to translate the text data into different languages ​​and convert the text in the different languages ​​into audio data, and means for transmitting the audio data to the user's terminal and playing it back.
Owner:SOFTBANK GROUP CORP

Speech translation with performance characteristics

An expressive speech translation system may process source speech in a source language and output synthesized speech in a target language while retaining vocal performance characteristics such as intonation, emphasis, rhythm, style, and / or emotion. The system may receive a transcript of the source speech, translate it, and generate transcript data. To generate the synthesized speech, the system may process the transcript data with a language embedding representing language-dependent speech characteristics of the target language, a speaker embedding representing speaker-dependent voice identity characteristics of a speaker, and a performance embedding representing the vocal performance characteristics of the source speech. The system may control the duration of segments of the synthesized speech to better align with corresponding segments of the source speech for the purpose of dubbing multimedia content with synthesized speech in a language different from that of the original audio.
Owner:AMAZON TECH INC

Methods, apparatuses and computer program products for providing large language models to facilitate live translations

A system and method for translating speech are provided. The system may detect audio signals associated with speech data of a first user(s) and determine the speech data is associated with a first language. The system may translate words of the speech data associated with the first language to other words associated with a second language. The system may present the other words translated in the second language as text items to a display of a device of a second user(s) or output the other words in the second language as audio to the second user(s). The system may generate, in response to the speech data, a first set of replies in the first language and a corresponding second set of other replies in the second language presented via the display. The other replies may include a translation in the second language of the replies in the first language.
Owner:META PLATFORMS INC

Method, device, electronic device and storage medium for speech translation

The application discloses a speech translation method and device, electronic equipment and storage medium, wherein the method comprises: acquiring continuous audio segments of a source language audio stream; for each audio segment in the continuous audio segments, performing feature extraction on each audio segment through a large language model (LLM) to obtain a first feature representation of each audio segment; based on the first feature representation of each audio segment, generating a source language text corresponding to each audio segment through a decoder, and displaying the source language text; the decoder is a decoder independent of the LLM; obtaining a second feature representation of the source language audio stream based on the first feature representation of each audio segment; and based on the second feature representation of the source language audio stream, generating a target language text corresponding to the source language audio stream through the LLM. The application embodiment can improve the real-time performance of ASR recognition result output while taking into account the device power consumption.
Owner:HUAWEI TECH CO LTD

A head-mounted real-time speech translation control method and system based on smart glasses

PendingCN122635381AData controlSmartglasses
The application relates to the technical field of intelligent wearable devices and voice translation, and provides a head-mounted real-time voice translation control method and system based on intelligent glasses. Environmental voice signals in a head-mounted scene are collected through double-channel microphones, and pure voice data corresponding to a target speaker is extracted; the pure voice data is sent into a translation processing link to generate translated audio data and translated text data in a target language; an audio playing module of the intelligent glasses is controlled to output the translated audio data, and a display module of the intelligent glasses is controlled to synchronously display translated subtitle content which is time-aligned with the translated audio data; in the real-time translation operation process, preset language switching instructions in the collected voice signals are continuously monitored, and when an effective switching instruction is identified, the language configuration of the translation processing link is directly adjusted to complete quick switching of a multilingual translation mode. The method can accurately extract target speaker voice in a head-mounted scene, effectively filter out background interference, and improve the accuracy of voice collection.
Owner:SHENZHEN ZHILIAN SHENGYA ELECTRONIC TECH CO LTD

Sound equipment control system for speech translation and sound equipment

The invention relates to the field of speech translation, in particular to a sound equipment control system for speech translation and sound equipment. The sound equipment control system comprises the main microphone module and the auxiliary microphone module, the main microphone module and the auxiliary microphone module can pick up voices, transmit the voices to the translation equipment for processing, and then output the voices through the voice output unit, so that through the cooperation of the main microphone module and the auxiliary microphone module, the pickup range can be expanded, and voice information in the surrounding environment can be collected more clearly; voice data missing or inaccuracy is avoided; the sound equipment comprises the main microphone mechanism and the auxiliary microphone mechanism, the auxiliary microphone mechanism is detachably installed on the main microphone mechanism, the pickup position can be adjusted according to actual requirements, voice information in the surrounding environment can be collected more clearly, and the voice translation quality is improved.
Owner:SHENZHEN TEANA TECH CO LTD

Speech translation method, server, storage medium and program product

The invention provides a speech translation method, a server, a storage medium and a program product. The method relates to the field of artificial intelligence. The method comprises the following steps: acquiring original voice of a to-be-translated original language and a to-be-translated target language; and searching translation knowledge matched with the original voice in a translation knowledge base, inputting the original voice and the translation knowledge matched with the original voice into a voice translation model, translating the original language voice in the translation knowledge appearing in the original voice into a corresponding target language text through the voice translation model, and generating a translation text of the target language. On the basis of a cross-modal translation knowledge base, external knowledge intervention is carried out on the speech translation model, translation errors of customized vocabularies are effectively avoided on the premise that parameters of the speech translation model do not need to be retrained, the translation accuracy of the speech translation model to customized vocabularies such as proper nouns, terminologies and new words is improved, and the translation efficiency is improved. Therefore, the accuracy of the speech translation result is improved.
Owner:ALIBABA (CHINA) CO LTD

Translator (S3)

1. The name of the product of this design: Translator (S3). 2. Application of the product of this design: A translation machine used for voice translation, text translation, simultaneous interpretation, instant audio translation, photo translation, and offline translation. 3. The key point of the design of this product lies in its shape. 4. The picture or photo that best illustrates the design points: three-dimensional picture. 5. Other situations that require explanation: This design product supports more than 138 languages ​​worldwide, covering voice, text, conversation, and simultaneous interpretation. It can be used across third-party social software and has instant translation functions such as audio and video calls, photo translation, and offline translation.
Owner:SHENZHEN CHENKUN TECHNOLOGY CO LTD

Speech translation method, electronic equipment and storage medium

The invention discloses a speech translation method, electronic equipment and a storage medium, and relates to the technical field of speech processing, the method comprises the following steps: converting a speech segment to be translated into a text sequence, and calculating the semantic coherence confidence of the text sequence; the acoustic features and the semantic coherence confidence coefficient of the to-be-translated speech segment are input into a boundary judgment model, candidate sentence boundaries in the to-be-translated speech segment and boundary judgment confidence coefficients corresponding to the candidate sentence boundaries are obtained, and the boundary judgment model is a model of a neural network structure and is obtained through pre-training; and determining a translation result output opportunity of the statement defined by the boundary of the candidate sentence according to the boundary judgment confidence coefficient. According to the method and the device, the problem of poor overall effect of real-time translation caused by inaccurate sentence boundary judgment and high judgment delay in speech translation can be avoided.
Owner:GOERTEK INC

Audio translation method and device based on large language model

The invention discloses a speech translation method and device based on a large language model, and the method comprises the steps: collecting a training corpus, and obtaining an audio and a corresponding text in a training stage, so as to extract the spectrum features of the audio; semantic features and global acoustic features of the spectrum features are extracted and coded; matching the semantic feature codes of the corresponding texts through a large number of translations to train a semantic feature translation large model so as to generate translated semantic feature codes from the semantic feature codes; training a vocoder basic model by using global acoustic feature codes and semantic feature codes of audios corresponding to a large number of texts; and finely adjusting the vocoder basic model through global acoustic feature information and semantic features of a small amount of audio of the target speaker to obtain a vocoder. According to the method, the problems of high processing delay, weak context understanding ability, complex system integration and the like of a traditional speech translation mode are solved, smooth and accurate translation audio can be synthesized even if the corpus of a certain language of a target speaker is few, and the consistency of timbres can be ensured.
Owner:PANOVASIC TECHNOLOGY CO LTD