Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

132 results about "Speech translation" patented technology

Speech translation is the process by which conversational spoken phrases are instantly translated and spoken aloud in a second language. This differs from phrase translation, which is where the system only translates a fixed and finite set of phrases that have been manually entered into the system. Speech translation technology enables speakers of different languages to communicate. It thus is of tremendous value for humankind in terms of science, cross-cultural exchange and global business. so when you type something in you can press a button and it will read it out and pronounce it for you so you know how to say it in a different language.

Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement

The invention relates to the technical field of artificial intelligence, and discloses a Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement, and the method comprises the steps: collecting a multi-channel audio signal through a built-in multi-microphone array of a Bluetooth device, carrying out the dynamic direction self-adaptive beam forming of the multi-channel audio signal, and carrying out the self-adaptive beam forming of the multi-channel audio signal; extracting a Mel spectrogram feature of the direction enhancement signal, identifying lip regions of a plurality of candidate speakers in each frame of real-time speaking video captured by a camera, performing time sequence convolution on the lip regions to obtain a lip movement time sequence embedded vector, calculating a correlation score with the Mel spectrogram feature, separating the direction enhancement signal, and obtaining a lip movement time sequence embedded vector; and performing text transcription and conversion on the high-confidence separation voice to obtain a translation language text, and sending the synthesized target translation voice to a preset mobile terminal through the Bluetooth device to obtain a target translation result. According to the method, the real-time performance and accuracy of speech translation are improved in a multi-person scene, far-field speech, noise interference and accent difference.
Owner:SHENZHEN DIE MICRO SEMICON CO LTD

Speech translation method and device of end cloud translation system, equipment and medium

The invention relates to a speech translation method, device and equipment of an end cloud translation system and a medium, and relates to the technical field of intelligent speech translations, the end cloud translation system comprises an audio receiving end and a cloud, the method is applied to the audio receiving end, and the method comprises the following steps: obtaining a first language audio of the audio receiving end; sending the first language audio to a cloud; and receiving and playing a second language audio, wherein the second language audio is translated by the cloud according to the first language audio and then is sent out. The method can be widely applied to cross-border conference scenes, and the cross-language communication efficiency and experience are improved.
Owner:SHENZHEN XINZHILIAN SOFTWARE CO LTD

Speech translation using a wearable device

A speech translation system may provide real-time or near real-time translation of speech uttered by a person or emitted from a media device. The speech translation system may include a device that may receive audio representing speech in a source language and output audio representing speech in a target language. The speech translation system may translate the speech in portions representing semantically cohesive speech segments such that the target speech reflects the semantic meaning of words, phrases, and / or clauses as used in the context of the source speech. The speech translation system may condense the speech segments prior to or during translation to reduce verbosity. The speech translation system may selectively translate some speakers and not others, and may determine voice characteristics of source speech and apply identifying characteristics to the target speech that allow a user to differentiate respective target speech from different speakers based on the identifying characteristics.
Owner:AMAZON TECH INC

Intelligent speech translation mobile phone and system capable of realizing multi-language inter-translation

The invention belongs to the technical field of mobile terminal communication and language translation, and particularly relates to an intelligent speech translation mobile phone and system capable of realizing multilingual inter-translation, which are characterized in that a two-way speech separation module and unit, a language recognition related module and unit, an end-cloud collaborative translation module and unit and a system are compatible with related module and unit collaboration; two-way voice collection, separation, language recognition and end-cloud collaborative translation are completed in a call scene, the system is compatible with a mainstream mobile operating system and communication application, a resource isolation and process linkage mechanism is adopted to guarantee operation compatibility, multilingual communication obstacles in the call scene are effectively solved, complex environment interference is adapted, the real-time performance and accuracy of translation are considered, and the system is suitable for large-scale popularization and application. And the native function of the terminal does not need to be modified.
Owner:SHENZHEN GUO ELECTRONIC INFORMATION CO LTD

Speech translation method and device

The invention provides a speech translation method and device, and the method comprises the steps: translating source language speech data based on a speech translation model, and obtaining a target language text; a training target of the speech translation model comprises minimizing a difference between a first target language prediction text generated based on source language sample speech data and a translation label corresponding to the source language sample speech data, and minimizing a difference between speech features of the source language sample speech data and text features of a first source language sample text. And minimizing the difference between a second target language prediction text generated based on the pseudo source language speech features and a translation label corresponding to the second source language sample text. According to the method and the device, the problem of scarcity of annotated voice data is solved by efficiently utilizing relatively rich text data in a low-resource language scene, so that the performance of a voice translation model is improved.
Owner:IFLYTEK CO LTD

Speech translation method, system, and storage medium

The application discloses a speech translation method and system and a storage medium, relates to the technical field of speech processing, and comprises the following steps: acquiring a real-time audio stream of source audio, and incrementally generating source language text corresponding to the real-time audio stream; detecting the semantic integrity of the generated source language text in real time; in the case where the semantic integrity is greater than or equal to the integrity threshold value corresponding to a minimum semantic unit, taking the generated source language text as a source language text block; determining target translation corresponding to the source language text block, and outputting the target translation. After the source language text is incrementally generated, the translation of the source language text block and the output are driven based on the semantic integrity, online segmentation translation of long sentences is realized, the translation waiting time is shortened, and the target translation and the original sound are synchronized.
Owner:ZHUHAI MOJIE TECH CO LTD

A crosswise intermodal transport obfuscation method and apparatus based on adaptive optimal transport

This invention relates to the field of speech translation technology, and particularly to a cross-transport obfuscation method and apparatus based on adaptive optimal transmission. The method includes: constructing an optimal transmission model within a multi-task general framework; performing attention-enhanced optimal transmission alignment on speech and text sequences; optimizing the attention-enhanced optimal transmission alignment based on a dynamic window strategy to obtain the alignment relationship between the speech and text sequences; fusing speech and text features of the speech and text sequences through a similarity-based adaptive fusion strategy; and combining a contrastive learning loss function with a multi-task learning framework, constructing positive and negative sample pairs, and combining the multi-task loss functions to obtain a unified optimization objective. This invention effectively reduces the representational differences between speech and text through dynamic window strategies, optimal transmission, and contrastive learning, achieving significant improvements in translation tasks for low-resource languages.
Owner:MINZU UNIVERSITY OF CHINA

Speech-to-speech translation

PendingUS20260154515A1Natural language translationSound input/outputSpeech to speech translationSpeech translation
A speech-to-speech translation method comprises transcribing speech spoken in a source language into transcribed text data in the source language using an on-premises speech recognition model. The transcribed text data is translated into translated text data in a target language using a first on-premises machine translation model. The translated text data is reverse translated into retranslated text data in the source language using a second, different on-premises machine translation model. The transcribed text and the retranslated text are displayed on a screen. The method also involves synthesizing, using an on-premises speech synthesis model, translated speech data in the target language based on the translated text data and play back, in response to a user confirmation, translated speech in the target language based on the translated speech data in the target language.
Owner:MABEL AI AB

Method and system for providing video conference service including artificial intelligence-based speech interpretation function

PCT designated stageWO2026146907A1Speech translationSpeech sound
Disclosed are a method and system for providing a speech conference service including an artificial intelligence-based speech interpretation function. According to one embodiment, the method for providing a video conference service may comprise the steps of: setting two or more interpretation bots participating as virtual participants in a video conference service; interpreting a speech of a first language input by participants of the video conference service into a speech of at least one other language different from the first language through processes of speech recognition, translation, and synthetic speech generation between the two or more interpretation bots and artificial intelligence; and providing the interpreted speech.
Owner:LINE PLUS

Voice real-time translation and proofreading system for cross-border e-commerce

The invention relates to the technical field of voice real-time translation, in particular to a voice real-time translation and proofreading system oriented to cross-border e-commerce, which comprises the following modules: a voice semantic processing module used for receiving voice input, converting the voice input into text data, performing semantic analysis on the text data, extracting commodity names, attribute information and logic relationships, and sending the extracted commodity names, attribute information and logic relationships to a server; semantic structure data is generated and stored in a context cache; and the anaphora analysis module is used for reading the semantic structure data from the context cache, detecting anaphora words in the semantic structure data, determining anaphora objects and updating the semantic data subjected to anaphora resolution to the context cache. According to the method, word segmentation, part-of-speech tagging, syntactic analysis and dependency analysis are carried out on a text, commodity names, the number, colors, sizes and attribute relations are extracted to generate structured semantic data, the structured semantic data are stored in a context cache, and when anaphora words are detected, anaphora objects are determined through semantic similarity and context logic analysis, and the semantic data are updated.
Owner:WUHAN CITY VOCATIONAL COLLEGE

system

We provide a system that enables more natural and accurate speech translation. [Solution] The system includes means for receiving the user's voice and acquiring audio data, means for providing speech recognition technology that analyzes the audio data and converts it into text data, means for motion recognition technology that captures the user's mouth movements and generates motion data, means for transmitting the text data and motion data to a server, means for the server to translate the text data into different languages ​​and convert the text in the different languages ​​into audio data, and means for transmitting the audio data to the user's terminal and playing it back.
Owner:SOFTBANK GROUP CORP

Speech translation with performance characteristics

An expressive speech translation system may process source speech in a source language and output synthesized speech in a target language while retaining vocal performance characteristics such as intonation, emphasis, rhythm, style, and / or emotion. The system may receive a transcript of the source speech, translate it, and generate transcript data. To generate the synthesized speech, the system may process the transcript data with a language embedding representing language-dependent speech characteristics of the target language, a speaker embedding representing speaker-dependent voice identity characteristics of a speaker, and a performance embedding representing the vocal performance characteristics of the source speech. The system may control the duration of segments of the synthesized speech to better align with corresponding segments of the source speech for the purpose of dubbing multimedia content with synthesized speech in a language different from that of the original audio.
Owner:AMAZON TECH INC

Method, device, electronic device and storage medium for speech translation

The application discloses a speech translation method and device, electronic equipment and storage medium, wherein the method comprises: acquiring continuous audio segments of a source language audio stream; for each audio segment in the continuous audio segments, performing feature extraction on each audio segment through a large language model (LLM) to obtain a first feature representation of each audio segment; based on the first feature representation of each audio segment, generating a source language text corresponding to each audio segment through a decoder, and displaying the source language text; the decoder is a decoder independent of the LLM; obtaining a second feature representation of the source language audio stream based on the first feature representation of each audio segment; and based on the second feature representation of the source language audio stream, generating a target language text corresponding to the source language audio stream through the LLM. The application embodiment can improve the real-time performance of ASR recognition result output while taking into account the device power consumption.
Owner:HUAWEI TECH CO LTD

A head-mounted real-time speech translation control method and system based on smart glasses

PendingCN122635381AData controlSmartglasses
The application relates to the technical field of intelligent wearable devices and voice translation, and provides a head-mounted real-time voice translation control method and system based on intelligent glasses. Environmental voice signals in a head-mounted scene are collected through double-channel microphones, and pure voice data corresponding to a target speaker is extracted; the pure voice data is sent into a translation processing link to generate translated audio data and translated text data in a target language; an audio playing module of the intelligent glasses is controlled to output the translated audio data, and a display module of the intelligent glasses is controlled to synchronously display translated subtitle content which is time-aligned with the translated audio data; in the real-time translation operation process, preset language switching instructions in the collected voice signals are continuously monitored, and when an effective switching instruction is identified, the language configuration of the translation processing link is directly adjusted to complete quick switching of a multilingual translation mode. The method can accurately extract target speaker voice in a head-mounted scene, effectively filter out background interference, and improve the accuracy of voice collection.
Owner:SHENZHEN ZHILIAN SHENGYA ELECTRONIC TECH CO LTD

Speech translation method, server, storage medium and program product

The invention provides a speech translation method, a server, a storage medium and a program product. The method relates to the field of artificial intelligence. The method comprises the following steps: acquiring original voice of a to-be-translated original language and a to-be-translated target language; and searching translation knowledge matched with the original voice in a translation knowledge base, inputting the original voice and the translation knowledge matched with the original voice into a voice translation model, translating the original language voice in the translation knowledge appearing in the original voice into a corresponding target language text through the voice translation model, and generating a translation text of the target language. On the basis of a cross-modal translation knowledge base, external knowledge intervention is carried out on the speech translation model, translation errors of customized vocabularies are effectively avoided on the premise that parameters of the speech translation model do not need to be retrained, the translation accuracy of the speech translation model to customized vocabularies such as proper nouns, terminologies and new words is improved, and the translation efficiency is improved. Therefore, the accuracy of the speech translation result is improved.
Owner:ALIBABA (CHINA) CO LTD

Speech translation method, electronic equipment and storage medium

The invention discloses a speech translation method, electronic equipment and a storage medium, and relates to the technical field of speech processing, the method comprises the following steps: converting a speech segment to be translated into a text sequence, and calculating the semantic coherence confidence of the text sequence; the acoustic features and the semantic coherence confidence coefficient of the to-be-translated speech segment are input into a boundary judgment model, candidate sentence boundaries in the to-be-translated speech segment and boundary judgment confidence coefficients corresponding to the candidate sentence boundaries are obtained, and the boundary judgment model is a model of a neural network structure and is obtained through pre-training; and determining a translation result output opportunity of the statement defined by the boundary of the candidate sentence according to the boundary judgment confidence coefficient. According to the method and the device, the problem of poor overall effect of real-time translation caused by inaccurate sentence boundary judgment and high judgment delay in speech translation can be avoided.
Owner:GOERTEK INC

Speech translation method and device, readable storage medium and program product

Embodiments of the invention provide a speech translation method and device, a readable storage medium and a program product. The method comprises the steps of determining a source language corresponding to a to-be-processed speech; generating a translation speech and a translation text of a translation language; the translated speech and the speech to be processed have the same acoustic characteristics; the translation language is determined by a user or a use environment; and displaying the translated speech and the translated text. And the information transmission efficiency in the cross-language interaction scene is improved.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Multimodal LLM that learns to self-correct, with a focus on ASR

Systems and methods are disclosed that are designed for an assistive large language model (LLM) with self-correction. An assistive LLM can operate autoregressively and perform language-related tasks such as automatic speech recognition, speech translation, and / or machine translation.
Owner:GOOGLE LLC

Audio translation system and method based on translation earphone

The audio translation system comprises the translation earphone and an intelligent terminal which communicate with each other, and the intelligent terminal is provided with third-party audio software and audio translation software; after receiving original audio data from the third-party audio software, the translation earphone copies the original audio data and sends the copied audio data to the audio translation software for voice translation, and the intelligent terminal sends the audio data translated by the audio translation software to the translation earphone for playing. According to the system and the method, the audio data of the third-party audio software is copied through the translation earphone and then sent to the audio translation software to be translated, and finally the translated audio data is sent to the translation earphone to be played, so that real-time translation of the audio data of the third-party audio and video playing software is realized; and the translation efficiency is improved. In addition, the system reduces the audio signal transmission sampling rate of the audio data copied by the translation earphone through the sampling rate conversion module, so that the data processing complexity of audio translation software is reduced.
Owner:SHENZHEN TIMEKETTLE TECH CO LTD

Face-translator: end-to-end system for speech-translated lip-synchronized and voice preserving video generation

A neural end-to-end system is provided for the face and voice preserving translation of videos. The system is a pipeline of multiple models that produces a video of the original speaker speaking in the target language with modified lip movement to match the target speech, while preserving emphases and prosody of the original speech, and voice characteristics of the original speaker. The pipeline starts with automatic speech recognition including emphasis detection, followed by the translation model. The translated text is then synthesized by a Text-to-Speech model that recreates the original emphases in the target sentence. The resulting synthetic speech is then converted back to the original speakers' voice using a voice conversion model. Finally, to synchronize the lips of the speaker with the translated audio, a generative model generates frames of adapted lip movements which are combined with the audio to produce the final output. The disclosure further describes several use-cases and configurations that apply these techniques to video conferencing, dubbing, low-bandwidth transmission, speech enhancement and assistive technology for the hearing impaired.
Owner:WAIBEL ALEXANDER

Audio segmentation method and device, electronic equipment and storage medium

The application provides an audio segmentation method and device, electronic equipment and storage medium, wherein the method comprises: obtaining audio to be segmented; extracting acoustic features of each frame in the audio to be segmented, and based on the acoustic features of each frame, performing semantic boundary sequence labeling on the audio to be segmented to obtain semantic boundary labeling results of each frame; and based on the semantic boundary labeling results of each frame, performing segmentation on the audio to be segmented. The method, device, electronic equipment and storage medium provided by the application can assist semantic segmentation based on tone and pause information in the acoustic features of each frame, retain complete semantic information of the audio, and avoid punctuation recognition errors, thereby improving the accuracy and reliability of audio segmentation. Furthermore, the method can be applied to a cascaded speech translation system and an end-to-end speech translation system, thereby expanding the application range of audio segmentation.
Owner:HKUST IFLYTEK (SHANGHAI) TECH CO LTD

A voice translation screen control method and system for interaction

The application relates to a voice translation screen control method and system for interaction, and relates to the field of intelligent interaction technology, which comprises the following steps: acquiring a region detection image; performing feature recognition on the region detection image to determine user lip features, and determining user lip positions according to the user lip features; determining effective sound pickup distances according to the user lip positions and positions of sound pickup devices; determining sound pickup sensitive ranges corresponding to the effective sound pickup distances according to sound pickup matching relationships; randomly selecting a sound pickup sensitive value in each sound pickup sensitive range to define an effective sensitive value, and controlling the sound pickup devices to work at the effective sensitive value to acquire external voice volumes; determining a user representative volume according to the external voice volumes, determining a sensitive adjustment coefficient according to the user representative volume and an effective recognition volume, and adjusting and updating the effective sensitive values according to the sensitive adjustment coefficient. The application has the effect of facilitating subsequent analysis of collected voices.
Owner:NINGBO LANKE INTELLIGENT ENG

Game speech translation method and system

The embodiment of the invention provides a game voice translation method and system, and the method comprises the steps: capturing and segmenting a sound signal detected by a terminal microphone in a game process through human voice detection (VAD), and obtaining a first voice signal; and translating the first voice signal into a first text character string through an artificial intelligence AI translation model, and converting the first text character string into a second voice signal. Through the embodiment of the invention, the problem of difficult communication caused by language impassability in the game process in the related art is solved.
Owner:ZTE CORP

A translation interaction device and electronic kit for use with smart wearable devices

This utility model relates to the field of smart device technology, and more particularly to a translation interaction device and electronic kit for use with smart wearable devices. The translation interaction device includes a main body for use with an external smart wearable device. The main body includes a housing, a main control module disposed within the housing, and a display screen disposed on the surface of the housing. The main control module is communicatively connected to both the display screen and the external smart wearable device. The main control module includes a voice translation module, which includes a language receiving unit, a voice translation unit, and a voice output unit. The main control module also includes one or any combination of a recording module, a text translation module, an audio / video playback module, and a camera module. This utility model, through collaboration with smart wearable devices, can meet users' translation needs in different language environments, thereby expanding the functionality of external smart wearable devices.
Owner:深圳目渡科技有限公司

A method and apparatus for modeling an end-to-end speech translation model based on cross-language CTC

The application relates to an end-to-end speech translation model modeling method and device based on cross-language CTC, and belongs to the technical field of natural language processing; the method solves the problem that the speech translation method in the prior art ignores the guidance of target language text to an encoder and the monotone hypothesis and conditional independent hypothesis problems of CTC; the modeling method comprises the following steps: constructing an initial speech translation model; the initial speech translation model comprises an acoustic encoder, a text encoder and a decoder; obtaining a speech data set; the speech data set comprises source language speech data, source language labeled text corresponding to the speech data and target language labeled text; the initial speech translation model is trained by using the speech data set, is iteratively updated by using a loss function, and the speech translation model is obtained.
Owner:XIAONIU FANYI

Speech translation, model training method and device, equipment and storage medium

ActiveCN114783428BVoice translation process is simpleHigh efficiency of voice translationSpeech recognitionSpeech synthesisSpeech translationSpeech sound
The present disclosure provides a speech translation, a model training method, device and equipment and a storage medium, relates to the technical field of data processing, and particularly relates to the technical field of audio data processing. The specific implementation scheme is as follows: audio features of each audio frame in source language audio to be translated are extracted; based on the audio features of each audio frame, speech units of a target language corresponding to each audio frame are respectively determined as target speech units, wherein each speech unit is audio data of an acoustic category corresponding to the audio; and based on the time sequence order of each audio frame in the source language audio and the target speech units corresponding to each audio frame, target language audio is generated. When the scheme provided by the embodiment of the present disclosure is applied to speech translation, the efficiency of speech translation can be improved.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Voice processing method and device, electronic equipment and storage medium

The invention discloses a voice processing method and device, electronic equipment and a storage medium, and relates to the technical field of voice processing, and the method comprises the steps: extracting a language-independent intermediate representation from an input voice, the language-independent intermediate representation being a feature vector reflecting acoustic details and semantic information of the input voice; and executing the speech recognition task and the speech translation task in parallel according to the language-independent intermediate representation to obtain a source language text and a target language text corresponding to the input speech. According to the invention, the efficiency and accuracy of the multi-voice task processing flow can be improved.
Owner:GOERTEK INC

Automatic transcription-assisted speech translation using language models

Disclosed are apparatuses, systems, and techniques that implement training and deployment of automatic transcription-assisted translation systems that use language models. The techniques include processing, using a first speech-to-text (S2T) model, a first input that includes a speech in a first language to generate a transcription of the speech. The techniques further include processing, using a second S2T model, a second input to generate a translation of the speech to a second language. The second input includes at least a representation of the speech, and the transcription of the speech.
Owner:NVIDIA CORP