System for bidirectional real-time speech translation
The wearable unit with a microphone and smartphone-based translation system provides real-time, simultaneous translation and noise cancellation, addressing the limitations of existing devices by ensuring clear, bidirectional communication with separated language outputs.
Patent Information
- Application Number
- DE202025002838
- Authority / Receiving Office
- DE · DE
- Patent Type
- Utility models
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2026-01-15
- Estimated Expiration
- 2035-09-30
AI Technical Summary
Existing translation devices and apps either require users to wait for their speech to be translated or mix original and translated speech, leading to reduced intelligibility and acoustic confusion, especially in noisy environments, and fail to provide real-time, simultaneous translation of both input and output languages.
A wearable unit with a mouth-mounted microphone, optional camera, and active noise cancellation, combined with a smartphone for real-time speech recognition, translation, and directional loudspeaker output, ensuring only the translated language is audible externally and the user's voice is audible internally, while suppressing external noise.
Enables clear, real-time, bidirectional communication without acoustic overlap, allowing quiet speech capture and natural conversation in noisy conditions, with low cognitive load and effective language separation.
Abstract
Description
Technical field
[0001] The invention relates to portable translation systems, in particular head-worn devices (e.g. headsets and smart glasses) or comparable communication devices that enable virtually delay-free bidirectional communication between speakers with different native languages. State of the art
[0002] Well-known translation headphones and apps already enable live translation via a device (for example, devices based on US patent US 10 1 08612B2). However, no device currently exists that translates input and output language simultaneously and aloud in real time. Instead, there are either headphones that translate spoken language from the environment for the user in real time, or devices and applications that translate and output the user's speech. There are also apps where the user has to wait for their speech to be translated before they can speak again, but these do not translate aloud in real time. Existing solutions also have other disadvantages: The speaker's original speech remains audible and overlays the translation; quiet speech is often poorly captured in noisy environments; moreover, the speaker frequently hears both their own and others' speech, which reduces intelligibility.A technology that completely replaces the original voice and exclusively reproduces the target language is not yet available on the market. Task
[0003] The aim is to provide a portable system that enables natural communication across language barriers without either party needing to hear the foreign language. The speaker should be able to articulate their native language very quietly, without it being audible to others. The recipient should only hear the selected, already translated target language immediately in real time. Responses from the recipient should be relayed back to the speaker exclusively as translations in their own language. Solution
[0004] The invention comprises a wearing unit (for example, a headset or smart glasses) with headphones and active noise cancellation, a mouth-mounted microphone that can also capture very quiet sounds and whispers and is optionally supplemented by a camera with lip-reading support, a connected terminal device (for example, a smartphone) with a signal processing unit (speech recognition, machine translation, speech synthesis), and an outward-facing loudspeaker attached to the microphone. • The user speaks very quietly into the microphone of the wearable unit, which is positioned in front of the speaker's mouth; optionally, a small camera on the microphone captures lip movements to improve recognition when whispering. • The device (for example, a smartphone) recognizes the whispering or lip movements, translates them immediately, and outputs the translation as a spoken voice. • A loudspeaker, preferably directional, positioned on the outside of the microphone, outputs only the target language. The original language remains practically inaudible to the outside. • The receivers of the wearing unit (for example, headphones or bone conduction receivers) suppress the target language when speaking and reproduce only the user's own voice as far as possible in order to avoid acoustic confusion. • When the other person speaks, their speech is captured via the microphone, immediately translated into the user's native language and played back exclusively in their listeners.
[0005] The invention consistently separates external and internal output: outside, only the target language is audible; inside, the user hears only their own language when answering and only the translation when listening. Example of implementation
[0006] A wearing unit (for example, a headset) with active noise cancellation connects wirelessly to a device (such as a smartphone). A microphone positioned close to the mouth is attached to the microphone arm for voice input. A small camera pointed at the mouth can capture lip movements; optionally, infrared illumination assists with detection in low light. The app on the smartphone continuously processes the incoming audio signal and, optionally, the image data. The smartphone immediately converts the speech into a fluent target language tone and outputs it directly through the external speaker on the microphone. Simultaneously, it controls the headphones so that the target language being emitted outwards is actively suppressed, ensuring that only the user's own whisper is audible as much as possible.The microphone picks up the conversation partner's voice, translates it, and plays it back exclusively in the user's native language through the headphones. The desired language can be easily selected; the device automatically switches between speaking and listening depending on the situation, or allows manual switching. Advantages • Anyone can communicate with anyone else, in real time and regardless of origin or existing language skills. • No acoustic overlap between original and translated voices. • Clear communication even with very quiet articulation in noisy environments. • Natural conversational feel, as each side only hears the relevant language. • Low cognitive load for the user, as foreign language components are filtered out in the headphones. Design variants
[0007] The microphone for signal capture can be designed in various ways. In one version, it comprises a microphone module with two separate recording channels, one for the user's voice and the other for the other person's. Another version can contain two microphones, one for the user's voice and the other for the other person's. Alternatively, two independent, separately installed microphones can be used, each capturing the voice of the user and the other person. Instead of or in addition to the microphone, sensors can be used on the neck or head that detect vibrations during speech, thus reliably capturing even very quiet speech. The translation can be performed entirely on the smartphone or partially on the wearable device.The speaker can be designed so that the sound is directed specifically to the conversation partner, instead of radiating throughout the room. The camera can be covered to meet data protection requirements. The device can be configured for a fixed language combination or for multiple languages and is also suitable for use with different conversation partners speaking different languages.
[0008] Alternatively, the wearing unit can be designed as smart glasses. In this version, a camera and microphone on the smart glasses detect whispering, and the translated speech is spoken aloud by a speaker and the smart glasses. The user's own whisper is returned to them via integrated bone conduction headphones. Commercial applicability
[0009] The invention can be used in all areas where people without a common language need to communicate with each other. It can be used worldwide in almost any context – from travel and international business meetings to healthcare, education, administration, industry, security, and everyday conversations.
[0010] By enabling the complete overcoming of language barriers, the invention makes a significant contribution to global understanding. It allows direct communication regardless of origin and language skills, thus representing a major advancement in interpersonal understanding – both on the small scale of personal conversations and on the large scale of international cooperation. QUOTES INCLUDED IN THE DESCRIPTION
[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature
[0000] US 10 1 08612B2
[0002]
Claims
[1] System for bidirectional real-time speech translation, comprising ◯ a head-worn carrying unit with receivers (headphones or bone conduction receivers) and active noise cancellation, ◯ a microphone positioned close to the mouth for capturing speech at low volume levels, including whispers. ◯ an outward-facing loudspeaker for speech output, arranged on the carrying unit, and ◯ a terminal device connected to the wearer with a signal processing unit (speech recognition, machine translation, speech synthesis), characterized by , that the carrying unit is designed in such a way that: - when the user speaks very quietly or whispers, the speech signal captured by the microphone is immediately translated into a target language and output exclusively via the outward-facing speaker, making the user's original speech practically imperceptible to the outside, and - When a conversation partner makes a voice input, their speech is immediately translated into the user's native language via the microphone and played back exclusively through the user's headphones, whereby the user's headphones actively suppress the outwardly transmitted target language and optionally play only their own voice to the user as far as technically possible in order to avoid acoustic confusion. [2] System according to claim 1, characterized by , that it includes a camera pointed at the user's mouth to support speech recognition through lip reading. [3] System according to any one of the preceding claims, characterized by that the camera pointing towards the mouth is covered to protect privacy when not in use. [4] System according to any one of the preceding claims, characterized by , that additional sensors are provided on the user's neck or head to detect vibrations generated during speech in order to reliably record even very quiet speech. [5] System according to any one of the preceding claims, characterized by , that the wearing unit is designed as smart glasses, in which a microphone and a camera capture the user's whisper, the translation is output to the outside via a speaker integrated into the smart glasses and the user's own whisper is returned to them via integrated bone conduction headphones. [6] System according to any one of the preceding claims, characterized by, that the translation processing is optionally carried out entirely on the connected terminal device or at least partially in the carrying unit. [7] System according to any one of the preceding claims, characterized by , that the loudspeaker arranged at the microphone is designed as a directional loudspeaker, so that the emitted sound essentially only reaches the conversation partner and does not radiate into the entire room. [8] System according to any one of the preceding claims, characterized by that the system can be configured either for a fixed language combination or for several different languages. [9] System according to any one of the preceding claims, characterized by , that it can also be used in conversations with multiple interlocutors with different native languages, whereby the language of each interlocutor is translated into the user's native language and vice versa.
Citation Information
Patent Citations
Mobile device having human language translation capability with positional feedback
US10108612B2