Multi-modal translation equipment based on scene perception and transparent display
By using scene-aware and transparent multimodal translation devices, environmental image information is used to assist semantic understanding, solving the problem of translation errors of polysemous words and realizing natural eye-tracking interaction and efficient communication.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUHAN CITY VOCATIONAL COLLEGE
- Filing Date
- 2026-01-08
- Publication Date
- 2026-05-12
AI Technical Summary
Existing translation devices lack multimodal environment awareness, leading to translation errors of polysemous words in different contexts, and the interaction method affects natural communication.
By combining natural language processing and computer vision technologies, using environmental image information to assist semantic understanding, employing transparent display technology for bidirectional interaction, identifying scene features through an image acquisition device, using a neural machine translation model for contextual constraints, and combining transparent display units to achieve natural eye-tracking interaction.
It improves the contextual adaptability and naturalness of the translation, solves the problem of translation errors of polysemous words, and maintains the naturalness and convenience of eye contact with users.
Smart Images

Figure CN122021670A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of natural language processing technology, and more specifically, to a multimodal translation device based on scene awareness and transparent display. Background Technology
[0002] With the acceleration of global economic integration, cross-language communication is becoming increasingly frequent. To overcome language barriers, various translation aids (such as smartphone translation apps, handheld translators, and translation headsets) have been widely used. Existing mainstream translation devices typically employ a technical approach of "ASR (Automatic Speech Recognition) + Machine Translation (MT) + Text-to-Speech (TTS)," which involves capturing speech through a microphone, converting it into source language text, translating it into target language text, and then playing or displaying it.
[0003] Although existing technologies have met basic communication needs to a certain extent, the following significant technical shortcomings still exist in practical use: 1. Lack of multimodal environment awareness: Traditional translation devices rely solely on voice input and cannot perceive the physical context in which the dialogue takes place, leading to translation errors of polysemous words in different contexts. For example, the English word "Check" means "pay the bill" in a restaurant context, "check-in / check-out" in a hotel context, and "examination" in a medical context. Its meaning varies significantly in different contexts, and a single voice input cannot accurately determine its semantics.
[0004] 2. Interaction methods affect natural communication: Most existing translation devices use opaque screens to display translation results. Users need to look down or shift their gaze when viewing the translation, which interrupts eye contact between the two parties and reduces the naturalness and trust in the communication.
[0005] Therefore, this application proposes a bidirectional translation device that can integrate environmental visual information for semantic understanding and support natural eye-to-eye interaction. Summary of the Invention
[0006] In view of this, the purpose of this invention is to provide a multimodal translation device based on scene awareness and transparent display, which combines natural language processing (NLP) and computer vision (CV) technologies, enabling the translation device to use environmental image information to assist semantic understanding, thereby improving the contextual adaptability and reliability of translation. In addition, by adopting transparent display technology for two-way interaction, it helps to reconstruct a natural face-to-face communication experience and effectively improve the naturalness of communication.
[0007] To achieve the above objectives, the present invention provides a multimodal translation device based on scene awareness and transparent display, characterized in that it includes a base, a transparent display unit installed on the top of the base, two image acquisition units symmetrically installed on both sides of the base along its length, two voice interaction modules symmetrically installed on both sides of the base along its width, and a central processing unit installed inside the base. The base includes a bottom shell and a top cover mounted on top of the bottom shell. The top cover includes a flat plate located above the bottom shell and an inclined panel disposed outside the flat plate and connected to the top edge of the bottom shell. The transparent display unit is vertically mounted on the top surface of the tablet. The transparent display unit is transparent on both sides, allowing the line of sight to pass through and supporting the display of text information. The image acquisition device is installed on both sides of the inclined panel along its length. The image acquisition device is used to acquire environmental image data around the device or to perform face tracking. The voice interaction module is used to collect voice signals and play audio. The central processing unit is installed inside the base and is electrically connected to the transparent display unit, the image acquisition unit, and the voice interaction module; a neural machine translation model is provided inside the central processing unit. The central processing unit is configured to perform the following steps: The system receives environmental image data from the image acquisition device, identifies key object features in the image, and determines the domain weight coefficient of the current dialogue scene based on the identification results. Receive the voice signal from the voice interaction module and convert it into source text; The source text is input into the neural machine translation model, and the domain weight coefficients are introduced as contextual constraints to generate the target text; The transparent display unit is driven to display the target text on the side facing the receiver.
[0008] Furthermore, when determining the domain weight coefficients, the central processing unit specifically performs the following steps: The identified key object features are matched with a pre-stored industry feature library, which includes at least one of the following scene tags: catering, medical, business meetings, transportation, and daily social interaction. If a specific industry characteristic is matched, the probability of selecting the corresponding term in the translation model for that industry is increased.
[0009] Furthermore, the multimodal translation device based on scene awareness and transparent display also includes a manually corrected interactive interface; When a user negates the current domain weight coefficient by touching the transparent display unit or using a gesture command, the central processing unit switches to general translation mode, or manually locks to another specific domain weight coefficient according to the user's selection.
[0010] Furthermore, the central processing unit uses beamforming technology to distinguish the direction of the sound source of the speech signal, automatically determines the identity of the current speaker, and controls the transparent display unit to display the target text on the side facing away from the speaker.
[0011] Furthermore, when the transparent display unit displays the target text to the receiver, the central processing unit is also configured to: Based on the image data acquired by the image acquisition device, the facial positions of both parties in the conversation are tracked in real time, and the display coordinates of the target text on the transparent display unit are dynamically adjusted so that the text display position and the speaker's facial position maintain a preset relative spatial relationship.
[0012] Furthermore, the multimodal translation device based on scene awareness and transparent display also includes an attitude sensor, which is installed inside the base and used to detect changes in the physical attitude of the device; The central processing unit is configured to automatically adjust the display orientation of the text on the transparent display unit based on the monitoring data of the attitude sensor when the attitude sensor detects that the base has flipped or rotated, so as to ensure that the text is always facing the observer.
[0013] Furthermore, the multimodal translation device based on scene awareness and transparent display also includes a physical interaction interface, which includes a power button, a connection port, a switching button, and a display control button. The power button is used to control the start and stop of the device, the connection port is used to connect to external devices, the switching button is used to control mode switching to switch the current domain weight coefficient to another specific domain weight coefficient, and the display control button is used to switch the transparent display unit between transparent and opaque states.
[0014] Furthermore, the voice interaction module includes two microphone arrays symmetrically mounted on both sides of the width direction of the inclined panel and four speakers symmetrically mounted on both sides of the width direction of the bottom housing, with the two speakers located on the same side of the width direction of the bottom housing symmetrically arranged with the transparent display unit as the boundary.
[0015] Furthermore, the speaker is used for auxiliary voice interaction and system control, or for playing translated audio in conjunction with the transparent display unit.
[0016] Compared with the prior art, the present invention has the following advantages and effects: 1. The multimodal translation device based on scene awareness and transparent display in this invention can introduce visual modalities (environmental images) into the translation device using an image acquisition device, so as to determine the domain weight coefficient of the current dialogue scene based on the key features in the image information. Thus, the domain weight coefficient is used as a contextual constraint during translation to select the interpretation that conforms to the current scene, thereby achieving the effect of assisting semantic understanding through visual modalities, effectively improving the contextual adaptability of translation, solving the problem of polysemy, and enhancing the practicality and effectiveness of the device.
[0017] 2. The multimodal translation device based on scene awareness and transparent display in this invention eliminates the obstruction of the view by the physical screen by adopting a transparent display unit. Users can maintain eye contact with the other party while reading the translated subtitles, and capture the other party's facial expressions and body language, which greatly improves the naturalness of communication.
[0018] 3. The multimodal translation device based on scene awareness and transparent display in this invention combines face tracking and directional display technology, so that the text always floats in the most comfortable reading area, without requiring the user to manually adjust the device angle, thus improving the convenience and technological feel of use. Attached Figure Description
[0019] Figure 1 This is a three-dimensional structural diagram of a multimodal translation device based on scene awareness and transparent display in an embodiment of the present invention; Figure 2 This is a three-dimensional structural diagram of the other side of the multimodal translation device based on scene awareness and transparent display in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the working principle of a multimodal translation device based on scene awareness and transparent display in an embodiment of the present invention.
[0020] Explanation of reference numerals in the attached figures: 1-Base; 11-Bottom shell; 12-Top cover; 121-Flat plate; 122-Slanted panel; 2-Transparent display unit; 3-Image acquisition device; 4-Central Processing Unit; 5-Voice interaction module; 51-Microphone array; 52-Speaker; 6-Physical interface; 61-Power button; 62-Toggle button; 63-Toggle button; 64-Display control button. Detailed Implementation
[0021] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can also refer to the internal connection of two components; and they can refer to a wireless connection or a wired connection. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0023] Please see Figure 1-3 As shown, this embodiment of the invention provides a multimodal translation device based on scene perception and transparent display, including a base 1, a transparent display unit 2, an image acquisition unit 3, a voice interaction module 5, and a central processing unit 4. The transparent display unit 2 is installed on the top of the base 1, two image acquisition units 3 are symmetrically installed on both sides of the base 1 in the length direction, two voice interaction modules 5 are symmetrically installed on both sides of the base 1 in the width direction, and the central processing unit 4 is installed inside the base 1.
[0024] The base 1 includes a bottom shell 11 and a top cover 12. The top cover 12 is mounted on the top of the bottom shell 11. The top cover 12 includes a flat plate 121 located above the bottom shell 11 and an inclined plate 122 disposed outside the flat plate 121 and connected to the top edge of the bottom shell 11.
[0025] The transparent display unit 2 is vertically mounted on the top surface of the flat panel 121. The transparent display unit 2 is transparent on both sides, allowing the line of sight to pass through and supporting the display of text information. The image acquisition unit 3 is mounted on both sides of the inclined panel 122 along its length. The image acquisition unit 3 is used to acquire environmental image data around the device or to perform face tracking. The voice interaction module 5 is used to acquire voice signals and play audio.
[0026] As a further description of the face tracking function of the image acquisition device 3 in the above scheme, when the user on one side of the device moves his head left or right, the central processing unit 4 will calculate the projection coordinates of the user's face on the screen based on the monitoring data of the image acquisition device 3, and adjust the display position of the subtitles in real time so that the subtitles are always kept 5cm below the user's face, so as to avoid the subtitles obscuring the user's face and make it easier for the user on the other side of the device to observe the user's expression.
[0027] As a preferred option of the above solution, the transparent display unit 2 adopts a double-sided transparent OLED or Micro-LED display panel. The light transmittance of the transparent display unit 2 in the area where no pixels are displayed is greater than 40%, which can ensure that users located on both sides of the device can clearly see each other's facial expressions through the screen, ensuring a balance between high definition and transparency.
[0028] As a further description of the above solution, mounting the image acquisition unit 3 on the inclined panel 122 of the base 1 helps to enable the two image acquisition units 3 to perform face tracking of the users on both sides of the device when the two sides of the transparent display unit 2 are facing the users on both sides respectively. Furthermore, mounting the image acquisition unit 3 on the inclined panel 122 with a certain degree of inclination also helps to further increase the acquisition area of the image acquisition unit 3.
[0029] As a preferred embodiment of the above solution, the image acquisition device 3 is a wide-angle camera, which can cover the desktop environment and part of the background environment around the device during the acquisition of environmental image data, while meeting the face tracking requirements during use, thereby obtaining environmental image data containing object information.
[0030] The central processing unit 4 is installed inside the base 1, and the central processing unit 4 is electrically connected to the transparent display unit 2, the image acquisition unit 3, and the voice interaction module 5. The central processing unit 4 is equipped with a neural machine translation model, which is a machine translation method that directly maps source language sentences to target language sentences.
[0031] Central processing unit 4 is configured to perform the following steps: It receives environmental image data from image acquisition device 3, identifies key object features in the image, and determines the domain weight coefficient of the current dialogue scene based on the identification results. Receive the voice signal from the voice interaction module 5 and convert it into source text; The source text is input into the neural machine translation model, and domain weight coefficients are introduced as contextual constraints to generate the target text. The transparent display unit 2 is driven to display the target text on the side facing the receiver.
[0032] As a preferred embodiment of the above scheme, the central processing unit 4 stores a computer vision algorithm. A computer vision algorithm is a technology that enables computers to understand and interpret visual information in images or videos, so that the central processing unit 4 can use the computer vision algorithm to identify key objects in the image (such as menus, medical devices, office supplies, etc.) and generate the domain weight coefficient of the current dialogue scene accordingly.
[0033] Specifically, when determining the domain weight coefficients, the central processing unit 4 performs the following steps: The identified key object features are matched with a pre-stored industry feature library, which contains at least one of the following scenario tags: catering, medical, business meetings, transportation, and daily social interaction. If a specific industry characteristic is matched, the probability of selecting the corresponding term in the translation model for that industry is increased.
[0034] Specifically, this multimodal translation device based on scene awareness and transparent display also includes a manually corrected interactive interface; When a user negates the current domain weight coefficient by touching the transparent display unit 2 or by using a gesture command, the central processing unit 4 switches to the general translation mode, or manually locks to another specific domain weight coefficient according to the user's selection.
[0035] Specifically, the central processing unit 4 uses beamforming technology to distinguish the direction of the sound source of the speech signal, automatically determines the identity of the current speaker, and controls the transparent display unit 2 to display the target text on the side facing away from the speaker.
[0036] Specifically, when the transparent display unit 2 displays the target text to the receiver, the central processing unit 4 is also configured to: Based on the image data acquired by the image acquisition device 3, the facial positions of both parties in the dialogue are tracked in real time, and the display coordinates of the target text on the transparent display unit 2 are dynamically adjusted so that the text display position and the speaker's facial position maintain a preset relative spatial relationship.
[0037] Specifically, the scene-aware and transparent display-based multimodal translation device also includes an attitude sensor, which is installed inside the base 1 and used to detect the physical attitude of the device.
[0038] The central processing unit 4 is configured to automatically adjust the display direction of the text on the transparent display unit 2 based on the monitoring data of the attitude sensor when the attitude sensor detects that the base 1 has flipped or rotated, so as to ensure that the text is always facing the observer.
[0039] Please see Figure 1-2 As shown, the multimodal translation device based on scene awareness and transparent display also includes a physical interaction interface 6. The physical interaction interface 6 includes a power button 61, a connection port 62, a switching button 63, and a display control button 64. The power button 61 is used to control the start and stop of the device. The connection port 62 is used to connect external devices to facilitate charging, data interaction, system maintenance, and other operations. The switching button 63 is used to control mode switching to switch the current domain weight coefficient to another specific domain weight coefficient. The display control button 64 is used to switch the transparent display unit 2 between transparent and opaque states.
[0040] As a preferred embodiment of the above scheme, the power button 61 and the connection port 62 in the physical interaction interface 6 are located on one side of the width direction of the base 1, and the switch button 63 and the display control button 64 are located on the other side of the width direction of the base 1.
[0041] As a further description of the above solution, the user can issue a privacy command by issuing a voice command or manually pressing the display control key 64 to switch the transparent display unit 2 between transparent and opaque states. When the transparent display unit 2 is switched to the opaque state, the transparent display unit 2 can block the line of sight and meet the needs of temporary private communication.
[0042] Please see Figure 1-2 As shown, the voice interaction module 5 includes two microphone arrays 51 symmetrically mounted on both sides of the width direction of the inclined panel 122 and four speakers 52 symmetrically mounted on both sides of the width direction of the bottom housing 11. The two speakers 52 located on the same side of the width direction of the bottom housing 11 are symmetrically arranged with the transparent display unit 2 as the boundary.
[0043] As a preferred embodiment of the above scheme, the microphone array 51 is composed of several microphones arranged in a linear manner. The central processing unit 4 stores a beamforming algorithm. Beamforming is a technique that processes the signal received by the microphone array 51 to form a spatial filter pointing in a specific direction. It can enhance the sound from a specific direction and suppress noise and interference from other directions, so as to facilitate the directional acquisition of voice signals from users on both sides of the device.
[0044] As a further description of the above solution, the voice interaction module 5 can use beamforming technology to automatically determine the direction of the sound source and control the text to be displayed directionally on the corresponding side of the transparent display unit.
[0045] Specifically, speaker 52 is used for auxiliary voice interaction and system control, or to play translated audio in conjunction with transparent display unit 2.
[0046] It should be noted that the central processing unit 4 of this application also stores a lightweight object detection algorithm, an ASR engine, and a machine translation (MT) engine. The lightweight object detection algorithm is a type of object detection model that can achieve fast inference and low power consumption while maintaining relatively high accuracy. The ASR engine is used to automatically and accurately convert human speech into text. The machine translation engine is a technical system that uses computer software to automatically translate text or speech from one natural language into another natural language.
[0047] The workflow of this scene-aware and transparent display-based multimodal translation device is as follows: Step S1: Environmental perception, the image acquisition module 3 periodically (e.g., 1 frame per second) acquires environmental images; Step S2: Scene modeling. The central processing unit 4 runs a lightweight object detection algorithm (such as YOLOv8-Nano) to identify key objects in the image. Example: If the system recognizes "stethoscope", "white coat", or "medicine packaging", it will generate a feature vector for the [medical scenario]; if it recognizes "stethoscope", "menu", or "plate", it will generate a feature vector for the [dining scenario]. Step S3: Weight generation. The system retrieves the corresponding domain weight coefficients from the preset industry terminology library based on the feature vector. Step S4, Semantic Fusion Translation: When the microphone captures speech, the ASR engine transcribes it into source text. Then, the machine translation engine introduces the domain weight coefficient during the translation process to perform probability re-ranking on polysemous words in the source text, and selects the interpretation that best fits the current scenario. Step S5: Transparent display. The target text generated by the translation is transmitted to the transparent display unit 2. The system determines the speaker based on the direction of the sound source and displays the translation on one side of the screen in the direction of the listener's line of sight. The font color and size can be adaptively adjusted according to the loudness of the tone.
[0048] Example 1: A specific application example of this scene-aware and transparent display-based multimodal translation device in a Western restaurant; 1. Scene initialization: User A (English user) and User B (Chinese user) are seated face to face across the device of this invention; the device's rear-facing camera captures the menu and cutlery placed on the table.
[0049] 2. Context Locking: After the central processing unit identifies the above objects, it automatically locks the translation model to "dining / tourism mode".
[0050] Conversation: After finishing their meal, User A asked the waiter, "Can I have the check?" Technical Comparison: Traditional translation machines, lacking visual information, might translate "check" as "May I check it?" or "May I have that check?" based on a general corpus, causing confusion for the waiter.
[0051] The device of this invention: Since the "dining mode" weight has been loaded, the translation engine knows that the high-frequency meaning of "check" in the restaurant environment is "Bill"; therefore, the system outputs the translation: "Can I pay the bill?" 5. Interactive Experience: The Chinese translation of the sentence is displayed on a transparent screen, floating next to User A's face; while User B (the waiter) reads the subtitle, he can clearly see User A's waving gesture and friendly facial expression through the screen, and the two parties complete a natural and smooth cross-language communication.
[0052] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of this disclosure, and all such changes and modifications will fall within the scope of protection of this invention.
Claims
1. A multimodal translation device based on scene awareness and transparent display, characterized in that, include The base (1), the transparent display unit (2) installed on the top of the base (1), the two image acquisition units (3) symmetrically installed on both sides of the base (1) in the length direction, the two voice interaction modules (5) symmetrically installed on both sides of the base (1) in the width direction, and the central processing unit (4) installed inside the base (1). The base (1) includes a bottom shell (11) and a top cover (12) mounted on the top of the bottom shell (11). The top cover (12) includes a flat plate (121) located above the bottom shell (11) and a sloping panel (122) disposed outside the flat plate (121) and connected to the top edge of the bottom shell (11). The transparent display unit (2) is vertically mounted on the top surface of the flat plate (121). The transparent display unit (2) is transparent on both sides, allowing the line of sight to pass through and supporting the display of text information. The image acquisition device (3) is installed on both sides of the inclined panel (122) along its length. The image acquisition device (3) is used to acquire environmental image data around the device or to perform face tracking. The voice interaction module (5) is used to collect voice signals and play audio; The central processing unit (4) is installed inside the base (1), and the central processing unit (4) is electrically connected to the transparent display unit (2), the image acquisition unit (3), and the voice interaction module (5); a neural machine translation model is provided inside the central processing unit (4); The central processing unit (4) is configured to perform the following steps: Receive environmental image data from the image acquisition device (3), identify key object features in the image, and determine the domain weight coefficient of the current dialogue scene based on the identification results; Receive the voice signal from the voice interaction module (5) and convert it into source text; The source text is input into the neural machine translation model, and the domain weight coefficients are introduced as contextual constraints to generate the target text; The transparent display unit (2) is driven to display the target text on the side facing the receiver.
2. The multimodal translation device based on scene awareness and transparent display according to claim 1, characterized in that, When determining the domain weight coefficients, the central processing unit (4) specifically performs the following steps: The identified key object features are matched with a pre-stored industry feature library, which includes at least one of the following scene tags: catering, medical, business meetings, transportation, and daily social interaction. If a specific industry characteristic is matched, the probability of selecting the corresponding term in the translation model for that industry is increased.
3. The multimodal translation device based on scene awareness and transparent display according to claim 1, characterized in that, This also includes manual modification of the user interface; When a user negates the current domain weight coefficient by touching the transparent display unit (2) or by gesture command, the central processing unit (4) switches to the general translation mode, or manually locks to another specific domain weight coefficient according to the user's selection.
4. The multimodal translation device based on scene awareness and transparent display according to claim 1, characterized in that, The central processing unit (4) uses beamforming technology to distinguish the direction of the sound source of the speech signal, automatically determines the identity of the current speaker, and controls the transparent display unit (2) to display the target text on the side facing away from the speaker.
5. The multimodal translation device based on scene awareness and transparent display according to claim 1, characterized in that, When the transparent display unit (2) displays the target text to the receiver, the central processing unit (4) is further configured to: Based on the image data acquired by the image acquisition device (3), the facial positions of both parties in the dialogue are tracked in real time, and the display coordinates of the target text on the transparent display unit (2) are dynamically adjusted so that the text display position and the speaker's facial position maintain a preset relative spatial relationship.
6. The multimodal translation device based on scene awareness and transparent display according to claim 1, characterized in that, It also includes an attitude sensor, which is installed inside the base (1) and is used to detect changes in the physical attitude of the device; The central processing unit (4) is configured to automatically adjust the display direction of the text on the transparent display unit (2) according to the monitoring data of the posture sensor when the posture sensor detects that the base (1) is flipped or rotated, so as to ensure that the text is always facing the observer.
7. The multimodal translation device based on scene awareness and transparent display according to claim 1, characterized in that, It also includes a physical interaction interface (6), which includes a power button (61), a connection port (62), a switch button (63), and a display control button (64). The power button (61) is used to control the start and stop of the device. The connection port (62) is used to connect to external devices. The switch button (63) is used to control the mode switching to switch the current domain weight coefficient to another specific domain weight coefficient. The display control button (64) is used to switch the transparent display unit (2) between transparent and opaque states.
8. The multimodal translation device based on scene awareness and transparent display according to claim 1, characterized in that, The voice interaction module (5) includes two microphone arrays (51) symmetrically installed on both sides of the width direction of the inclined panel (122) and four speakers (52) symmetrically installed on both sides of the width direction of the bottom housing (11). The two speakers (52) located on the same side of the width direction of the bottom housing (11) are symmetrically arranged with the transparent display unit (2) as the boundary.
9. The multimodal translation device based on scene awareness and transparent display according to claim 8, characterized in that, The speaker (52) is used for auxiliary voice interaction and system control, or for playing translated audio in conjunction with the transparent display unit (2).