Instant intelligent translation method and translation machine based on clinical doctor-patient double-screen translation machine
By addressing the technical challenges faced by healthcare professionals in a doctor-patient setting, the system has implemented advanced technologies and methods for real-time intelligent translation between the doctor and patient screens. It has also enabled simultaneous editing and control on the healthcare side and focused, easy-to-use presentation on the patient side. Furthermore, it has resolved issues of terminological ambiguity, unit misunderstanding, and contextual breaks, and achieved precise alignment and stable presentation of cross-modal content on the same timeline.
Patent Information
- Application Number
- CN202511278280.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-09-09
AI Technical Summary
Existing translation technologies suffer from terminological ambiguity, unit misunderstanding, or contextual breaks in medical-patient scenarios. They struggle to simultaneously meet the differentiated needs of medical staff for editing and control while patients require focused and easy-to-read presentation. Furthermore, they lack precise alignment of the content and audio on both screens along the same timeline, making frame-level synchronization of the translation, broadcast, and the interpreted material difficult. This can lead to flickering, skipping, and audio-visual asynchrony when doctors correct text or replace materials.
The real-time intelligent translation method based on the clinical doctor-patient dual-screen translator initializes the session identifier, target language and display synchronization parameters in the translator host to achieve unified arrangement of text, audio and visualization materials, automatically selects the translation direction, calls up multilingual translation engines and terminology resources for the medical field, performs terminology alignment and context disambiguation, and synchronously presents mirrored visualization materials on the patient screen, maintaining the continuity of the session time sequence and playback queue.
It achieves synchronized control and aligned display effects across multiple terminals, resolves issues of terminological ambiguity, unit misunderstanding, and contextual breaks, and enables medical staff to experience efficient, editable, and controllable synchronized data.
Smart Images

Figure CN120764564B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of machine translation, semantic processing, and more particularly to an instant intelligent translation method and translation machine based on a clinical doctor-patient dual-screen translation machine. Background Technology
[0002] In doctor-patient scenarios, communication between healthcare professionals and non-native speakers typically relies on general translation software or on-site human translators. General translation software has limited capabilities in handling medical terminology, abbreviations, and dosage units, and lacks contextual understanding of electronic medical records, examination reports, etc., easily leading to ambiguity in terminology, misunderstanding of units, or gaps in context. While human translation can alleviate terminology issues to some extent, it is costly, has poor usability, and is detrimental to privacy protection and standardized record keeping. On the other hand, clinical communication involves not only speech but also the interpretation of multimodal materials such as images, charts, and medication instructions. Existing translation devices struggle to simultaneously meet the differentiated needs of both healthcare professionals for editing and control, and patients for focused and readable presentation; they lack precise alignment of content and audio on both screens on the same timeline, failing to achieve frame-level synchronization between the translation, broadcast, and the interpreted material. This can cause flickering, skipping, and audio-visual asynchrony when doctors correct text or change materials.
[0003] In summary, existing translation technologies suffer from the following technical problems: ambiguous terminology, misunderstanding of units, or fragmentation of context; difficulty in simultaneously meeting the differentiated needs of medical staff for editing and control and patients for focused and easy-to-read presentation; lack of precise alignment of content and audio on both screens on the same timeline, making it impossible to achieve frame-level synchronization between the translation, broadcast, and the interpreted material, which can easily lead to flickering, skipping, and audio-visual asynchrony when doctors correct text or change materials. Summary of the Invention
[0004] To address the shortcomings of the existing technologies, this invention provides an instant intelligent translation method and translation machine based on a clinical doctor-patient dual-screen translation device. This addresses issues such as multi-terminal conversation mismatch, cross-modal content asynchrony, inaccurate translation of medical terminology, and difficulties in visualization and sharing, thereby improving the accuracy and efficiency of clinical doctor-patient communication.
[0005] In a first aspect, the present invention provides an instant intelligent translation method based on a clinical doctor-patient dual-screen translation device, comprising:
[0006] In a clinical diagnosis and treatment scenario, a diagnosis and treatment session is established, and the medical staff screen, the patient screen and the translation machine host are grouped and bound together. The session identifier, target language and display synchronization parameters are initialized in the translation machine host, and the current patient identifier is written into the session context. During the session, the translation machine host maintains the synchronization control relationship and time axis number between the two screens and the host, which is used to drive the unified arrangement of text, audio and visualization materials.
[0007] During the conversation, the translator determines the speaker's location and role, automatically selects the translation direction, and enters medical mode when a medical staff member is detected speaking. The translator's microphone channel picks up the medical staff member's speech, displays the source language transcription on the medical staff screen, and simultaneously displays the corresponding translation on the patient screen and plays it back. When a patient member is detected speaking, the translator enters patient mode, displays the patient member's speech through the microphone channel, displays the source language transcription on the patient screen, and simultaneously displays the corresponding translation on the medical staff screen and plays it back. During mode switching, the conversation time sequence and playback queue are kept continuous, and the current speaking role and translation direction are clearly marked on both screens.
[0008] During the translation process, the translation machine host calls upon a multilingual translation engine and terminology resources for the medical field. Based on the current patient's medical records, examination reports, and shared content selected by the medical staff, it performs terminology alignment and context disambiguation. The medical staff screen initiates the sharing control of visual materials, establishes timestamps or fragment ID anchors between the selected visual materials and corresponding sentence fragments, and presents them synchronously on the patient screen in a mirror manner, so that the two screens complete the aligned display of the source text, translation, and visual materials on the same timeline.
[0009] Secondly, the present invention provides a translator that uses the above-mentioned real-time intelligent translation method based on a clinical doctor-patient dual-screen translator to perform real-time intelligent translation in doctor-patient scenarios.
[0010] Compared with the prior art, the beneficial effects of this invention are as follows:
[0011] This invention provides a real-time intelligent translation method and translator based on a clinical dual-screen translator for doctors and patients. The method includes: establishing a clinical consultation session in a clinical setting; binding the medical staff screen, the patient screen, and the translator host together; initializing the session identifier, target language, and display synchronization parameters in the translator host; and writing the current patient identifier into the session context. During the session, the translator host maintains the synchronization control relationship and timeline number between the two screens and the host to drive the unified arrangement of text, audio, and visual materials. When medical staff speech is detected, the system enters medical staff mode, where the translator host's audio pickup channel captures the medical staff's speech, presents the source language transcription on the medical staff screen, and simultaneously presents the corresponding translation on the patient screen with synchronized audio playback. When patient speech is detected, the system enters patient mode. The system uses the microphone channel of the translator to collect the patient's speech, presenting the transcribed source language on the patient's screen and simultaneously displaying the corresponding translation on the medical staff's screen with synchronized audio playback. During mode switching, the conversation time sequence and playback queue are kept continuous, and the current speaker role and translation direction are clearly marked on both screens. During the translation process, the translator calls upon a multilingual translation engine and terminology resources for the medical field, performing terminology alignment and context disambiguation based on the current patient's medical records, examination reports, and shared content selected by the medical staff. The medical staff's screen initiates the sharing control of visual materials, establishing timestamps or segment ID anchors between the selected visual materials and corresponding sentence fragments, and presenting them synchronously on the patient's screen in a mirror manner, so that the two screens complete the aligned display of source text, translation, and visual materials on the same timeline.This invention addresses the issues of lack of unified control and session mismatch among multiple terminals. It achieves individualized parameter loading and end-to-end consistent orchestration centered on session identifiers, resolves the asynchrony between text, audio, and video, enables precise alignment and stable presentation of cross-modal content on the same timeline, solves the problems of unclear speech attribution and translation errors in complex acoustic environments, achieves low-latency role determination and correct translation routing for both medical staff and patients, addresses the issue of delayed patient comprehension, enables dual-screen collaborative display of medical staff's statements and real-time audio / readability for patients, resolves the issue of doctors' inability to understand patients' statements in real time, enables real-time translation and feedback of patient statements and synchronized grasp of key points by medical staff, resolves content loss, repetitive broadcasts, and role confusion during switching, and achieves smooth transitions and dialogue across rounds. By clarifying the boundaries of responsibility, this approach addresses the issue of low accuracy in general translation of medical terminology, abbreviations, and dosage units. It establishes a high-precision professional translation capability for medical scenarios, resolving issues of terminological ambiguity, unit misunderstanding, and contextual breaks. It enables translation selection and continuous semantic expression consistent with the patient's medical context, addresses the lack of control over the pace and scope of information sharing among doctors, establishes an editable and controllable data delivery mechanism and content boundary management for healthcare professionals, resolves the inability to establish precise correspondences between text, audio, and images, achieves one-to-one anchoring at the sentence and image levels, and provides a basis for subsequent error correction. It addresses the issues of incomplete information acquisition and distracted attention for patients, enabling a focused presentation and intuitive understanding path for patients, and ensuring a continuous, stable, and synchronized experience during and after editing. Attached Figure Description
[0012] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. Some specific embodiments of the invention will be described in detail below with reference to the accompanying drawings in an exemplary and non-limiting manner. The same reference numerals in the drawings designate the same or similar parts or components. It should be understood by those skilled in the art that these drawings are not necessarily drawn to scale. In the drawings:
[0013] Figure 1 This is a flowchart illustrating an embodiment of the real-time intelligent translation method based on a clinical doctor-patient dual-screen translator according to the present invention.
[0014] Figure 2 This is a schematic diagram of a real-time intelligent translation method based on a clinical doctor-patient dual-screen translator according to an embodiment of the present invention;
[0015] Figure 3 This is an exploded structural diagram of an embodiment of the real-time intelligent translation method based on a clinical doctor-patient dual-screen translator according to the present invention;
[0016] Figure 4This is a schematic diagram of the structure of the outer side of the housing of the translator main unit according to an embodiment of the present invention, showing the translation bar configuration slot and the translation bar assembly state;
[0017] Figure 5 This is a schematic diagram illustrating a structure of the charging head configured in the slot of the translation stick according to an embodiment of the present invention;
[0018] Figure 6 This is a schematic diagram of a structure of the top hinge assembly of the screen bracket according to an embodiment of the present invention.
[0019] Explanation of reference numerals in the attached figures:
[0020] 1. Translator main unit; 10. Translator stick configuration slot; 101. First configuration slot; 102. Second configuration slot; 11. Charging head;
[0021] 2. Translation stick; 20. Charging port; 21. First translation stick; 22. Second translation stick;
[0022] 3. Pickup array;
[0023] 4. Adjust the concave surface of the screen; 40. First concave surface; 41. Second concave surface;
[0024] 5. Screen bracket; 50. Hinge assembly;
[0025] 6. Touchscreen. Detailed Implementation
[0026] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0027] Example 1
[0028] See Figure 1-6 This embodiment provides a real-time intelligent translation method based on a clinical doctor-patient dual-screen translation device, including the following steps:
[0029] S101. In a clinical diagnosis and treatment scenario, establish a diagnosis and treatment session, bind the medical staff screen, the patient screen and the translator host in a group, initialize the session identifier, target language and display synchronization parameters in the translator host, and write the current patient identifier into the session context; during the session, the translator host maintains the synchronization control relationship and time axis number between the two screens and the host, which is used to drive the unified arrangement of text, audio and visual materials.
[0030] S102. During the conversation, the translator host determines the speaker's location and role, automatically selects the translation direction, and enters medical mode when a medical staff member is detected speaking. The translator host's audio pickup channel collects the medical staff member's speech, displays the source language transcription on the medical staff screen, and simultaneously displays the corresponding translation on the patient screen and plays it aloud. When a patient member is detected speaking, the translator host enters patient mode, collects the patient member's speech through the audio pickup channel, displays the source language transcription on the patient screen, and simultaneously displays the corresponding translation on the medical staff screen and plays it aloud. During mode switching, the conversation time sequence and playback queue are kept continuous, and the current speaking role and translation direction are clearly marked on both screens.
[0031] S103. During the translation process, the translation machine host calls upon a multilingual translation engine and terminology resources for the medical field, and performs terminology alignment and context disambiguation based on the current patient's medical records, examination reports, and shared content selected by the medical staff. The medical staff screen initiates the sharing control of visual materials, establishes timestamps or fragment ID anchors between the selected visual materials and corresponding sentence fragments, and presents them synchronously on the patient screen in a mirror manner, so that the two screens complete the aligned display of the source text, translation, and visual materials on the same timeline.
[0032] In this embodiment, a clinical diagnosis and treatment session is established, binding the medical staff screen, the patient screen, and the translator host together. The translator host initializes the session identifier, target language, and display synchronization parameters, and writes the current patient's identifier into the session context. This solves the problems of lack of unified control and session mismatch between multiple terminals, achieving individualized parameter loading and end-to-end consistent orchestration centered on the session identifier. During the session, the translator host maintains the synchronization control relationship and timeline numbering between the two screens and the host, driving the unified orchestration of text, audio, and visual materials. This solves the problem of text, audio, and video being out of sync, achieving precise alignment and stable presentation of cross-modal content on the same timeline. The translator host determines the speaker's position and role, automatically selecting the translation direction. This solves the problems of unclear speaker attribution and translation errors in complex acoustic environments, achieving low-latency role determination and correct translation routing on both the medical staff and patient sides. When a medical staff member speaks, the system enters medical staff mode. The translator's main unit captures the medical staff's speech through its audio pickup channel, displays the transcribed source language on the medical staff's screen, and simultaneously displays the corresponding translation on the patient's screen with synchronized audio playback. This solves the problem of delayed patient comprehension and enables dual-screen collaborative display of medical staff speech while allowing the patient to hear and read it instantly. When a patient speaks, the system enters patient mode. The translator's main unit captures the patient's speech through its audio pickup channel, displays the transcribed source language on the patient's screen, and simultaneously displays the corresponding translation on the medical staff's screen with synchronized audio playback. This solves the problem of doctors not being able to understand the patient's statements in real time, enabling real-time translation and feedback of the patient's speech and simultaneous grasp of key points by the medical staff. During mode switching, the conversation time sequence and playback queue are maintained continuously, and the current speaking role and translation direction are clearly marked on both screens. This solves the problems of content loss, repeated playback, and role confusion during switching, achieving smooth transitions across rounds and clear boundaries of dialogue responsibility. By utilizing multilingual translation engines and terminology resources tailored to the medical field, the accuracy of general translation for medical terminology, abbreviations, and dosage units can be addressed, enabling high-precision professional translation capabilities in medical scenarios. Performing terminology alignment and context disambiguation based on current patient medical records, examination reports, and selected shared content from the healthcare provider resolves issues of terminological ambiguity, unit misunderstandings, and contextual breaks, ensuring consistent terminology selection and continuous semantic expression within the context of the consultation. Controlling the sharing of visual materials initiated by the healthcare provider's screen addresses the lack of control over the pace and scope of sharing by doctors, enabling an editable and controllable data push mechanism and content boundary management for the healthcare provider. Establishing timestamps or segment IDs as anchor points between selected visual materials and corresponding sentence fragments solves the problem of establishing precise correspondences between text, audio, and images, achieving sentence-level and screen-level one-to-one anchoring and providing a basis for subsequent error correction. Simultaneous mirroring on the patient's screen addresses the issues of incomplete information acquisition and distracted attention for the patient, achieving focused presentation and an intuitive understanding path for the patient.By aligning the source text, translation, and visualizations on both screens along the same timeline, the problem of flickering, skipping, and audio-visual asynchrony when doctors correct text or change data can be solved, achieving a continuous and stable synchronized experience during and after editing.
[0033] Preferably, the translator host performs environmental adaptive sound pickup channel selection before entering medical mode or patient mode. The selection includes: estimating the noise intensity and spectral distribution of the current sound field of the medical session, forming a channel score based on the speaker's position, and prompting the removal of the translator stick mounted on the translator host as the sound pickup channel and entering the corresponding mode if the score points to the near-field sound pickup channel; if the score points to the desktop sound pickup channel, the microphone array of the translator host is used as the sound pickup channel and the corresponding mode is entered. When switching channels, a seamless transition is set for the corresponding translated text segment that has not yet been played out, based on the time axis number, and the switching event is recorded in the channel log corresponding to the session identifier to avoid voice interruption and text misalignment caused by mode and channel asynchrony. In this embodiment, the translator host performs environmental adaptive sound pickup channel selection before entering medical or patient mode. It estimates the sound field noise intensity and spectrum distribution and forms a channel score based on the speaker's position. When the score points to the near channel, it prompts the user to remove the translation stick mounted on the translator host as the sound pickup channel. When the score points to the desktop channel, the microphone array of the translator host picks up the sound. Furthermore, a seamless transition is set according to the time axis number and the channel log is recorded when switching channels. This can solve problems such as improper sound pickup path selection and low translation accuracy in noisy environments, and achieve stable translation and translation channel management under high signal-to-noise ratio sound pickup.
[0034] Preferably, when the translator host determines the speaker's position and role, upon detecting a continuous sound source, it first divides the speech into segments using a preset minimum speaking unit, calculates the role confidence of each segment, and links it with the start and end silence thresholds. When both the medical staff and the patient speak simultaneously, the ownership of the current round is determined according to the rule that the medical staff has priority and must meet the preset minimum occupation time. On the other side, the timeline number is kept continuous by a waiting broadcast placeholder. If repeated contention occurs within the threshold window, the established role is maintained until the current segment ends before being released. In this embodiment, when determining the speaker's position and role, the translator host first divides the segment into segments using a preset minimum speaking unit, calculates the role confidence, and links it with the start and end silence thresholds. When both the medical staff and the patient speak simultaneously, the medical staff takes priority and must meet the minimum occupation time requirement to determine the ownership of the current round, while the other side maintains the continuity of the timeline by waiting for broadcast. Repeated contention within the threshold window maintains the established role until the segment ends. This can solve the problems of frequent direction switching and round disorder caused by multiple concurrent users and interruptions, and achieve round control with clear role boundaries, stable direction, and low latency.
[0035] Preferably, the terminology alignment and context disambiguation include: the translator host extracts structured entries from the current patient's medical records and examination reports as contextual constraints, generates multiple candidate translations for identified medical terms in the source language, rearranges them according to departmental relevance, entry matching degree, and historical selection frequency, and provides an interactive entry point for locking terminology translations on the medical staff screen; once locked, subsequent segments reuse the same translation and are effective within the session identifier range; if the doctor revokes the lock, the rearrangement strategy is restored, and both locking and revocation are written into the version index with timeline numbers for retrospective purposes. In this embodiment, in the terminology alignment and context disambiguation, the translator host extracts structured entries from the patient's medical records and examination reports as context, generates multiple candidate translations for identified medical terms and rearranges them according to departmental relevance, entry matching degree, and historical selection frequency, and the medical staff screen provides an entry point for locking translations and writes the locking and revocation into the version index with timeline numbers, which can solve the problems of medical terminology ambiguity and inconsistency in translations within the same session, and achieve terminology selection consistent with the clinical context and retrospective consistency control.
[0036] Preferably, when the two screens are displayed in alignment, when the medical staff screen performs text correction on a certain source language transcription or performs manual revision on the corresponding translation, the translation machine host generates a new version with the same segment ID as the original segment, and seamlessly replaces the rendered area of the patient screen and the currently playing audio queue while maintaining timestamp alignment. After the replacement action is completed, the old version is downgraded to a retrospective state but no longer participates in rendering, and only the latest version is output when the session is exported. In this embodiment, when the two screens are displayed in alignment, the medical staff screen generates a new version with the same segment ID when correcting the source language transcription or corresponding translation, seamlessly replacing the rendered area of the patient screen and the currently playing audio queue while maintaining timestamp alignment, and downgrading the old version to only retrospective without rendering, and exporting only the latest version, can solve the problems of flickering, segment skipping, and audio-visual asynchrony during the correction process, and achieve uninterrupted real-time error correction and stable presentation.
[0037] Preferably, the visualized data includes images, charts, and medication guidance. In the sharing control of the visualized data, when the medical staff selects the visualized data to be shared on the medical staff screen, the translation machine host performs masking or blurring of the identity identification area, contact information area, and free text area according to preset desensitization rules, and generates a segment ID anchor point corresponding to the current sentence segment for the desensitized data; the patient screen only mirrors the desensitized view layer and automatically focuses on the anchor point area when the corresponding segment appears. If the medical staff changes the data or changes the aligned segment again, the translation machine host updates the anchor point according to the timeline number and maintains consistency between the two screens. In this embodiment, in the control of shared visual data (images, charts, medication guidance), the translation machine host performs masking or blurring of the identity identification area, contact information area, and free text area, and generates a segment ID anchor point corresponding to the current sentence segment for the desensitized data; the patient screen only mirrors the desensitized view and automatically focuses on the anchor point when the corresponding segment appears. When the data or aligned segment changes, the anchor point is updated according to the timeline number, which can solve the privacy leakage and text-image synchronization problems in mirrored display, and achieve accurate anchoring and consistency between the two screens under compliant desensitization.
[0038] Preferably, when the network fails, the online translation service of the translator host is unavailable, or the latency exceeds a threshold, it automatically switches to offline fault-tolerant mode. The translation is completed using a locally cached multilingual engine and terminology resources, and the offline-generated segments are marked as pending review. When the network is restored, the segments to be reviewed are submitted in batches to the online engine for translation comparison. If the new translation is superior to the old one, a one-click replacement is prompted on the medical staff screen while maintaining the timestamp and segment ID, and a silent update is performed on the patient screen. If the medical staff refuses the replacement, the offline version is retained and the review decision is recorded. In this embodiment, when online translation is unavailable or the latency exceeds a threshold, the system automatically switches to offline fault-tolerant mode, uses a local multilingual engine and terminology resources for translation, and marks the segments as pending review. After the network is restored, batch translation comparisons are performed. If the new translation is superior to the old one, it is replaced with one click while maintaining the timestamp and segment ID, and a silent update is performed on the patient screen. If the old translation is rejected, the offline version is retained and the decision is recorded. This can solve the problems of session interruption and decreased consistency caused by weak or outdated networks, achieving improved clinical availability by enhancing translation continuity and post-translation consistency.
[0039] Preferably, when the source language transcription includes dosage, frequency, course of treatment, and unit information, structured slots are parsed out. These structured slots are represented as: drug—dosage—unit—frequency—course of treatment, and are cross-validated against the patient's allergy history and existing medications in their medical record. If potentially confusing units or possible contraindications are detected, a double confirmation dialog box pops up on the medical staff screen, providing at least two standardized expressions for selection and confirmation. After confirmation, the patient screen only displays the confirmed standardized translation, while unconfirmed candidates are not broadcast simultaneously. In this embodiment, when the source language transcription includes dosage, frequency, course of treatment, and unit information, structured slots containing drug—dosage—unit—frequency—course of treatment are parsed out and cross-validated against the patient's allergy history and existing medications in their medical record. If unit confusion or potential contraindications are detected, a double confirmation dialog box pops up on the medical staff screen, providing standardized expressions for selection. After confirmation, the patient screen only displays the confirmed standardized translation, and unconfirmed candidates are not broadcast. This can solve the problem of unit misunderstanding and safety risks in medication communication, achieving standardized and verifiable medication information transmission and misuse prevention.
[0040] Preferably, the translation device host supports the medical staff to select an easy-to-understand mode during session initialization. Once enabled, the easy-to-understand mode automatically reduces the speaking speed, increases the proportion of pauses, and raises the word spacing. Simultaneously, the patient screen increases the font size, line spacing, and contrast, and provides a soft key to replay the sentence at the bottom of each segment. When the patient is determined to have difficulty understanding based on reading time and the number of repeated requests, the medical staff screen receives a prompt to enable the easy-to-understand mode. Once enabled, the easy-to-understand mode remains active within the session identifier range until manually disabled or the session ends. In this embodiment, the medical staff can select an easy-to-understand mode during session initialization. Once enabled, it automatically reduces the speaking speed, increases the proportion of pauses, and raises the word spacing. The patient screen simultaneously increases the font size, line spacing, and contrast and provides a replay function. The prompt to enable the easy-to-understand mode based on reading time and repeated requests, and its continued effectiveness within the session identifier range, can solve the problem of insufficient accessibility in reading and listening comprehension, achieving adaptive human-computer interaction and higher information comprehensibility.
[0041] Preferably, when the physical combination of the two screens and the translation machine host changes due to the patient's transfer to an examination room or change of consultation room, the medical staff restores the existing state on the new device combination by scanning the session identifier or entering the session code. The restored content includes the locked terminology list, timeline number, segment ID and its version index, broadcast parameters, and visualization anchor points. After restoration, the first segment continues numbering from the point of last interruption to ensure that the source text, translation, and shared data are still presented aligned on the same timeline in the new location. In this embodiment, when the device combination changes due to the patient's transfer to an examination room or change of consultation room, the medical staff restores the existing state on the new device combination by scanning the session identifier or entering the session code, restoring the locked terminology list, timeline number, segment ID and version index, broadcast parameters, and visualization anchor points, and continues numbering from the point of interruption. This can solve the problem of context loss and repeated communication caused by cross-space migration, and achieve seamless connection and timeline continuity across multiple spaces.
[0042] Preferably, when the translation host determines that the recognition or translation confidence of a certain segment is below a threshold, it pauses the audio push to the patient screen, only displays the placeholder text being confirmed on the patient screen, and simultaneously highlights uncertain words on the medical staff screen and provides at least two alternative translations or a "please restate" option; if the confirmation is completed within the time limit, the confirmed text and audio are played synchronously and the timestamp remains unchanged; if the confirmation is not completed within the time limit, a simplified conservative version is pushed and recorded as an item requiring review. In this embodiment, when the recognition or translation confidence is below the threshold, the audio push to the patient screen is paused, only the placeholder text being confirmed is displayed, and uncertain words are highlighted on the medical staff screen with alternative translations or a "please restate" option; if the confirmation is completed within the time limit, the confirmed text and audio are played synchronously and the timestamp remains unchanged; if the time limit is exceeded, a simplified conservative version is pushed and recorded as an item requiring review. This can solve the risk of misunderstanding caused by directly presenting low-confidence output to the patient, and achieve a safe-first human-in-the-loop confirmation and a conservative backtrack without disrupting the timeline.
[0043] Example 2
[0044] See Figures 1-6This invention provides a translator that uses the aforementioned real-time intelligent translation method based on a clinical doctor-patient dual-screen translator for real-time intelligent translation in doctor-patient scenarios. The translator includes: a translator host 1, the host's housing having a translation stick configuration slot 10, a charging head 11 inside the slot, and the charging head electrically connected to a power module inside the host's housing; dual touchscreens 6, connected and communicating with the host, respectively positioned above two screen adjustment concave curved surfaces formed on the outer side of the host's housing; the concave direction of the screen adjustment concave curved surfaces faces the interior of the host's housing, forming a screen adjustment accommodating space on the outer side of the host's housing; each screen adjustment concave curved surface has a corresponding screen bracket for mounting one of the dual touchscreens 6, and each touchscreen is rotatably connected to its corresponding screen bracket; wherein, when a touchscreen rotates on its corresponding screen bracket, the screen adjustment accommodating space accommodates the rotation trajectory of the touchscreen. When the device touches the concave curved surface of the screen, it rotates to the end of its rotation trajectory. In the dual touchscreen setup, one serves as the medical screen and the other as the patient screen. The translation stick 2 is mounted in the translation stick configuration slot 10 and is used by the user to remove it in noisy environments and place it near the sound source to receive voice signals, translate, and play them. The translation stick has a charging interface 20. When the translation stick is mounted in the configuration slot, the charging interface is electrically connected to the charging head located in the configuration slot. The power module charges the translation stick, and the translation stick's microphone is turned off while charging. The inner wall of the translation stick configuration slot can be provided with an oblique guide groove matching the shape of the translation stick 2. Physical error prevention is achieved through an asymmetrical contour (such as a trapezoidal cross-section), ensuring that the translation stick 2 can only be inserted in a single direction. A permanent magnet array can also be embedded in the bottom of the configuration slot. The permanent magnet array resonates with the ferromagnetic material on the translation stick 2, providing initial adsorption force to assist the user in aligning and assembling the translation stick 2.
[0045] The translator host monitors the electrical connection status between the translator stick and the charging head located in the translator stick's configuration slot. When the electrical connection between the translator stick and the charging head in the translator stick's configuration slot is broken, the translator stops translating and playing the user's voice signal. When the electrical connection between the translator stick and the charging head in the translator stick's configuration slot is maintained, the translator host receives the user's voice signal for translation and playback. The bottom of the translator host's housing is provided with an interface slot, the bottom of which is provided with a high-definition video interface, and the side wall of the interface slot is provided with a cable passage groove. When the high-definition video interface is connected to a high-definition video transmission line, the high-definition video transmission line passes through the cable passage groove, transmitting the high-definition video signal output by the translator host to a display terminal with a screen size larger than the dual touch screen for large-screen display. Furthermore, the housing of the translator host 1 is provided with a microphone array 3. The microphone array is used to locate and select sounds from different directions when the microphone of the translator stick itself is turned off, picking up the sound from the speaker's direction and separating the main speaker's voice signal for the translator host to translate and play.
[0046] In this embodiment, the translator host 1 has a translator stick configuration slot 10 on its housing. Inside the translator stick configuration slot 10 is a charging head 11 electrically connected to the power module. This solves the problem of desktop translators lacking physical support and power supply interfaces for detachable peripherals, making integrated peripheral management inconvenient. It enables the orderly storage, interface alignment, and stable power supply of the translator stick 2 on the translator host 1. The translator stick 2 is assembled in the translator stick configuration slot 10, and its charging interface 20 is electrically connected to the charging head 11. This solves the problem of the translator stick needing a separate charging cable, and the cumbersome usage process caused by separating charging and storage. It optimizes the process by integrating charging, storage, charging, and use of the translator stick into a single, integrated process. The power module charges the translator stick 2, solving the problem of limited battery life and frequent power outages affecting on-site translation, enabling continuous power replenishment and improved usability of the translator stick 2 at the desktop workstation. The translator stick 2, placed close to the sound source in noisy environments, allows users to receive speech signals. This addresses the issues of low signal-to-noise ratio (SNR) and decreased speech recognition and translation accuracy in high-noise environments caused by far-field microphones. It improves SNR and robustness in transcription and translation through near-field microphone pickup. SNR, usually expressed in decibels (dB), is the ratio of speech signal intensity to background noise intensity. In this embodiment, when the translator stick 2 is placed close to the sound source (the speaker's mouth), the speech signal amplitude increases significantly (due to the short distance and minimal energy attenuation). Background noise remains relatively unchanged or decreases (because the near-field microphone has a small pickup range). This significantly improves SNR, making speech clear and indistinguishable. Transcription robustness refers to the accuracy, stability, and low error rate of the transcription results when converting speech signals into text under different noise conditions, accents, and speaking speeds. Translation robustness refers to the accuracy and consistency of translation results under different scenarios and noise backgrounds. For machine translation, the more accurate the input text and the fewer errors or omissions, the more reliable the translation result.
[0047] In this embodiment, the translation stick 2 translates and plays the received voice signal, which solves the problem of lack of real-time feedback and playback when the device is far from the host, and realizes real-time translation closed-loop and local broadcasting close to the sound source. When the translation stick 2 is installed in the configuration slot and is in the charging state, the translation stick 2 turns off its own microphone. At this time, the translation and playback are provided by the translation host 1. This can solve the problems of parallel channel conflict, crosstalk and howling caused by simultaneous sound pickup at the host and stick ends, realize automatic channel management according to charging / parking status, avoid acoustic interference and save power consumption of the stick end.
[0048] In this embodiment, the translator host 1 monitors the electrical connection status between the translator stick 2 and the charging head 11. This solves the problem of sensing the electrical connection status of the translator stick and enabling mode linkage based on the translator stick's presence / absence. It achieves translation drive with electrical connection as a reliable trigger, providing a criterion for automatically switching translation playback paths. When the electrical connection between the translator stick 2 and the charging head 11 is broken, the translator host 1 stops translating and playing the user's voice signal. This solves the problem of channel contention and acoustic loop superposition when the translator stick 2 is picked up for near-field communication, achieving mutual exclusion switching between near-field communication and far-field communication, avoiding dual-path parallel interference. When the translator stick 2 is charging in its slot and its microphone is off, the translator host 1 translates and plays, enabling adaptive scene switching: when the stick is in place (charging), the host performs far-field communication; when the stick is out of place (near-field communication), the host stops translating and playing.
[0049] In this embodiment, the modular hardware architecture consisting of the host unit, the translation stick, the physical housing of the translation stick, and the charging interface 20 solves the problems of simple hardware structure, poor expandability, and difficulty in coupling with software functions in desktop translators. This enables hardware and software collaboration for multiple scenarios and provides an expandable hardware base. Through automatic power on / off based on connection status and translation playback path management (i.e., two control logics: charging turns off the stick's audio pickup, disconnecting turns off the host unit), the problems of cumbersome manual switching, high error rates, and translation interruptions or errors caused by inconsistent statuses are solved. This achieves automated scenario adaptation with zero learning cost and a stable and consistent user experience.
[0050] In this embodiment, a screen adjustment and accommodating space is formed on the outer side of the translator's main unit casing. Each screen adjustment concave curved surface is equipped with a corresponding screen bracket to install one of the dual touchscreens. Each touchscreen and its corresponding screen bracket are rotatably connected, and the angle and orientation are adjustable to adapt to different heights, sitting postures, and reflective environments, avoiding fatigue from prolonged use. When rotating and adjusting the screen, the screen adjustment concave curved surface can accommodate the screen without obstructing its rotation within a preset trajectory range. This allows the overall structure of the translator to be as compact as possible, reducing the overall size of the device. Furthermore, the screen and main unit are separated by the bracket, allowing for separate maintenance and facilitating easy upkeep. The integration of the translation stick effectively expands the hardware structure and functions of the translator, adapting to different translation scenarios in conjunction with the main unit.
[0051] Preferably, the translation stick configuration slot 10 includes a first configuration slot 101 and a second configuration slot 102; the translation stick 2 includes a first translation stick 21 and a second translation stick 22 for corresponding configuration within the first configuration slot 101 and the second configuration slot 102; when the first translation stick 21 and the second translation stick 22 are correspondingly assembled within the first configuration slot 101 and the second configuration slot 102, the charging interfaces 20 of the first translation stick 21 and the second translation stick 22 are electrically connected to the charging heads 11 provided within the first configuration slot 101 and the second configuration slot 102, respectively; the power module charges the first translation stick 21 and the second translation stick 22; and the first translation stick 21 and the second translation stick 22 turn off their microphones while charging.
[0052] In this embodiment, the translation stick configuration slot 10 includes a first configuration slot 101 and a second configuration slot 102. The translation stick 2 includes a first translation stick 21 and a second translation stick 22, which are configured into the first configuration slot 101 and the second configuration slot 102 respectively. When the two are assembled in the corresponding slots, their charging interfaces 20 are electrically connected to the charging heads 11 in the configuration slots respectively, and are charged by the power module. In the charging state, the microphones of the two translation sticks are turned off, which can solve the problem of frequent switching and low efficiency of a single translation stick 2 in multiple scenarios, multiple languages or multiple people translating at the same time. In this way, two translation sticks 2 can be configured at the same time in a desktop translator, which can support multiple people to speak and translate at the same time, reducing waiting time. At the same time, the automatic shutdown of the microphones of the translation sticks 2 in the charging state can avoid channel interference and acoustic crosstalk caused by the simultaneous pickup of two translation sticks 2, realize the collaborative management of multiple translation sticks 2, automated translation playback path control, and higher device utilization and scene coverage.
[0053] Preferably, when the microphones of the first translator stick 21 and the second translator stick 22 are both turned off, the microphone array 3 locates and selects sounds from different directions, picks up the sound from the speaker's direction, and separates the main speaker's voice signal for the translator host 1 for translation and playback. The translator host 1 monitors the connection status of the first translator stick 21, the second translator stick 22, and the charging heads 11 disposed in the first configuration slot 101 and the second configuration slot 102. When the connection is broken, the microphone array 3 is controlled to turn off.
[0054] In this embodiment, the translator host 1 monitors the electrical connection status of the first translator stick 21, the second translator stick 22, and the charging head 11 in the first configuration slot 101 and the second configuration slot 102. When a disconnection occurs, the microphone array 3 is shut down. This solves the problem of the multi-translator stick 2 system lacking an automated microphone control mechanism based on hardware connection status, and the inability to prevent resource conflicts and acoustic feedback caused by simultaneous microphone pickup from the array and stick ends. In this embodiment, the connection status of the charging interface 20 is used as a reliable hardware trigger signal to achieve automatic start / stop control of the array microphones: when a translator stick 2 is removed from use, the array shuts down to avoid dual-channel interference; when the stick end is returned to charging, the array can restart for far-field microphone pickup. This improves the system's adaptability and stability in multiple scenarios, reduces manual switching steps, and lowers the risk of misoperation.
[0055] Preferably, the disconnection status includes: the electrical connection between the first translator 21 and the charging head 11 disposed in the first configuration slot 101 is disconnected; or the electrical connection between the second translator 22 and the charging head 11 disposed in the second configuration slot 102 is disconnected; or the electrical connections between the first translator 21 and the second translator 22 and the charging heads 11 disposed in the first configuration slot 101 and the second configuration slot 102 are both disconnected. In this embodiment, the disconnection status includes: the first translator 21 is disconnected from the charging head 11 in the first configuration slot 101, or the second translator 22 is disconnected from the charging head 11 in the second configuration slot 102, or both are disconnected. This can solve the defect that a single connection status detection logic cannot cover the usage scenarios of multiple translator 2s, and avoid the pickup conflict caused by the inability to shut down the array in time when one translator 2 is removed. In this embodiment, multiple determination conditions are provided to ensure that as long as any translator 2 is in the removed working state, the host can respond and shut down the array pickup, preventing sound source aliasing and echo problems caused by the simultaneous operation of the array and the proximity microphone. This more refined connection disconnection determination mechanism can improve the accuracy of translation playback path switching and scene adaptability of the MultiTranslator 2 system, making hardware collaboration more intelligent and reliable.
[0056] It should be noted that when only one of the first translation stick 21 or the second translation stick 22 is electrically disconnected from its corresponding charging head 11, the microphone array 3 remains off and the disconnected translation stick 2 is designated as the main voice channel, while the unconnected translation stick 2 is placed in a charging and microphone-off state. When both the first translation stick 21 and the second translation stick 22 are electrically disconnected from their respective charging heads 11, the microphone array 3 remains off and the status indicator light is illuminated to prompt the user that the translation stick 2 has been selected as the proximity device to enter the conversation process, thereby refining the channel strategy in the case of disconnection and improving operability.
[0057] Preferably, the translation stick configuration slot 10 is located beside the concave curved surface 4 for screen adjustment, and is located at different parts of the outer side of the translator host 1 housing, respectively, from the screen adjustment accommodating space. In this embodiment, the concave curved surface 4 for screen adjustment and the translation stick configuration slot 10 are arranged in separate sections to avoid mutual interference, improve the overall appearance integration of the device, enhance the convenience of peripheral device management, and optimize operation management in desktop scenarios.
[0058] Preferably, the screen adjustment concave surface 4 includes a first concave surface 40 and a second concave surface 41; the first concave surface 40 and the second concave surface 41 are respectively formed at different parts of the outer side of the housing of the translator host 1, so as to form screen adjustment accommodating spaces at different positions on the outer side of the housing of the translator host 1. A screen support 5 is provided in the screen adjustment concave surface 4, and the screen support 5 is used to fix the touch screen so that the touch screen is located in the screen adjustment accommodating space. Each touch screen can be rotatably connected to the corresponding screen support by setting a pivot assembly 50 at the top of the screen support 5, so that the touch screen can rotate relative to the screen support about an axis parallel to its long side. It is understood that the specific structure of the pivot assembly that realizes the rotatable connection between the touch screen and the corresponding screen support is conventional technology in the art, and those skilled in the art can choose a suitable specific structure according to specific needs, which will not be described in detail in this embodiment.
[0059] It should be noted that the above embodiments are merely preferred embodiments of the present invention, and the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention, and the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A real-time intelligent translation method based on a clinical doctor-patient dual-screen translation device, characterized in that, include: In a clinical diagnosis and treatment scenario, a diagnosis and treatment session is established, and the medical staff screen, the patient screen and the translation machine host are grouped and bound together. The session identifier, target language and display synchronization parameters are initialized in the translation machine host, and the current patient identifier is written into the session context. During the session, the translation machine host maintains the synchronization control relationship and time axis number between the two screens and the host, which is used to drive the unified arrangement of text, audio and visualization materials. During the conversation, the translator host determines the speaker's location and role, automatically selects the translation direction, and enters the medical staff mode when the medical staff speaks. The voice pickup channel of the translator host collects the medical staff's voice, displays the source language transcription on the medical staff screen, and displays the corresponding translation on the patient screen and broadcasts it synchronously. When a patient speaks, the system enters patient mode. The voice is picked up by the microphone of the translator and the original language is transcribed on the patient screen. At the same time, the corresponding translation is displayed on the medical staff screen and broadcast simultaneously. During mode switching, the session time sequence and playback queue are kept continuous, and the current speaker role and translation direction are clearly marked on both screens; During the translation process, the translation machine host calls upon a multilingual translation engine and terminology resources for the medical field, and performs terminology alignment and context disambiguation based on the current patient's medical records, examination reports, and shared content selected by the medical staff. The sharing control of visual materials initiated by the medical staff screen establishes a timestamp or fragment ID anchor point between the selected visual materials and the corresponding sentence fragments, and presents them synchronously on the patient screen in a mirror manner, so that the two screens complete the aligned display of the source text, translation and visual materials on the same timeline. Before entering medical or patient mode, the translator host performs an environment-adaptive audio pickup channel selection. This selection includes: estimating the noise intensity and spectral distribution of the current sound field in the medical session, forming a channel score based on the speaker's position, and prompting the user to remove the translator stick mounted on the translator host as the pickup channel and enter the corresponding mode if the score points to the near-field pickup channel; and using the microphone array of the translator host as the pickup channel and entering the corresponding mode if the score points to the desktop pickup channel. During channel switching, a seamless transition is set for the corresponding translated text segment that has not yet been played out, based on the timeline number, and the switching event is recorded in the channel log corresponding to the session identifier to avoid voice interruption and text misalignment caused by mode and channel asynchrony. When the source language transcription includes dosage, frequency, course of treatment, and unit information, structured slots are parsed out. The structured slots are represented as: drug—dosage—unit—frequency—course of treatment, and are cross-validated with the patient's allergy history and existing medications in the medical record. If a potentially confusing unit or a combination of possible contraindications is detected, a double confirmation dialog box pops up on the medical staff screen and provides at least two standardized expressions for selection and confirmation. After confirmation, the patient screen only displays the confirmed standardized translation, while the unconfirmed candidates are not broadcast synchronously.
2. The method according to claim 1, characterized in that, When the translator determines the speaker's location and role, upon detecting a continuous sound source, it first divides the speech into segments using a preset minimum speaking unit, calculates the role confidence of each segment, and links it to the start and end silence thresholds. When both the medical staff and the patient speak simultaneously, the ownership of the current round is determined according to the rule that the medical staff takes priority and must meet the preset minimum occupation time. On the other side, a waiting broadcast placeholder mark is used to maintain the continuity of the timeline number. If repeated contention occurs within the threshold window, the established role is maintained until the current segment ends before being released.
3. The method according to claim 1, characterized in that, The terminology alignment and context disambiguation include: the translator host extracts structured entries from the current patient's medical records and examination reports as contextual constraints, generates multiple candidate translations for medical terms identified in the source language, rearranges them according to departmental relevance, entry matching degree, and historical selection frequency, and provides an interactive entry point for locking terminology translations on the medical staff screen; once locked, subsequent segments reuse the same translation and are effective within the session identifier range; if the doctor cancels the lock, the rearrangement strategy is restored, and both locking and cancellation are written into the version index with timeline numbering for backtracking.
4. The method according to claim 1, characterized in that, When the two screens are aligned, when the medical staff screen performs text correction on a certain source language transcription or performs manual revision on the corresponding translation, the translation machine host generates a new version with the same segment ID as the original segment, and replaces the rendered area of the patient screen and the audio queue being played in a seamless splicing manner while maintaining timestamp alignment; after the replacement action is completed, the old version is downgraded to a retrospective state but no longer participates in rendering, and only the latest version is output when the session is exported.
5. The method according to claim 1, characterized in that, The visualized data includes images, charts, and medication instructions. In the sharing control of visualized data, when the medical staff selects the visualized data to be shared on the medical staff screen, the translation machine host performs masking or blurring of the identity identification area, contact information area, and free text area according to preset desensitization rules, and generates a segment ID anchor point corresponding to the current sentence segment for the desensitized data; the patient screen only mirrors the desensitized view layer and automatically focuses on the anchor point area when the corresponding segment appears. If the medical staff changes the data or changes the aligned segment again, the translation machine host updates the anchor point according to the timeline number and maintains consistency between the two screens.
6. The method according to claim 1, characterized in that, When the network fails, and the online translation service of the translator host is unavailable or the latency exceeds the threshold, it automatically switches to offline fault-tolerant mode, uses the locally cached multilingual engine and terminology resources to complete the translation, and marks the offline-generated segments as pending review; when the network is restored, the segments to be reviewed are submitted in batches to the online engine for retranslation and comparison. If the new translation is better than the old translation, a one-click replacement is prompted on the medical staff screen while maintaining the timestamp and segment ID unchanged, and a silent update is performed on the patient screen; if the medical staff refuses to replace, the offline version is retained and the review decision is recorded.
7. The method according to claim 1, characterized in that, During session initialization, the translation device allows medical staff to select an easy-to-understand mode. Once enabled, the easy-to-understand mode automatically reduces the speech rate, increases the pause ratio, and raises the word spacing. The patient screen simultaneously increases the font size, line spacing, and contrast, and provides a soft key to replay the sentence at the bottom of each segment. When the patient is judged to have difficulty understanding based on the reading time and the number of repeated requests, the medical staff screen receives a prompt to enable the easy-to-understand mode. Once enabled, the easy-to-understand mode remains effective within the session identifier range until it is manually turned off or the session ends.
8. A translation machine, characterized in that, The translator uses the real-time intelligent translation method based on a clinical doctor-patient dual-screen translator as described in any one of claims 1-7 to perform real-time intelligent translation in doctor-patient scenarios.
Citation Information
Patent Citations
Method and equipment for realizing double-screen different-display real-time translation machine
CN119669427A
Automatic interpretation and translation and dialogue assistance system using transparent display
KR102557092B1