Medical record report generation method for outpatient scenario and related device
By collecting and translating image and sound data through the intelligent badge system, the problem of communication barriers between doctors and patients in outpatient settings has been solved, standardized medical record generation has been achieved, and diagnostic efficiency and treatment quality have been improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN PEOPLES HOSPITAL
- Filing Date
- 2026-05-09
- Publication Date
- 2026-06-05
AI Technical Summary
In outpatient settings, communication barriers between doctors and patients and the lack of standardized medical record generation lead to low diagnostic efficiency, which existing technologies cannot effectively address.
The system uses an intelligent name tag system to collect image and sound data, identify and translate the synchronous translation types of intentions between doctors and patients, and generate standardized medical record reports.
It enables efficient medical record generation under various communication barriers, reducing doctors' paperwork burden and improving diagnostic efficiency and treatment quality.
Smart Images

Figure CN122157940A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method and related apparatus for generating medical record reports in outpatient settings. Background Technology
[0002] As the initial stage of medical services, efficient communication between outpatient doctors and patients, along with standardized medical record keeping, is crucial for ensuring the quality and efficiency of outpatient consultations. In actual outpatient settings, several issues remain regarding doctor-patient communication and medical record keeping: First, some patients have special circumstances such as dialect barriers, hearing impairments, or speech impairments, creating communication barriers between doctors and patients, potentially leading to information transmission errors and affecting diagnostic accuracy. Second, patients often lack a reserve of professional medical terminology, making it difficult to accurately and systematically describe symptom details. Doctors need to invest extra time in summarizing and organizing the consultation content, significantly reducing consultation and diagnostic efficiency. Third, in the current treatment process, unstructured consultation dialogues cannot be directly converted into standardized electronic medical records. Doctors still need to manually organize and record consultation information, adding to their paperwork burden, further compressing treatment decision-making time, and hindering the improvement of outpatient treatment efficiency.
[0003] Therefore, how to improve the efficiency of outpatient diagnosis and treatment is an urgent issue that needs to be addressed. Summary of the Invention
[0004] The purpose of this application is to provide a method and related apparatus for generating medical record reports in outpatient settings, which can solve the problem of low diagnostic efficiency of outpatient doctors due to communication barriers between outpatient doctors and patients and the lack of simultaneous generation of standardized medical records during the consultation process.
[0005] To achieve the objectives of this application, the following technical solution is provided: Firstly, this application provides a method for generating medical record reports in outpatient settings, applied to a server of a smart badge system. The smart badge system includes the server and smart badges communicating with the server. The method includes: The smart badges worn by outpatient doctors collect image and sound data from the consultation process. Based on the image data and the sound data, the type of synchronous translation of intent between the outpatient doctor and the patient is determined. The type of synchronous translation of intent includes one-way dialect translation, two-way dialect translation, and sign language-speech two-way translation. The interaction information between the outpatient doctor and the patient is processed according to the intent synchronous translation type to obtain processing result information; and the patient's consultation dataset is created based on the interaction information and the processing result information. Generate the patient's target medical record report based on the consultation dataset.
[0006] Secondly, this application provides a medical record report generation device for outpatient scenarios, applied to a server of a smart badge system. The smart badge system includes the server and smart badges communicating with the server. The medical record report generation device includes: The acquisition unit is used to collect image and sound data of the consultation scene through the smart badge worn by the outpatient doctor; The calculation unit is used to determine the type of synchronous translation of intent between the outpatient doctor and the patient based on the image data and the sound data. The type of synchronous translation of intent includes dialect one-way translation, dialect two-way translation, and sign language-speech two-way translation. The unit processes the interaction information between the outpatient doctor and the patient according to the type of synchronous translation of intent to obtain the processing result information. The control unit is configured to create a patient's medical history dataset based on the interaction information and the processing result information; and to generate a target medical record report for the patient based on the medical history dataset.
[0007] Thirdly, this application provides a smart badge, including a processor, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, and the programs include instructions for performing steps in any of the methods in the first aspect of the embodiments of this application.
[0008] Fourthly, this application provides a computer-readable storage medium storing a computer program for electronic data interchange, wherein the computer program causes a computer to perform some or all of the steps described in any of the methods of the first aspect of the embodiments of this application.
[0009] Fifthly, this application provides a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps described in any method of the first aspect of the embodiments of this application. The computer program product may be a software installation package.
[0010] This application provides a method and related apparatus for generating medical record reports in outpatient settings. By implementing the embodiments of this application, the following beneficial effects can be achieved: The method for generating medical record reports in outpatient settings first collects image and audio data from the consultation scenario using smart badges worn by outpatient doctors; second, it determines the type of simultaneous translation of intent between the outpatient doctor and the patient based on the image and audio data; third, it processes the interaction information between the outpatient doctor and the patient according to the type of simultaneous translation of intent to obtain processing result information; fourth, it creates a patient's consultation dataset based on the interaction information and processing result information; and finally, it generates the patient's target medical record report based on the consultation dataset. Among them, the synchronous translation types include dialect one-way translation, dialect two-way translation, and sign language-speech two-way translation. Specifically, dialect one-way translation can adapt to communication barriers where patients speak dialects that doctors cannot understand; dialect two-way translation can adapt to communication barriers where patients speak dialects that doctors cannot understand; and sign language-speech two-way translation can adapt to communication barriers where patients express themselves in sign language that doctors cannot understand. This allows for comprehensive and flexible adaptation to various communication barriers in actual consultation scenarios, improving the system's comprehensiveness, intelligence, and scenario applicability in processing communication information in consultation scenarios. In addition, the consultation dataset created based on interaction information and processing results covers as comprehensive an unstructured consultation content as possible, improving the comprehensiveness and intelligence of the system's medical record report creation, and further enhancing the diagnostic efficiency of outpatient doctors. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is an architecture diagram of an intelligent name tag system provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of an AI smart badge provided in an embodiment of this application; Figure 3 This is a flowchart illustrating the method for generating medical record reports in outpatient settings provided in this application embodiment; Figure 4 This is a schematic diagram of the appearance of an AI smart name tag provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of a case report provided in an embodiment of this application; Figure 6 This is a backend interface diagram of an intelligent name tag system provided in an embodiment of this application; Figure 7 This is a block diagram of the functional modules of a medical record report generation device for outpatient scenarios provided in an embodiment of this application. Detailed Implementation
[0013] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0014] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0015] It should be understood that the term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this document indicates that the preceding and following related objects are in an "or" relationship. In the embodiments of this application, "multiple" refers to two or more.
[0016] In the embodiments of this application, "at least one item" or its similar expression refers to any combination of these items, including any combination of a single item or a plurality of items. "One or more" means one or more, while "multiple" means two or more. For example, "at least one item" of a, b, or c can represent the following seven cases: a, b, c; a and b; a and c; b and c; a, b, and c. Each of a, b, and c can be an element or a set containing one or more elements.
[0017] In this application, the term "connection" refers to various connection methods, such as direct connection or indirect connection, to achieve communication between devices. This application does not impose any limitations on this.
[0018] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0019] The following is an explanation of the relevant terms used in this application: AI Smart Badge: The AI Smart Badge is a miniaturized smart terminal that can be worn by medical staff. It is an integrated smart hardware that can record data like a body camera worn by police officers, and also has intelligent medical analysis functions.
[0020] The following is combined with Figure 1 The architecture of an intelligent name tag system according to an embodiment of this application will be described. Figure 1 This is an architecture diagram of an intelligent name tag system provided in an embodiment of this application. The intelligent name tag system 100 includes an AI intelligent name tag 110, a server 120, and a hospital information system 130.
[0021] The AI smart badge 110 includes a camera module 111, a voice interaction module 112, a control module 113, and a communication module 114. Deployed at the outpatient doctor's end, the AI smart badge 110 can be worn by the doctor or placed on the desktop during consultations to collect voice and video data from both the doctor and patient. As a front-end sensing and interaction carrier for doctor-patient interaction, the AI smart badge 110, through multimodal data collection, real-time translation and output, and interactive feedback, simultaneously completes the front-end collection and preliminary preprocessing of consultation interaction information, providing high-quality raw data input for back-end algorithm operation.
[0022] In one possible embodiment, the camera module 111 includes a high-definition wide-angle camera for real-time acquisition of visual image data such as the patient's facial expressions, gestures, and body posture during a consultation, providing visual feature input for sign language recognition and intent translation type determination; the voice interaction module 112 is equipped with a high-sensitivity array microphone and a directional speaker, which on the one hand acquires the voice signals of both the doctor and the patient, completes the acquisition of raw sound data and front-end noise reduction preprocessing, and on the other hand, plays the translated voice in real time and pushes interactive prompts according to the translation results sent by the server, realizing the visualization and audible output of two-way voice-sign language translation; the control module 1 13 is the control unit of the AI smart badge 110, which coordinates the collaborative work of the camera module, voice interaction module and communication module, and completes preprocessing operations such as framing and normalization of front-end data to ensure low latency and high stability of each module. The communication module 114 adopts a wireless communication protocol that complies with medical data security standards (such as Wi-Fi 6, 5G private network) to realize bidirectional encrypted data transmission between the AI smart badge 110 and the server 120, uploads the collected multimodal interaction information to the server 120, and receives translation instructions, draft medical records and other data issued by the server 120 to complete efficient information interaction between the front end and the back end.
[0023] The server 120 receives and processes multimodal interactive information uploaded by the AI smart badge 110, executing algorithms such as intent translation type determination, interactive information translation processing, consultation dataset construction, and target medical record report generation. It also completes standardized data interaction with the hospital information system 130, enabling automatic archiving of medical records and interoperability of medical data. The front-end and back-end collaborative architecture built by the AI smart badge 110 and server 120 through the communication module 114 decouples the lightweight front-end perception interaction from the complex back-end computing power scheduling. This ensures both the portable and low-latency deployment of the AI smart badge 110, meeting the interactive needs of real-time outpatient consultations, and leverages the powerful computing power of the server 120 to complete complex computational tasks such as multimodal fusion and model inference, achieving an optimal balance between system performance and deployment cost.
[0024] In one possible embodiment, server 120 is equipped with algorithms such as a multimodal intent translation model fine-tuned from outpatient medical corpus, a dialect-to-Mandarin bidirectional translation model, a sign language-to-speech bidirectional translation model, and a medical entity relationship extraction model. Through image and sound data uploaded by the AI smart badge 110, it accurately determines the intent synchronization translation type, performs differentiated interactive information translation processing for different translation types, generates a time-aligned, timestamped standardized consultation dataset, and automatically generates a target medical record report conforming to hospital clinical standards based on the consultation dataset. Simultaneously, server 120 achieves bidirectional communication with hospital information system 130 (i.e., HIS system) through standardized medical data interfaces such as HL7, synchronously archiving the final doctor-confirmed target medical record report to hospital information system 130, and retrieving patient basic information, outpatient medical record templates, and standardized medical terminology databases from hospital information system 130 to provide a standardized framework and terminology support for medical record generation. Furthermore, the hospital information system 130, serving as the carrier of the hospital's existing information infrastructure, stores core medical data resources such as basic patient diagnosis and treatment information, outpatient electronic medical record templates, and a standardized medical terminology knowledge base. It provides standardized format specifications and terminology mapping support for the medical record generation process of server 120, while receiving target medical record reports uploaded by server 120, completing the digital archiving, storage, and clinical retrieval of medical records. This achieves seamless integration between the intelligent badge system and the hospital's existing information system, enabling rapid deployment without large-scale modifications to the existing hospital information system, significantly improving the system's scenario adaptability and engineering feasibility.
[0025] It should be noted that the architectural division of the AI smart badge 110, server 120, and hospital information system 130 is only an illustrative example. In practical applications, server 120 can be a distributed computing system, and the functions of each module can be flexibly adjusted according to the hospital's computing power deployment needs and information architecture. For example, some lightweight algorithm models can be deployed in the control module 113 of the AI smart badge 110 to achieve real-time inference on the terminal side, further reducing the system's communication latency. At the same time, the communication protocol of communication module 114 can be adapted according to the hospital's network security requirements, and an end-to-end encrypted transmission mechanism can be adopted to ensure the transmission security and privacy protection of medical data, complying with relevant medical data security regulations.
[0026] It is evident that through the three-tiered collaborative architecture of "front-end perception and interaction - back-end computing and scheduling - hospital system integration", the entire process of intelligent communication between doctors and patients, automatic collection of consultation data, and automatic generation of medical records in outpatient settings can be realized. This effectively breaks down communication barriers for special patients such as those with dialects or hearing impairments, significantly reduces doctors' paperwork workload, and improves outpatient diagnosis and treatment efficiency and service quality.
[0027] To more clearly illustrate the structure of the AI smart badge, the following will combine... Figure 2 The structure of the AI smart badge in the embodiments of this application will be described. Figure 2 This is a schematic diagram of the structure of an AI smart badge provided in an embodiment of this application, such as... Figure 2 As shown, the AI smart badge 110 includes a processor 210, a memory 220, a communication interface 230, and one or more programs 221. The processor 210 is connected to the memory 220 and the communication interface 230 via an internal communication bus.
[0028] The processor 210 is mainly used for: collecting image and sound data of the consultation scene through the smart badge worn by the outpatient doctor; determining the type of synchronous translation of intent between the outpatient doctor and the patient based on the image and sound data, including dialect one-way translation, dialect two-way translation, and sign language-speech two-way translation; processing the interaction information between the outpatient doctor and the patient according to the type of synchronous translation of intent to obtain the processing result information; creating the patient's consultation dataset based on the interaction information and the processing result information; and generating the patient's target medical record report based on the consultation dataset.
[0029] The one or more programs 221 are stored in the memory 220 and configured to be executed by the processor 210. The one or more programs 221 include instructions for performing any step in the above method embodiments.
[0030] The processor 210 can be a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, units, and circuits described in conjunction with the disclosure of this application. The processor can also be a combination that implements computational functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc. The communication unit can be a communication interface, transceiver, transceiver circuit, etc., and the storage unit can be a memory.
[0031] The memory 220 can be volatile memory or non-volatile memory, or it can include both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0032] It is understood that the AI smart badge 110 may include more or fewer structural components than those shown in the structural block diagram above, such as a power module, physical buttons, a Wi-Fi module, a speaker, a Bluetooth module, sensors, and a display module, etc., without limitation. It is also understood that the AI smart badge 110 may be equipped with... Figure 1 The architecture of the aforementioned intelligent badge system.
[0033] After understanding the software and hardware architecture of this application, the following will be combined with... Figure 3 This application describes a method for generating medical record reports in an outpatient setting. Figure 3 This is a flowchart illustrating a method for generating medical record reports in an outpatient setting, provided in an embodiment of this application. The method is applied to a server of a smart badge system, which includes the server and smart badges communicating with the server. The method specifically includes the following steps: Step S310: Collect image and sound data of the consultation scene through the smart badge worn by the outpatient doctor.
[0034] The smart badge, also known as the AI smart badge, is a human-computer interaction data collection terminal device used in outpatient consultation scenarios. It integrates a camera module, a voice acquisition module, a processor, and a wireless communication module. The camera module continuously collects visual information from the consultation environment, while the voice acquisition module collects audio signals in real time during doctor-patient interactions. The image data is a sequence of visual data formed by the camera module continuously sampling the consultation scene at a preset time resolution; the audio data is a sequence of audio data obtained by the voice acquisition module sampling, converting, and digitizing the voice signals generated by the doctor and patient during the consultation. The consultation scene includes the doctor's consultation behavior, patient facial expressions, body movements, and background information of the doctor-patient interaction environment, which are not limited here. The consultation scene represents the complete state of the outpatient interaction process.
[0035] Specifically, after the outpatient doctor conducts the consultation process, the smart badge worn on the doctor's chest first uses its camera module to capture images of the patient and the consultation environment within the doctor's field of view at a fixed frame rate, forming a continuous video frame sequence. Simultaneously, the voice acquisition module simultaneously captures the voice signals generated by the doctor and patient during the consultation in full-duplex mode, obtaining raw audio data including the doctor's questions and the patient's responses. During the acquisition process, the smart badge uses a built-in time synchronization module to timestamp-align the image and audio data, establishing a one-to-one correspondence between each frame of image data and the corresponding audio data within the time window, thus ensuring the temporal consistency of subsequent multimodal information fusion processing. After completing the basic acquisition, the processor inside the smart badge performs preliminary preprocessing operations on the acquired raw image and audio data. The image data undergoes denoising and resolution normalization to eliminate the impact of ambient lighting changes and motion blur on visual information quality. The audio data undergoes pre-emphasis filtering and frame segmentation to enhance high-frequency features in the audio signal and reduce environmental noise interference, thereby improving the stability and accuracy of audio feature extraction. Then, the preprocessed image and audio data are labeled with time indices and stored in the local cache unit of the smart badge in chronological order. They are then uploaded to the server in real time via a wireless communication module for subsequent multimodal fusion analysis. Throughout the acquisition process, the smart badge continuously runs a multi-threaded acquisition mechanism, enabling the camera module and voice acquisition module to work in parallel, thus achieving synchronous capture of visual and audio information in the consultation scenario. Simultaneously, to ensure the integrity and continuity of data acquisition, the smart badge automatically activates a local caching mechanism to temporarily store data when it detects data transmission delays or network fluctuations, and performs retransmission processing after the network recovers, ensuring the complete consistency of image and audio data in the time dimension. Furthermore, in scenarios involving multiple patients in continuous consultations, the smart badge can also perform session-level differentiated storage of the acquired data based on the doctor's identification, thereby avoiding data aliasing between different consultation sessions and ensuring the accuracy and traceability of subsequent consultation data processing.
[0036] For easier understanding, please refer to Figure 4 , Figure 4 This is a schematic diagram of the appearance of an AI smart badge provided in an embodiment of this application. As can be seen, the AI smart badge adopts an integrated lightweight design and is a rounded rectangular structure adapted to be worn on the chest of medical staff. It integrates multiple functional modules such as sensing, interaction, and power supply, and provides a hardware carrier for barrier-free doctor-patient interaction and intelligent medical record generation in outpatient scenarios.
[0037] Specifically, 401 is the display module, deployed in the upper left area of the front of the smart badge. It uses a high-definition touch screen to display in real time the results of doctor-patient interaction translation, consultation status prompts, initial medical record report previews, patient feedback interaction interfaces, etc., realizing the visual output of multimodal translation information and providing intuitive interactive support for communication between doctors and patients with special needs (such as dialect speakers or hearing impaired individuals). 402 is the camera module, located to the right of the display module. It uses a wide-angle high-definition camera to capture in real time visual images of patients' facial expressions, gestures, body postures, etc. during consultations. The data provides high-quality visual feature input for sign language recognition and simultaneous intent translation type determination. 403 is a recording switch button, located adjacent to the camera module. It's a physical press-type start / stop switch used to manually control the working status of the voice acquisition module. Recording can be turned off during non-consultation periods or in private communication scenarios to ensure the privacy and security of patient medical data. 404 is a wireless charging magnetic interface located at the bottom of the front of the smart badge. It adopts a magnetic wireless charging design, supporting convenient wireless charging and meeting the fast charging needs of medical staff in outpatient work, while also improving the device's waterproof and dustproof performance. In addition, the hardware parameters of the smart badge are also marked, including device model, charging input specifications, and built-in battery capacity: "Model:FW920 Input:5V-1.0A Capacity:400mAh". Here, Model:FW920 is the device's exclusive hardware model, Input:5V-1.0A is the rated input specification of the AI smart badge, and the rated capacity of the built-in battery ensures stable battery life for a full workday in the outpatient department. It should be noted that the microphone array and speaker array are integrated into the side of the AI smart badge. Figure 4 (Not marked in the text) The microphone array is used to collect the voice signals of both doctors and patients and complete the front-end noise reduction preprocessing, while the speaker array is used to play the translated voice in real time, realizing the audible output of two-way voice / sign language translation, forming a complete voice interaction link.
[0038] In some possible embodiments, outpatient doctors wear AI smart badges on their chests. During consultations, the camera module 402 and the side microphone array simultaneously collect the patient's visual and speech data. After the backend server completes the intent translation, consultation dataset construction, and medical record generation, the translated text and draft medical record are displayed on the display module 401, and the translated speech is played by the speaker array, enabling barrier-free communication between doctors and patients with dialects or hearing impairments. Doctors can control the recording start and stop via the recording switch button 403 to avoid unnecessary data collection and protect patient privacy. If the AI smart badge's battery is low, it can be quickly recharged via a wireless charging magnetic interface without plugging or unplugging cables, ensuring all-day battery life. The display module 401 allows doctors to directly proofread and confirm medical records on the badge itself, eliminating the need for additional operating terminals and significantly improving diagnostic efficiency. It can be seen that the AI smart badge achieves integrated sensing, interaction, and power supply, ensuring portability, privacy, and practicality, and providing reliable hardware support for the implementation of intelligent outpatient diagnosis and treatment systems.
[0039] It should be noted that the acquired image and audio data are synchronously collected and aligned on the same timeline, ensuring that the data not only reflects the static information of the consultation scenario but also fully represents the dynamic changes during the doctor-patient interaction. Furthermore, the AI smart badge can directly upload video frame data collected by the camera module and voice communication data between the doctor and patient collected by the voice module to the server via StarFlash technology.
[0040] It is evident that by using smart badges to synchronously collect image and audio data during outpatient consultations, and combining this with timestamp alignment and a multimodal synchronous acquisition mechanism, visual and audio information during doctor-patient interactions can be fully acquired within a unified time dimension. This effectively improves the consistency and completeness of multimodal data, reduces information loss and semantic deviation caused by asynchronous acquisition, enhances the stability of subsequent intent-based synchronous translation and structured medical record generation, and ultimately improves the accuracy of overall consultation information processing and the automation level of medical record generation.
[0041] Step S320: Determine the type of synchronous translation of intent between the outpatient doctor and the patient based on the image data and the sound data. The type of synchronous translation of intent includes dialect one-way translation, dialect two-way translation, and sign language-speech two-way translation.
[0042] Among them, the intent-synchronized translation type is a communication adaptation mode determined based on the differences in language expression ability, language comprehension ability, and interaction perception ability between doctors and patients. It is used to characterize the semantic conversion direction and interaction adaptation method between doctors and patients in different types of communication barrier scenarios. The image data is the video frames of the doctor-patient pre-question and answer session captured by the camera in the AI smart badge device; the audio data is the pre-question and answer speech captured by the microphone of the AI badge; the intent-synchronized translation types of dialect one-way translation, dialect two-way translation, and sign language-voice two-way translation are adapted to three communication scenarios: patients who can understand Mandarin but only use dialect to express themselves, patients who do not understand Mandarin, and deaf and mute patients.
[0043] Specifically, firstly, multimodal features of speech, posture, and face are mapped to a 256-dimensional unified feature space through a fully connected layer. The weights of each modality feature are adaptively calculated and weighted fusion is completed through an attention mechanism to avoid the one-sidedness of single-modality judgment. Then, the fused features are input into a lightweight Transformer classification model, which outputs the matching probability distribution of three translation types. The optimal translation type is determined according to the maximum probability decision rule. At the same time, the confidence of the maximum matching probability is checked. If the probability is lower than 0.85, the AI badge display and voice secondary confirmation process are triggered to complete the translation type calibration. Finally, the determined translation type directly drives the subsequent bidirectional real-time translation, without the need for manual settings by doctors, keeping the consultation process continuous and smooth.
[0044] It is evident that by fusing image and audio dual-modal features, accurate identification of outpatient communication abilities and automatic matching of translation types are achieved, solving the problems of low efficiency and error-proneness in traditional manual determination of dialects and communication types of deaf and mute patients. Simultaneously, the confidence verification and secondary confirmation mechanisms significantly improve the accuracy of translation type matching, avoiding communication interruptions caused by pattern mismatches. The lightweight model and adaptive weight fusion design are adapted to the embedded computing power of AI badges, meeting the real-time interaction needs of outpatient clinics. This provides stable support for subsequent bidirectional translation between doctors and patients, structured extraction of medical semantics, and automatic generation of medical records, thereby improving the accuracy of information transmission and the efficiency of diagnosis and treatment in outpatient consultations.
[0045] In some possible embodiments, determining the type of synchronized translation of intent between the outpatient doctor and the patient based on the image data and the audio data specifically includes the following steps: 321. Visual features of facial expressions, gestures, and body postures are extracted from the image data to obtain first visual feature data; 322. Extract voiceprint, semantic, and vocalization features from the sound data to obtain sound feature data; 323. Based on a preset fusion formula, the first visual feature data and the sound feature data, feature fusion is performed to obtain multimodal features that characterize the intentions between outpatient doctors and patients; 324. Input the multimodal features into a preset intent translation model to obtain the intent synchronous translation result; 325. Determine the intent translation type between the outpatient doctor and the patient based on the intent synchronization translation result, and obtain the intent synchronization translation type.
[0046] Among them, the preset fusion formula is a multimodal weighted fusion rule based on the attention mechanism, which can map speech features and video frame features to a 256-dimensional unified feature space for weighted calculation; the intent translation model is a pre-trained lightweight Transformer classification model with a confidence threshold of 0.85, which is used to quantify the credibility of translation type matching.
[0047] Specifically, firstly, visual feature analysis is performed on standardized image data to extract key features of facial expressions, gestures, and body postures, forming the first visual feature data. Simultaneously, voiceprint, semantic, and vocalization feature analysis is performed on preprocessed audio data to extract 13-dimensional MFCC features, resulting in audio feature data. Next, a fully connected layer maps the two types of features to a unified feature space, and an attention mechanism is used to adaptively calculate weights and perform weighted fusion, generating multimodal features that comprehensively reflect the doctor-patient communication intent. Then, the multimodal features are input into a lightweight Transformer intent translation model, which outputs the matching probability distribution of three translation types, forming the intent synchronous translation result. Finally, the optimal translation type is selected according to the maximum probability rule, and its probability is verified with confidence. If it is higher than the threshold, it is directly determined; if it is lower than the threshold, patient feedback is obtained through a smart badge to complete the calibration, thereby determining the target intent synchronous translation type.
[0048] It is evident that the hierarchical extraction and adaptive fusion of dual-modal features can comprehensively capture the characteristics of patients' communication abilities, avoid the bias of single-modal judgment, and the lightweight model is adapted to the embedded computing power of smart badges to ensure real-time judgment. In addition, confidence verification and feedback calibration improve matching accuracy, avoid mode mismatch, and provide precise support for subsequent interactive information processing and consultation dataset construction, ensuring the smoothness and effectiveness of outpatient consultation communication.
[0049] In some possible embodiments, the feature fusion based on a preset fusion formula, the first visual feature data, and the sound feature data to obtain multimodal features representing the intention between the outpatient doctor and the patient specifically includes the following steps: 3231. Map the first visual feature data and the sound feature data to Hilbert space to obtain the first mapped feature and the second mapped feature; 3232. Extract the statistical features corresponding to the first mapping feature and the second mapping feature to obtain the first statistical feature and the second statistical feature; 3233. Determine the first weight and the second weight corresponding to the first mapping feature and the second mapping feature based on the first statistical feature and the second statistical feature using an attention mechanism; 3234. Perform feature fusion on the first mapping feature, the second mapping feature, the first weight, and the second weight to obtain the multimodal feature.
[0050] The first visual feature data consists of a visual feature set comprising 68 facial key points, 17 posture key points of the human body, and 21 hand key points; the voice feature data is speech processed with pre-emphasis, i.e. Frames with Hanming Window The 13-dimensional MFCC features extracted later, among which, This represents the sampled value of the pre-emphasis processed speech signal at time t; This represents the sampled value of the speech signal at the current time t; This represents the sampled value of the speech signal at the previous time step (t-1); This is the pre-emphasis coefficient, used to control the amplitude of high-frequency enhancement; The Hamming window function is represented at the th... The time-domain weighted amplitude of each sampling point; N represents the frame length. The Hilbert spatial mapping employs a linear transformation. To achieve the projection of heterogeneous features onto a 256-dimensional unified space, where, The output mapping feature is represented by W; the weight coefficient is represented by W. Indicates initial features; This represents the bias coefficient. The first and second mapped features are standardized features after mapping; the first and second statistical features are the mean μ and variance of the corresponding mapped features, respectively. It is used to quantify feature distribution and representation ability; the attention mechanism adopts an adaptive weight calculation rule. ( =1,2) to obtain the first weight Second weight ,satisfy =1; where, Indicates the first Attention weight coefficients for each modality; Indicates the first Attention score for each modality; Indicates the first Attention score for each modality; Indicates the number of modes.
[0051] Specifically, in the feature mapping stage, the first visual feature data and the audio feature data are used as the raw input. The first visual feature data includes the two-dimensional coordinates of 68 key points on the patient's face, the three-dimensional coordinates of 17 posture key points on the human body, and the spatial features of 21 hand key points of the deaf-mute patient. The audio feature data consists of 13-dimensional Mel-frequency cepstral coefficients extracted after pre-emphasis, frame segmentation, and windowing preprocessing. The two types of features are significantly heterogeneous in dimensionality, numerical range, and physical meaning. Therefore, a linear transformation formula is used. This is uniformly projected onto a 256-dimensional Hilbert space, using a learnable weight matrix W and bias vector. By completing feature dimension normalization and distribution normalization, we obtain first and second mapping features with consistent dimensions and numerical compatibility. The completeness of Hilbert space and the inner product operation characteristics can effectively preserve the core semantic information of the two types of features, while eliminating the differences in physical dimensions between modes, thus laying the foundation for subsequent fusion operations.
[0052] Next, in the statistical feature extraction stage, global statistical analysis is performed on the first and second mapping features based on the mean and variance formulas, respectively. The mean and variance are extracted to form the first and second statistical feature quantities. The mean is used to characterize the overall response strength of the feature, and the variance is used to measure the dispersion and anti-interference ability of the feature. These two statistical quantities can objectively quantify the information effectiveness of the visual and audio modalities in the current consultation frame, avoiding subjective bias caused by manually setting weights. In the attention weight calculation stage, the first and second statistical feature quantities are input into the adaptive attention module. The module first calculates the attention scores of the two types of mapping features through a fully connected layer. and Then, the weights are calculated using the softmax normalization formula to obtain the first weight. With the second weight This mechanism dynamically allocates weights based on the patient's actual communication status. When the patient is deaf or mute, visual features are dominant, so the first weight is automatically increased; when the patient speaks a dialect, vocal features are dominant, so the second weight is automatically strengthened, achieving the focusing of effective information and the suppression of noise. In the final feature fusion stage, a weighted fusion formula is used... ,in, This represents the fused multimodal features; The first and second mapping features are represented respectively. Then, a weighted summation operation is performed on the first and second mapping features to generate a 256-dimensional multimodal feature. This feature fully integrates the core information of visual and auditory bimodality, and can accurately represent the patient's language ability, communication style and interaction intention. Moreover, the entire fusion process adopts a lightweight computing design, and the feature fusion time per frame is less than 10 milliseconds. It can fully adapt to the computing power limitations of the smart badge embedded terminal and the real-time interaction requirements of outpatient consultation, ensuring that the feature fusion process has no significant delay and does not interfere with the normal consultation process.
[0053] In some possible embodiments, determining the intent translation type between the outpatient doctor and the patient based on the intent synchronization translation result, and obtaining the intent synchronization translation type, specifically includes the following steps: 3251. Based on the intent synchronization translation results, determine multiple intent translation types and multiple probability values corresponding to the multiple intent translation types; 3252. Determine the maximum probability value among the plurality of probability values; 3253. Determine the intent translation type corresponding to the maximum probability value to obtain the first intent translation type; 3254. Calculate the confidence level corresponding to the first intent translation type based on the multiple probability values to obtain the first confidence level; 3255. If the first confidence level is greater than or equal to a preset confidence level threshold, the first intent translation type is determined to be the intent synchronous translation type; 3256. If the first confidence level is less than the confidence level threshold, obtain the patient's feedback information; determine the intention translation information between the outpatient doctor and the patient based on the feedback information; determine the intention synchronization translation type based on the intention translation information and the multiple intention translation types.
[0054] The intent-synchronous translation result is the probability distribution P=[p1,p2,p3] of three intent translation types output by the pre-trained lightweight Transformer classification model. The three translation types are dialect one-way translation, dialect two-way translation, and sign language-voice two-way translation, respectively adapted to three communication scenarios: patients who can understand Mandarin but only use dialect, patients who do not understand Mandarin, and deaf-mute patients. The probability value is a normalized value in the 0~1 interval of the model output, used to quantify the degree of matching between each translation type and the patient's actual communication ability. The maximum probability is selected through the argmax(P) decision function, which is the optimal matching probability in the probability distribution. The first confidence level directly adopts this maximum probability, and the pre-set confidence threshold θ=0.85 is the quantitative standard for judging the reliability of the model output results. Patient feedback information is collected through the interaction between the front display screen of the smart badge and the speaker, including multimodal information such as voice response and gesture confirmation, for manual calibration of translation types in low confidence scenarios.
[0055] Specifically, firstly, multimodal features are input into the Transformer intent translation model. The model is fine-tuned and optimized based on outpatient multimodal data to adapt to the computing power limitations of the smart badge embedded terminal. After forward inference, it outputs probability values corresponding to three intent translation types, forming a probability distribution to comprehensively characterize the matching probability of each translation type. Then, based on the probability distribution results, through... The function filters out the highest probability value and marks the translation type corresponding to this maximum probability value as the first intention translation type, completing the initial optimal matching of translation types. Simultaneously, this maximum probability value is defined as the first confidence level of the first intention translation type, based on a preset confidence threshold of 0.85. When the first confidence level is greater than or equal to the confidence threshold, it indicates that the model's judgment of the current patient's communication ability is highly reliable, with no significant judgment bias. The first intention translation type can be directly determined as the final intention-synchronized translation type, requiring no doctor intervention or additional patient feedback, ensuring the continuous and smooth outpatient consultation process. When the first confidence level is less than the confidence threshold, it indicates that the model's judgment result has uncertainty and is prone to translation type mismatch risk. In this case, a secondary confirmation and calibration process is triggered. Standardized voice prompts are played through the smart badge's speaker, and a visual interactive interface displaying candidate translation types is shown on the screen. Patients are guided to provide clear feedback information through simple methods such as voice and gestures. By performing multimodal analysis and feature extraction on the feedback information, and combining it with the original probability distribution results, the translation type matching weights are readjusted, ultimately determining the intention-synchronized translation type that best matches the patient's actual communication ability. Throughout the entire decision-making process, the time taken for a single frame to make a full-process judgment is less than 200ms, which is suitable for the interactive needs of real-time outpatient consultations. At the same time, it covers communication judgment scenarios for special patient groups such as those with dialects or who are deaf and mute, which greatly improves the adaptability and stability of the system and further enhances the interaction efficiency between doctors and patients.
[0056] It is evident that the decision-making mechanism, which maximizes probability in initial selection, quantifies confidence thresholds for verification, and performs low-confidence interactive calibration, enables the determination of intent translation type. This mechanism not only ensures efficiency in outpatient scenarios based on lightweight model inference but also avoids pattern mismatch issues through human-computer interaction calibration, effectively improving the accuracy and robustness of translation type matching. The standardized decision-making process is deeply adapted to smart badge hardware and can be stably deployed in actual outpatient clinic scenarios. This provides accurate pattern support for subsequent doctor-patient interaction information processing, consultation dataset construction, and automatic medical record generation, eliminating the adverse effects of communication type determination bias on the treatment process and further improving the efficiency of outpatient doctors.
[0057] Step S330: Process the interaction information between the outpatient doctor and the patient according to the intent synchronous translation type to obtain processing result information; and create the patient's consultation dataset according to the interaction information and the processing result information.
[0058] Among them, the interactive information consists of doctor's voice signal, patient's voice signal, or sign language video frame signal of deaf and mute patients collected in real time by the AI badge device throughout the outpatient consultation process; the processing result information consists of bidirectional translated voice, translated text, and sign language digital human video sequence generated in the corresponding translation mode; the consultation dataset is a set of doctor-patient dialogue text with timestamps after time-series alignment.
[0059] Specifically, the corresponding model is invoked to perform streaming translation processing according to the intended translation type. If the patient can understand Mandarin, in the dialect-only translation mode, the patient's dialect speech is preprocessed and input into the translation model, outputting Mandarin speech and text, which is then played and pushed in real time. If the patient cannot understand Mandarin, in the dialect-only translation mode, the patient's dialect is simultaneously translated into Mandarin, and the doctor's Mandarin into the target dialect, with real-time audio and video playback and synchronized text display. In the sign language-speech bidirectional translation mode, 21 key hand features of the patient's sign language video are extracted to generate speech text, while the doctor's speech is converted into a sign language digital human video displayed on the name tag screen. The translated doctor-patient dialogue text is then time-aligned and redundancy removed, and integrated according to the consultation timeline to form a timestamped consultation dataset, completely preserving the entire interaction content.
[0060] It is evident that low-latency real-time processing of doctor-patient interaction information based on differentiated translation logic avoids communication barriers for patients with dialects or who are deaf or mute, ensuring the complete and accurate transmission of consultation information. Its fully automated processing requires no manual intervention from doctors, effectively reducing the communication and recording workload of doctors and significantly improving the efficiency of the entire outpatient consultation process.
[0061] In some possible embodiments, the patient's medical history dataset is created based on the interaction information and the processing result information, specifically including the following steps: 331. Extract the patient's voice data from the interaction information to obtain first patient voice data; and extract the outpatient doctor's voice data from the interaction information to obtain first doctor voice data; 332. If the intent-synchronized translation type in the processing result information is dialect one-way translation, then the first patient voice data is input into the preset intent-synchronized translation model to obtain patient translated voice data; semantic analysis is performed on the patient translated voice data and the first doctor voice data to obtain first semantic interaction result data; the consultation dataset is determined based on the first semantic interaction result data. 333. If the intent-synchronized translation type in the processing result information is dialect bidirectional translation, then the first doctor's voice data is input into the intent-synchronized translation model to obtain the second doctor's voice data; the second patient's voice data, which is the patient's response to the second doctor's voice data, is obtained and input into the intent-synchronized translation model to obtain the third patient's voice data; semantic analysis is performed on the third patient's voice data and the first doctor's voice data to obtain the second semantic interaction result data; the consultation dataset is determined based on the second semantic interaction result data. 334. If the intention-synchronized translation type in the processing result information is the sign language-voice bidirectional translation, then visual features are extracted from the image data to obtain second visual feature data; the second visual feature data is input into the intention-synchronized translation model to obtain first sign language expression data; third doctor voice data of the doctor's response to the first sign language expression data is obtained; the third doctor voice data is input into the intention-synchronized translation model to obtain the consultation dataset.
[0062] Among them, the intent-to-synchronize translation model is a set of pre-trained models adapted for outpatient scenarios, including dialects. Mandarin one-way translation model, Mandarin Dialect reverse translation model, sign language recognition Speech generation models and speech The sign language digital human generation model adopts a streaming inference architecture with a single-frame processing latency of no more than 200ms. The first patient voice data and the first doctor voice data are speech signals for patients and outpatient doctors separated from the interaction information, and are preprocessed after removing environmental noise and interference signals. The second visual feature is the hand key point feature extracted from the image data collected from deaf and mute patients, which is used for sign language semantic parsing. The consultation dataset is a collection of doctor-patient dialogue texts with timestamps after time alignment and redundancy removal. Semantic analysis is used to complete the time sequence association and initial screening of medical semantics of the dialogue text, ensuring the integrity and standardization of the dataset.
[0063] Specifically, firstly, the interactive information collected by the smart badges undergoes signal separation and noise reduction processing to extract the first patient voice data and the first doctor voice data, completing the separation of doctor-patient voice signals and providing input for subsequent modal translation. For dialect-based one-way translation, the first patient voice data corresponding to patients who can only express themselves in their dialect is input into the dialect. The Mandarin translation model generates patient audio data that can be directly understood by doctors in real time, simultaneously converting speech to text. This translated text is then semantically aligned and correlated with the corresponding text of the first doctor's audio data, forming a first semantic interaction result data containing complete consultation logic. This data is used to construct a consultation dataset adapted to this scenario. For dialect-based bidirectional translation, the first doctor's audio data is first input into Mandarin. A dialect translation model generates second doctor's voice data that can be understood by patients and plays it in real time. After collecting second patient voice data based on the doctor's response, a third patient voice data with standard Mandarin semantics is generated again through the translation model. The original doctor's speech text and the translated patient text are semantically integrated and temporally corrected to obtain second semantic interaction result data, thus forming a standardized consultation dataset. (This is also relevant for sign language.) The two-way audio translation type involves refining the visual features of image data from the consultation scenario to obtain second visual features representing the patient's sign language actions. These second visual features are then input into a sign language recognition model to generate first sign language data that can be interpreted by doctors. Third, the doctor's voice data, representing their responses to the sign language, is collected and processed into text through speech recognition and semantic conversion. The translated sign language text and the doctor's voice text are then integrated to create a consultation dataset tailored for deaf and mute patients. All three dataset creation scenarios utilize real-time streaming processing, completing data acquisition, translation, semantic analysis, and integration synchronously with the consultation interaction. This results in a standardized consultation dataset with accurate timestamps, complete content, and logical coherence. The entire process eliminates the need for manual recording and organization by doctors, maintaining the continuity and efficiency of the outpatient consultation process.
[0064] As can be seen, by executing this embodiment, the differentiated translation of interactive information in multiple scenarios and the automatic construction of datasets are realized, ensuring the integrity and standardization of consultation data and improving the level of intelligence in the entire outpatient diagnosis and treatment process.
[0065] Step S340: Generate the patient's target medical record report based on the consultation dataset.
[0066] The target medical record report is a standardized structured electronic medical record that can be directly connected to the Hospital Information System (HIS). The medical record is generated based on an entity recognition and relation extraction model finely tuned to a medical corpus, a standardized medical terminology knowledge base, and a fixed outpatient medical record template. The generation process includes medical entity extraction, confidence calibration, template filling, doctor proofreading, and model closed-loop optimization. The confidence filtering threshold for medical entities is set to 0.85. The medical record correction range is determined by Levenstein edit distance. The proofreading data can be used as incremental samples for model optimization.
[0067] Specifically, firstly, the consultation dataset is input into a medical entity recognition and relation extraction model. The medical entity recognition model can be BiLSTM-CRF, BioBERT, etc.; the medical relation extraction model can be BiLSTM-Attention, BAMRE, etc., or a joint model of entity recognition and relation extraction, such as the Biaffine-based Joint Model. This model automatically extracts medical entities with attributes of symptoms, body parts, medical history, and time through the entity recognition model and constructs a set of medical knowledge triples relating these entities. Next, the arithmetic mean of the confidence scores from multiple recognitions of the same entity is taken, filtering out low-confidence entities with confidence scores below 0.85. Simultaneously, a standardized medical terminology knowledge base is used to map patients' colloquial expressions to standardized medical terminology, ensuring professional and accurate medical record descriptions. Then, the patient's name, age, and consultation number from the HIS system are synchronized, and the data is processed according to the pre-set chief complaint and present illness history in the outpatient medical record. Fields such as past medical history, physical examination, and treatment recommendations are automatically populated into a template with standardized medical entities and triples to generate an initial structured medical record. The initial medical record is then pushed to the AI badge and outpatient client for doctors to proofread and correct. The difference between the initial and corrected medical records is calculated using an edit distance algorithm. If the difference exceeds the correction threshold, the corrected medical record and the corresponding consultation text are used as incremental training samples. The translation model and semantic extraction model are then fine-tuned using the low-rank adaptation (LoRA) method. Finally, the final medical record, after being confirmed by the doctor, is automatically archived into the HIS system, completing the entire medical record generation process.
[0068] It is evident that by extracting medical semantics in a structured manner and automatically filling templates, fragmented consultation data is transformed into standardized and unified structured medical records, solving the problems of unstructured, incomplete, and time-consuming processing of traditional outpatient medical records. Entity confidence calibration and terminology standardization significantly improve the accuracy and professionalism of medical record content, ensuring the reliability of clinical diagnosis and data reuse. The mechanism of doctor proofreading combined with model closed-loop optimization can continuously improve the accuracy of system generation and gradually reduce the workload of doctors in making corrections. Structured medical records are directly connected to the HIS system, enabling rapid archiving and efficient retrieval of medical records, greatly shortening the time doctors spend writing medical records, reducing the workload of outpatient clinics, and improving the efficiency of outpatient doctors.
[0069] In some possible embodiments, generating the patient's target medical record report based on the consultation dataset specifically includes the following steps: 341. Obtain the preset medical record report template; 342. Based on a preset entity relationship extraction model, extract medical entities and their corresponding relationships from the consultation dataset; 343. Construct a consultation interaction knowledge graph based on the medical entities and the relationships; 344. Determine the medical record entity corresponding to the medical entity based on the consultation interaction knowledge graph; and determine the medical record attribution relationship of the medical record entity based on the consultation interaction knowledge graph; 345. Fill the medical record entity and the medical record attribution relationship into the medical record report template to obtain the target medical record report.
[0070] The consultation dataset consists of a complete collection of doctor-patient dialogue texts aligned to time sequence and timestamped, including symptom descriptions, medical history communication, and treatment inquiries from outpatient consultations. The preset medical record report template is a standardized electronic medical record template compatible with the hospital's HIS, including fixed structured fields such as chief complaint, present illness, past medical history, physical examination, and preliminary treatment suggestions, enabling direct medical record archiving and data interoperability. The consultation interaction knowledge graph is a topological structure with medical entities as nodes and relationships as directed edges, fully representing the inherent semantic logic of consultation information. Medical record entities are standardized medical expressions mapped from a standardized medical terminology knowledge base, eliminating colloquial and unprofessional descriptions. The medical record attribution relationship is the correspondence between medical record entities and template fields, ensuring accurate entity matching to medical record categories. The medical entity confidence filtering threshold is 0.85, used to filter low-confidence recognition results, ensuring the accuracy and professionalism of medical record content.
[0071] Specifically, firstly, through the communication interface with the hospital's HIS server, standardized outpatient medical record report templates conforming to clinical norms are retrieved in real time. Next, the consultation dataset is input into a medical entity relationship extraction model. The model parses the dialogue text frame by frame, automatically identifying and extracting four types of core medical entities. The model calculates the arithmetic average of the confidence scores of multiple identifications of the same entity, filtering out low-reliability entities with confidence scores below 0.85. Simultaneously, combined with a standardized medical terminology knowledge base, patients' colloquial expressions are mapped to professional medical terms, and semantic relationships between entities are extracted concurrently, forming a complete set of medical knowledge triples. Then, using standardized medical entities as graph nodes and entity relationships as directed edges, a consultation interaction knowledge graph is constructed. This graph clearly restores the semantic relationships and diagnostic logic of the consultation information, avoiding fragmented medical record content gaps caused by fragmented information. Then, based on the topological structure and semantic relationships of the knowledge graph, medical entities are transformed into medical record entities suitable for filling out medical records. Simultaneously, according to outpatient medical record field specifications and treatment logic, the medical record affiliation relationship corresponding to each entity is determined, clarifying the target template fields that each entity should fill, ensuring the logical rationality and field matching degree of the medical record content. Finally, according to the medical record affiliation relationship, various medical record entities are accurately filled into the corresponding fields of the preset template, automatically completing the splicing and integration of structured medical records, generating target medical record reports that meet clinical archiving requirements. Only a small amount of proofreading by doctors is needed to complete the medical record archiving, adapting to the needs of efficient outpatient diagnosis and treatment scenarios.
[0072] For easier understanding, please refer to Figure 5 , Figure 5 This is a schematic diagram of the structure of a case report provided in an embodiment of this application. As can be seen, the case report is a standardized electronic medical record for outpatients automatically generated based on an intelligent name tag system. The case report includes a medical record title bar, a patient basic information bar, a medical record core content bar, and a diagnosis and treatment bar. It fully covers key medical record elements such as chief complaint, present medical history, physical examination, preliminary diagnosis, and treatment suggestions for outpatient diagnosis and treatment. It is a standardized medical record document that is archived by connecting to HIS.
[0073] Specifically, Figure 5The top label of the electronic medical record is the title bar of the medical record document; below the title bar, the patient's basic information and the consultation time information are displayed in sequence. Among them, the patient: the test patient is the main information of the consultation corresponding to the medical record, the outpatient serial number: 0126 is a unique consultation identifier assigned by the hospital outpatient system, which is used to realize the accurate traceability and data association of the patient's medical records, and the time: 2026-1-26 15:31 is the specific timestamp of the consultation completion, reflecting the time node of the medical record generation, and providing a basis for the timeliness management of medical records and the time sequence analysis of medical data. The medical record content section is divided into several structured items according to the outpatient medical record writing standards: "Chief Complaint: Severe pain in the lower left right teeth at night (1 item)" is a description of the patient's symptoms, reflecting the main discomfort symptoms during this visit, and is the result of the consultation interaction data; "Present Illness, Personal History, Family History: Poor oral hygiene habits" is a comprehensive description of the patient's current condition, personal lifestyle habits, and family medical history. This content is derived from standardized expressions obtained after semantic analysis and medical entity extraction of doctor-patient consultation interaction information; "Drug Allergy History: Unknown" records information related to the patient's drug allergies; "Physical Examination: Poor oral hygiene, plaque, gingival redness and swelling" is... The physical examination results of the patient's oral condition, including the degree of oral hygiene, the amount of plaque buildup, and specific signs such as gingival redness and swelling, are medical entity information extracted from the consultation interaction data by the medical entity relationship extraction model. "Auxiliary Examination: To be supplemented" indicates the next steps in the diagnosis and treatment process, suggesting further examinations needed at the current consultation stage, providing guidance for the doctor's subsequent treatment. "Preliminary Diagnosis: Deciduous Tooth Caries" is a preliminary clinical diagnosis conclusion based on the patient's chief complaint and physical examination information. It is a medical record entity determined by the server through reasoning using the consultation interaction knowledge graph, completing the semantic transformation and logical association from consultation data to clinical diagnosis. The "Treatment: 1. Condition and treatment plan have been informed; 2. X-ray recommended" section at the bottom of the medical record outlines the treatment plan for the patient's condition, including information about the condition and treatment recommendations, fully reflecting the outpatient treatment process and is an important component of the consultation dataset generation and medical record report filling.
[0074] It can be seen that, Figure 5The electronic medical record report shown uses AI-powered smart badges to collect multimodal interaction information between doctors and patients. The server then performs feature extraction, intent translation type determination, and consultation dataset construction. Following medical entity relationship extraction, knowledge graph construction, and medical record template filling, a complete electronic medical record is generated, containing the patient's basic information, chief complaint, present illness history, physical examination, preliminary diagnosis, and treatment plan. This medical record achieves efficient conversion of unstructured consultation interaction data into structured medical text, preserving core clinical information from outpatient treatment while conforming to the hospital's HIS system's medical record archiving standards. It provides a high-quality, standardized medical record carrier for the storage, retrieval, analysis, and subsequent clinical research of medical data. Meanwhile, through standardized template support, intelligent extraction of medical entities, semantic integration of knowledge graphs, and automated field filling, the structured automatic generation of outpatient medical records is realized, transforming unstructured consultation dialogues into standardized and complete electronic medical records. This solves the problems of low efficiency, unprofessional expression, and fragmented information in traditional medical records. The fully automated processing significantly shortens the time doctors spend writing medical records and reduces the workload of outpatient diagnosis and treatment. The generated structured medical records can be directly connected to the HIS system, providing high-quality standardized data support for clinical diagnosis and treatment and medical big data analysis, while improving the efficiency of outpatient doctors and patients.
[0075] In some possible embodiments, filling the medical record entity and the medical record attribution relationship into the medical record report template to obtain the target medical record report specifically includes the following steps: 3451. Based on the medical record attribution relationship, map the medical record entity to the medical record report template to obtain the first medical record report; 3452. Send the first medical record report to the server and obtain the medical record feedback information from the doctor's side in response to the first medical record report; 3453. Adjust the first medical record report based on the medical record feedback information to obtain the second medical record report; 3454. Determine the semantic similarity parameters between the first medical record report and the second medical record report; 3455. When the semantic similarity parameter is greater than or equal to a preset similarity threshold, the first medical record report and the second medical record report are processed based on a preset text difference calculation model to obtain the target medical record report.
[0076] The first medical record report is the automatically generated initial structured medical record, and the second medical record report is the optimized medical record after being proofread and corrected by the doctor. Medical record feedback information includes correction, addition, and deletion instructions submitted by the doctor through the outpatient client and the smart badge front end. The semantic similarity parameter uses Levinstein edit distance quantization to characterize the content differences between the two medical records; the calculation formula is as follows: .
[0077] in: This represents the Levenstein edit distance value (i.e., semantic similarity). This represents the Levinstein edit distance function; This indicates the initial electronic medical record, i.e., the first electronic medical record report; This refers to the proofread electronic medical record report, also known as the second electronic medical record report.
[0078] The preset similarity threshold is a correction magnitude threshold δ=3. The text difference calculation model is a difference extraction and text fusion model based on edit distance. The difference data can be used as incremental samples to complete the model closed-loop optimization through the low-rank adaptation (LoRA) method. The target medical record report is the final standardized medical record that can be directly archived to the HIS system.
[0079] Specifically, firstly, based on the attribution of medical records, standardized medical record entities, extracted through medical semantics and integrated with a knowledge graph, are precisely mapped to the corresponding columns of the medical record report template according to field matching logic. This automatically fills in the chief complaint, present illness history, and other content, generating a first medical record report with a standardized format and complete information. This report fully complies with the medical industry's medical record writing standards and HIS system data standards. Next, the first medical record report is uploaded to the server via a wireless communication module. The server simultaneously pushes the report to the outpatient client and the smart badge front-end display. Doctors review the initial medical record based on their clinical experience and submit feedback information regarding issues such as expression deviations, missing information, and content errors. The system collects and transmits this feedback data back to the server in real time. Then, the server performs targeted corrections on the first medical record report based on the feedback information, including replacing erroneous content, supplementing key information, and standardizing colloquial expressions, generating a second medical record report that is more in line with clinical practice. Based on this, the Levenstein edit distance algorithm is used to calculate the semantic similarity parameters of the two medical records. A formula quantifies the minimum number of operations required for character replacement, insertion, and deletion, using objective numerical values to characterize the extent of medical record correction and avoiding subjective bias from manual judgment. Finally, the semantic similarity parameters are compared with a preset threshold δ=3. When the parameters are greater than or equal to the threshold, it indicates that the correction extent meets the model optimization conditions. The text difference calculation model extracts the differences between the two medical records, and the second medical record report is identified as the target medical record report. Simultaneously, the difference data and the corresponding consultation dataset are used as incremental samples, and the low-rank adaptation (LoRA) method is used to perform lightweight fine-tuning of the translation model and entity extraction model, forming a technical closed loop of "generation-proofreading-optimization". If the parameters are below the threshold, the second medical record report is directly archived as the target medical record report to ensure the real-time generation of outpatient medical records. Simultaneously, the target medical record report is uploaded to the HIS system for archiving, and can be directly used for clinical diagnosis and treatment, medical record retrieval, and medical big data analysis.
[0080] It is evident that the progressive processing of automatic filling, manual proofreading, difference quantification, and closed-loop optimization significantly improves the efficiency of medical record generation while ensuring the clinical accuracy and professionalism of medical record content through physician proofreading. The Levinstein edit distance enables objective quantification of the correction magnitude, providing a precise basis for model optimization. The lightweight fine-tuning mechanism based on incremental samples continuously improves the accuracy of system translation and entity extraction, gradually reducing the workload of physician proofreading. The standardized generation and archiving process is deeply adapted to the hospital HIS system, realizing the digital and standardized management of outpatient medical records, providing high-quality support for medical data reuse and clinical research, reducing the burden of medical documents for outpatient physicians, and thus improving the efficiency of outpatient services.
[0081] For easier understanding, please refer to Figure 6 , Figure 6This is a backend interface diagram of an intelligent name tag system provided in an embodiment of this application. As can be seen, this interface is a visual interactive portal for the backend management of the server in the intelligent name tag system. It adopts a modular and layered design and integrates modules such as system status monitoring, function operation, and operation and maintenance data visualization analysis. It provides hospital administrators with a unified and convenient system operation and maintenance and data management entry point, realizing the control of AI intelligent name tags. Specifically, the left side of the interface is the system function navigation bar, which sequentially includes function entrances such as the workbench, patient list, AI smart badge, permission management, and system settings. The workbench is the default homepage, used to display an overview of system operation. The patient list entrance is used to retrieve patient medical records associated with consultation data collected by the smart badge, enabling traceability and linkage between consultation interaction records and basic patient information. The AI smart badge entrance is used to perform maintenance operations such as remote device configuration, firmware upgrades, status inspections, and parameter distribution. The permission management entrance is used to assign system operation permissions to different roles (such as system administrators, outpatient doctors, and maintenance personnel), controlling the scope of access to medical data based on the principle of least privilege to ensure patient privacy and data security. The system settings entrance is used to configure system communication protocols, data storage rules, and hospital information system (HIS) interface adaptation parameters, achieving seamless integration between the smart badge system and the hospital's existing information system. The right side of the interface's workbench area is divided into three functional modules: statistical data, frequently used functions, and recording trends for the past 15 days. The statistics display operational metrics for the total number of devices, the number of online devices, and the total number of audio files. The total number of devices refers to the total number of AI smart badge 110 units deployed throughout the hospital. The number of online devices refers to the number of devices currently connected and working normally, used to monitor device online rate and operational health status in real time. The total number of audio files refers to the total number of files such as doctor-patient consultation voice and translated text stored in the system, used to quantify the scale of the system's consultation service and data storage capacity. Common functions include device management, audio files, doctor management, and device logs. Device management is used to perform maintenance operations such as binding, unbinding, remote restarting, and troubleshooting of AI smart badges. Recording files are used to retrieve, play back, and export consultation interaction recordings and transcribed texts, enabling full lifecycle traceability of medical data. Doctor management is used to maintain information on medical staff using smart badges and the binding relationship between associated devices and doctors. Device logs are used to record information such as the operating status of AI smart badges, communication logs, and fault alarms, providing data support for device maintenance. The recording trend module for the past 15 days uses a curve graph to visually display the changing trend of the daily recording file generation volume of the system over the past 15 days. The vertical axis is marked with four levels of quantity scales: 20, 30, 40, and 50, which intuitively reflects the fluctuation pattern of outpatient consultation service volume and provides data basis for hospital outpatient scheduling optimization and equipment maintenance scheduling.
[0082] As can be seen, the backend management interface realizes the visualized control of the intelligent badge system, which not only meets the needs of hospital administrators for equipment operation and maintenance, data security, and access control, but also provides data support for hospital operation optimization through data trend analysis, protects patient privacy through access management and audio file tracing, and provides a complete operation and maintenance management solution for the long-term stable operation and service quality improvement of the outpatient intelligent diagnosis and treatment system. It achieves deep integration of technical solutions with the actual operation and maintenance needs of the hospital and improves the overall efficiency of the outpatient department.
[0083] As can be seen, the above-described method for generating medical record reports for outpatient scenarios first involves collecting image and audio data of the consultation scenario using the smart badges worn by outpatient doctors; secondly, determining the type of simultaneous translation of intent between the outpatient doctor and the patient based on the image and audio data; thirdly, processing the interaction information between the outpatient doctor and the patient based on the type of simultaneous translation of intent to obtain processing result information; fourthly, creating a patient's consultation dataset based on the interaction information and processing result information; and finally, generating the patient's target medical record report based on the consultation dataset. Among them, the synchronous translation types include dialect one-way translation, dialect two-way translation, and sign language-speech two-way translation. Specifically, dialect one-way translation can adapt to communication barriers where patients speak dialects that doctors cannot understand; dialect two-way translation can adapt to communication barriers where patients speak dialects that doctors cannot understand; and sign language-speech two-way translation can adapt to communication barriers where patients express themselves in sign language that doctors cannot understand. This can comprehensively and flexibly adapt to various communication barriers in actual consultation scenarios, improving the comprehensiveness, intelligence, and scenario applicability of the system in processing communication information in consultation scenarios. In addition, the consultation dataset created based on interaction information and processing results information covers as comprehensive an unstructured consultation content as possible, improving the comprehensiveness and intelligence of the system in creating medical record reports.
[0084] The above mainly describes the solutions of the embodiments of this application from the perspective of the method execution process. It is understood that, in order to achieve the above functions, the smart badge includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments provided herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0085] This application embodiment can divide the smart badge into functional units according to the above method example. For example, each function can be divided into a separate functional unit, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0086] When dividing each function into modules according to its corresponding function. Figure 7 This is a functional module block diagram of a medical record report generation device for outpatient scenarios provided in an embodiment of this application. The medical record report generation device 700 for outpatient scenarios is applied to a server of a smart badge system. The smart badge system includes the server and smart badges communicating with the server. The medical record report generation device 700 for outpatient scenarios includes: The acquisition unit 710 is used to collect image and sound data of the consultation scene through the smart badge worn by the outpatient doctor; The calculation unit 720 is used to determine the intention synchronization translation type between the outpatient doctor and the patient based on the image data and the sound data. The intention synchronization translation type includes dialect one-way translation, dialect two-way translation, and sign language-speech two-way translation. The calculation unit 720 is used to process the interaction information between the outpatient doctor and the patient based on the intention synchronization translation type to obtain processing result information. The control unit 730 is configured to create a patient's medical history dataset based on the interaction information and the processing result information; and generate a target medical record report for the patient based on the medical history dataset.
[0087] In one possible embodiment, the computing unit 720, in determining the type of synchronized translation of intent between the outpatient doctor and the patient based on the image data and the sound data, is specifically configured to: Visual features of facial expressions, gestures, and body postures are extracted from the image data to obtain first visual feature data; The sound data is subjected to voiceprint, semantic and vocalization sound feature extraction to obtain sound feature data; Based on a preset fusion formula, the first visual feature data and the sound feature data, feature fusion is performed to obtain multimodal features that characterize the intentions between outpatient doctors and patients. The multimodal features are input into a preset intent translation model to obtain the intent synchronous translation result; Based on the intent synchronization translation results, the intent translation type between the outpatient doctor and the patient is determined, and the intent synchronization translation type is obtained.
[0088] In one possible embodiment, the computing unit 720, in performing feature fusion based on a preset fusion formula, the first visual feature data, and the sound feature data to obtain multimodal features characterizing the intention between the outpatient doctor and the patient, is specifically used for: The first visual feature data and the sound feature data are mapped to Hilbert space to obtain the first mapped feature and the second mapped feature. Extract the statistical features corresponding to the first mapping feature and the second mapping feature to obtain the first statistical feature and the second statistical feature; The first weight and the second weight corresponding to the first mapping feature and the second mapping feature are determined by the attention mechanism based on the first statistical feature and the second statistical feature. The first mapping feature, the second mapping feature, the first weight, and the second weight are fused to obtain the multimodal feature.
[0089] In one possible embodiment, the computing unit 720, in determining the intent translation type between the outpatient doctor and the patient based on the intent synchronization translation result, and obtaining the intent synchronization translation type, is specifically configured to: Based on the intent synchronization translation results, multiple intent translation types and multiple probability values corresponding to the multiple intent translation types are determined; Determine the maximum probability value among the plurality of probability values; Determine the intent translation type corresponding to the maximum probability value to obtain the first intent translation type; The confidence level corresponding to the first intent translation type is calculated based on the multiple probability values to obtain the first confidence level; If the first confidence level is greater than or equal to the preset confidence level threshold, the first intent translation type is determined to be the intent synchronous translation type; If the first confidence level is less than the confidence threshold, obtain the patient's feedback information; determine the intention translation information between the outpatient doctor and the patient based on the feedback information; determine the intention synchronization translation type based on the intention translation information and the multiple intention translation types.
[0090] In one possible embodiment, the control unit 730, in creating the patient's consultation dataset based on the interaction information and the processing result information, is specifically configured to: The patient's voice data is extracted from the interaction information to obtain first patient voice data; and the outpatient doctor's voice data is extracted from the interaction information to obtain first doctor voice data; If the intent-synchronized translation type in the processing result information is dialect one-way translation, then the first patient voice data is input into the preset intent-synchronized translation model to obtain patient translated voice data; semantic analysis is performed on the patient translated voice data and the first doctor voice data to obtain first semantic interaction result data; the consultation dataset is determined based on the first semantic interaction result data. If the intent-synchronized translation type in the processing result information is dialect bidirectional translation, then the first doctor's voice data is input into the intent-synchronized translation model to obtain the second doctor's voice data; the second patient's voice data, which is the patient's response to the second doctor's voice data, is obtained and input into the intent-synchronized translation model to obtain the third patient's voice data; semantic analysis is performed on the third patient's voice data and the first doctor's voice data to obtain the second semantic interaction result data; the consultation dataset is determined based on the second semantic interaction result data. If the intent synchronization translation type in the processing result information is the sign language-voice bidirectional translation, then visual features are extracted from the image data to obtain second visual feature data; the second visual feature data is input into the intent synchronization translation model to obtain first sign language expression data; third doctor voice data of the doctor's response to the first sign language expression data is obtained; the third doctor voice data is input into the intent synchronization translation model to obtain the consultation dataset.
[0091] In one possible embodiment, the control unit 730, in generating the patient's target medical record report based on the consultation dataset, is specifically configured to: Retrieve preset medical record report templates; Based on a preset entity relationship extraction model, medical entities and their corresponding relationships are extracted from the consultation dataset. Construct a consultation interaction knowledge graph based on the medical entities and the relationships; Based on the consultation interaction knowledge graph, the medical entity corresponding to the medical entity is determined; and the medical record attribution relationship of the medical record entity is determined based on the consultation interaction knowledge graph. The medical record entity and the medical record attribution relationship are filled into the medical record report template to obtain the target medical record report.
[0092] In one possible embodiment, the control unit 730 is specifically configured to: fill the medical record entity and the medical record attribution relationship into the medical record report template to obtain the target medical record report. Based on the medical record attribution relationship, the medical record entity is mapped to the medical record report template to obtain the first medical record report; The first medical record report is sent to the server, and the doctor's response to the first medical record report is obtained; The first medical record report is adjusted based on the medical record feedback information to obtain the second medical record report; Determine the semantic similarity parameters between the first medical record report and the second medical record report; When the semantic similarity parameter is greater than or equal to a preset similarity threshold, the first medical record report and the second medical record report are processed based on a preset text difference calculation model to obtain the target medical record report.
[0093] As can be seen, this embodiment provides a medical record report generation device for outpatient scenarios. This device collects image and sound data from the consultation scene using a smart badge worn by outpatient doctors. Based on the image and sound data, it determines the type of synchronous translation of intent between the outpatient doctor and the patient. This type of synchronous translation includes one-way dialect translation, two-way dialect translation, and sign language-speech two-way translation. It processes the interaction information between the outpatient doctor and the patient according to the synchronous translation type to obtain processing result information. Based on the interaction information and processing result information, it creates a patient's consultation dataset. Finally, it generates the patient's target medical record report based on the consultation dataset. This improves the processing efficiency of outpatient services.
[0094] This application also provides a computer-readable storage medium storing a computer program for electronic data interchange, which causes a computer to perform some or all of the steps of any of the methods described in the above method embodiments, wherein the computer includes a smart badge.
[0095] This application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps of any of the methods described in the above method embodiments. The computer program product may be a software installation package, and the computer may include a smart badge.
[0096] It should be noted that, for the sake of simplicity, the above embodiments are all described as a series of actions. Those skilled in the art should understand that this application is not limited to the described order of actions, as some steps in the embodiments of this application can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions, steps, modules, or units involved are not necessarily essential to the embodiments of this application.
[0097] In the above embodiments, the descriptions of each embodiment in this application have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0098] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
[0099] The steps of the methods or algorithms described in the embodiments of this application can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in RAM, flash memory, ROM, EPROM, electrically erasable programmable read-only memory (EEPROM), registers, hard disk, portable hard disk, read-only optical disk (CD-ROM), or any other form of storage medium well known in the art. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Furthermore, the ASIC can reside in a terminal device or management device. Alternatively, the processor and storage medium can exist as discrete components in the terminal device or management device.
[0100] Those skilled in the art will recognize that, in one or more of the examples above, the functions described in the embodiments of this application can be implemented, in whole or in part, by software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital video discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).
[0101] The various devices and products described in the above embodiments include modules / units that can be software modules / units, hardware modules / units, or a combination of both. The specific embodiments described above further illustrate the purpose, technical solutions, and beneficial effects of the embodiments of this application. It should be understood that the above descriptions are merely specific implementations of the embodiments of this application and are not intended to limit the scope of protection of the embodiments of this application. Any modifications, equivalent substitutions, improvements, etc., made based on the technical solutions of the embodiments of this application should be included within the scope of protection of the embodiments of this application.
Claims
1. A method for generating medical record reports in outpatient settings, characterized in that, A server for an intelligent badge system, the intelligent badge system including the server and intelligent badges communicating with the server, the method comprising: The smart badges worn by outpatient doctors collect image and sound data from the consultation process. Based on the image data and the sound data, the type of synchronous translation of intent between the outpatient doctor and the patient is determined. The type of synchronous translation of intent includes one-way dialect translation, two-way dialect translation, and sign language-speech two-way translation. The interaction information between the outpatient doctor and the patient is processed according to the intent synchronous translation type to obtain processing result information; and the patient's consultation dataset is created based on the interaction information and the processing result information. Generate the patient's target medical record report based on the consultation dataset.
2. The method as described in claim 1, characterized in that, The step of determining the type of synchronized translation of intent between the outpatient doctor and the patient based on the image data and the audio data includes: Visual features of facial expressions, gestures, and body postures are extracted from the image data to obtain first visual feature data; The sound data is subjected to voiceprint, semantic and vocalization sound feature extraction to obtain sound feature data; Based on a preset fusion formula, the first visual feature data and the sound feature data, feature fusion is performed to obtain multimodal features that characterize the intentions between outpatient doctors and patients. The multimodal features are input into a preset intent translation model to obtain the intent synchronous translation result; Based on the intent synchronization translation results, the intent translation type between the outpatient doctor and the patient is determined, and the intent synchronization translation type is obtained.
3. The method as described in claim 2, characterized in that, The feature fusion based on the preset fusion formula, the first visual feature data, and the sound feature data yields multimodal features representing the intentions between outpatient doctors and patients, including: The first visual feature data and the sound feature data are mapped to Hilbert space to obtain the first mapped feature and the second mapped feature. Extract the statistical features corresponding to the first mapping feature and the second mapping feature to obtain the first statistical feature and the second statistical feature; The first weight and the second weight corresponding to the first mapping feature and the second mapping feature are determined by the attention mechanism based on the first statistical feature and the second statistical feature. The first mapping feature, the second mapping feature, the first weight, and the second weight are fused to obtain the multimodal feature.
4. The method as described in claim 2, characterized in that, The step of determining the intent translation type between the outpatient doctor and the patient based on the intent synchronization translation result, and obtaining the intent synchronization translation type, includes: Based on the intent synchronization translation results, multiple intent translation types and multiple probability values corresponding to the multiple intent translation types are determined; Determine the maximum probability value among the plurality of probability values; Determine the intent translation type corresponding to the maximum probability value to obtain the first intent translation type; The confidence level corresponding to the first intent translation type is calculated based on the multiple probability values to obtain the first confidence level; If the first confidence level is greater than or equal to the preset confidence level threshold, the first intent translation type is determined to be the intent synchronous translation type; If the first confidence level is less than the confidence threshold, obtain the patient's feedback information; determine the intention translation information between the outpatient doctor and the patient based on the feedback information; determine the intention synchronization translation type based on the intention translation information and the multiple intention translation types.
5. The method according to any one of claims 1-4, characterized in that, Creating the patient's consultation dataset based on the interaction information and the processing result information includes: The patient's voice data is extracted from the interaction information to obtain first patient voice data; and the outpatient doctor's voice data is extracted from the interaction information to obtain first doctor voice data; If the intent-synchronized translation type in the processing result information is dialect one-way translation, then the first patient voice data is input into the preset intent-synchronized translation model to obtain patient translated voice data; semantic analysis is performed on the patient translated voice data and the first doctor voice data to obtain first semantic interaction result data; the consultation dataset is determined based on the first semantic interaction result data. If the intent-synchronized translation type in the processing result information is dialect bidirectional translation, then the first doctor's voice data is input into the intent-synchronized translation model to obtain the second doctor's voice data; the second patient's voice data, which is the patient's response to the second doctor's voice data, is obtained and input into the intent-synchronized translation model to obtain the third patient's voice data; semantic analysis is performed on the third patient's voice data and the first doctor's voice data to obtain the second semantic interaction result data; the consultation dataset is determined based on the second semantic interaction result data. If the intent synchronization translation type in the processing result information is the sign language-voice bidirectional translation, then visual features are extracted from the image data to obtain second visual feature data; the second visual feature data is input into the intent synchronization translation model to obtain first sign language expression data; third doctor voice data of the doctor's response to the first sign language expression data is obtained; the third doctor voice data is input into the intent synchronization translation model to obtain the consultation dataset.
6. The method according to any one of claims 1-4, characterized in that, The step of generating the patient's target medical record report based on the consultation dataset includes: Retrieve preset medical record report templates; Based on a preset entity relationship extraction model, medical entities and their corresponding relationships are extracted from the consultation dataset. Construct a consultation interaction knowledge graph based on the medical entities and the relationships; Based on the consultation interaction knowledge graph, the medical entity corresponding to the medical entity is determined; and the medical record attribution relationship of the medical record entity is determined based on the consultation interaction knowledge graph. The medical record entity and the medical record attribution relationship are filled into the medical record report template to obtain the target medical record report.
7. The method as described in claim 6, characterized in that, The step of filling the medical record entity and the medical record attribution relationship into the medical record report template to obtain the target medical record report includes: Based on the medical record attribution relationship, the medical record entity is mapped to the medical record report template to obtain the first medical record report; The first medical record report is sent to the server, and the doctor's response to the first medical record report is obtained; The first medical record report is adjusted based on the medical record feedback information to obtain the second medical record report; Determine the semantic similarity parameters between the first medical record report and the second medical record report; When the semantic similarity parameter is greater than or equal to a preset similarity threshold, the first medical record report and the second medical record report are processed based on a preset text difference calculation model to obtain the target medical record report.
8. A medical record report generation device for outpatient settings, characterized in that, A server for an intelligent name tag system, the intelligent name tag system including the server and intelligent name tags communicating with the server, and the medical record report generation device including: The acquisition unit is used to collect image and sound data of the consultation scene through the smart badge worn by the outpatient doctor; The computing unit is used to determine the type of synchronous translation of intent between the outpatient doctor and the patient based on the image data and the sound data. The type of synchronous translation of intent includes dialect one-way translation, dialect two-way translation, and sign language-speech two-way translation. The unit processes the interaction information between the outpatient doctor and the patient according to the type of synchronous translation of intent to obtain processing result information. The control unit is configured to create a patient's medical history dataset based on the interaction information and the processing result information; and to generate a target medical record report for the patient based on the medical history dataset.
9. A smart name tag, characterized in that, include: Processor, memory, communication interface, and one or more programs; The one or more programs are stored in the memory and configured to be executed by the processor, the programs including instructions for performing the steps of the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program including program instructions that, when executed by a processor, cause the processor to perform the method as described in any one of claims 1-7.