Medical information generation method and system based on model processing
By detecting voice signals and generating text analysis results through smart watches, the problems of low efficiency and cross-infection in medical document recording are solved, and real-time and efficient document entry and understanding of medical conditions are achieved.
Patent Information
- Application Number
- CN202510838563.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-19
AI Technical Summary
The existing medical document recording method is inefficient, takes up core diagnosis and treatment time, and poses a risk of cross-infection. Ward rounds are time-consuming and cannot effectively improve medical work efficiency.
A smart watch is used to detect voice signals, determine the direction of the sound source, collect video information, and generate text analysis results through a text analysis model, which are integrated into the smart watch and transmitted to the computer.
It realizes the real-time generation of medical information, improves the efficiency of document entry, reduces human errors and information loss, and reduces the risk of cross infection.
Smart Images

Figure CN120673958A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of text analysis technology, and in particular to a method and system for generating medical information based on model processing. Background Art
[0002] In the medical field, doctors and nurses frequently involve the process of documenting in their daily work. The current method of documenting relies on manual input, which has the following problems: Low efficiency of medical record writing: Doctors often need to spend a lot of time on documenting, squeezing out core diagnosis and treatment time, such as asking about the condition, physical examination, doctor-patient communication, etc.; when nurses take documents at the bedside, they need to input word for word and face the risk of cross-infection. Time-consuming ward rounds: When doctors make ward rounds at the bedside, in addition to asking about the condition and conducting physical examinations, they also need to record relevant data in documents. Traditional handwriting methods are time-consuming, and the use of electronic devices cannot completely eliminate this problem, especially in a busy clinical environment.
[0003] With the increasing demand for medical services, how to improve medical work efficiency, reduce time waste and return time to patients has become an urgent problem that needs to be solved. Summary of the Invention
[0004] In response to the technical problems existing in the prior art, the present invention provides a medical information generation method and system based on model processing, and integrates the method and system into a smart watch.
[0005] The technical solution of the present invention to solve the above technical problems is as follows: The present invention provides a method for generating medical information based on model processing, the method comprising a smart watch: The smart watch detects voice signals emitted from the outside; the voice signals are voice signals emitted by different speakers; determining a sound source direction according to the voice signal; Controlling the camera of the smartwatch to face the direction of the sound source to collect video information in the direction of the sound source; Inputting the video information into a text analysis model to obtain a text analysis result; Reanalyzing the text analysis result to obtain a reanalysis result; The reanalysis results are transmitted to a computer.
[0006] Optionally, before the smart watch detects the voice signal emitted from the outside, the smart watch also receives a wake-up signal from the user; in response to the wake-up signal, the smart watch detects the voice signal emitted from the outside.
[0007] Optionally, the wake-up signal includes a touch operation or a voice command.
[0008] Optionally, the step of inputting the video information into a text analysis model to obtain a text analysis result includes: Splitting the video information to create one or more sub-videos; Extracting text data from each of the sub-videos; wherein the text data corresponds to the sub-videos in a one-to-one manner; The text data is input into a text analysis model to obtain a text analysis result.
[0009] Optionally, splitting the video information to create one or more sub-videos includes: extracting voiceprint features of different speakers in the video information; Extracting audio information and facial features corresponding to each of the voiceprint features; Synchronize the audio information of different speakers with their facial features; Based on the synchronized audio information and facial features, sub-video information of each speaker is generated.
[0010] Optionally, after splitting the video information to create one or more sub-videos, the method further includes: The sub-videos whose time length is less than the preset time threshold are deleted to obtain valid sub-videos.
[0011] Optionally, the text analysis model includes: a contrast analysis model and a contextual reasoning model.
[0012] Optionally, inputting the text data into a text analysis model to obtain a text analysis result includes: Inputting the text data into a comparative analysis model to obtain a conflicting text and a first credible text; Inputting the conflicting text into the contextual reasoning model to obtain a second credible text; The first credible text and the second credible text are used as text analysis results.
[0013] Optionally, reanalyzing the text analysis result to obtain a reanalysis result includes: Get the format information of each field of medical record documents; Matching the text analysis result with the field to obtain a matching result; The text analysis result is modified based on the matching result to obtain a re-analysis result; the re-analysis result is text data that conforms to the format information.
[0014] The present invention also provides a medical information generation system based on model processing, which is integrated into a smart watch; the system comprises: A video acquisition module is configured to detect external voice signals, wherein the voice signals are voice signals emitted by different speakers; determine the direction of the sound source based on the voice signals; and control the camera of the smart watch to face the direction of the sound source to acquire video information in the direction of the sound source; A text analysis module, configured to input the video information into a text analysis model to obtain a text analysis result; and reanalyze the text analysis result to obtain a reanalysis result; The data transmission module is used to transmit the reanalysis results to the computer.
[0015] In addition, to achieve the above-mentioned purpose, the present invention also proposes an electronic device, comprising: a memory for storing a computer software program; a processor for reading and executing the computer software program, thereby implementing the medical information generation method based on model processing as described above.
[0016] In addition, to achieve the above-mentioned purpose, the present invention also proposes a non-transitory computer-readable storage medium, in which a computer software program is stored. When the computer software program is executed by a processor, the medical information generation method based on model processing as described above is implemented.
[0017] The beneficial effects of the present invention are: (1) The smartwatch of the present invention detects externally emitted voice signals; the voice signals are voice signals emitted by different speakers; the direction of the sound source is determined based on the voice signals; the camera of the smartwatch is controlled to face the direction of the sound source to collect video information in the direction of the sound source; the video information is input into a text analysis model to obtain a text analysis result; the text analysis result is re-analyzed to obtain a re-analysis result; and the re-analysis result is transmitted to a computer. Based on this, the watch device transcribes the content of documents through real-time voice, greatly improving the real-time nature of document entry, enabling doctors and nurses to complete document recording more quickly and improve work efficiency.
[0018] (2) Furthermore, the present invention splits video information based on the voiceprint information of different speakers to create sub-videos corresponding to different speakers; extracts text data from each sub-video; the text data corresponds to the sub-video one-to-one; inputs the text data into a comparative analysis model to obtain a conflicting text and a first credible text; inputs the conflicting text into the contextual reasoning model to obtain a second credible text; and uses the first credible text and the second credible text as text analysis results. Thus, the patient's condition can be accurately understood, reducing human errors and information loss. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 A scene diagram of the medical information generation method based on model processing provided by the present invention; Figure 2 A flowchart of the medical information generation method based on model processing provided by the present invention; Figure 3 A schematic diagram of the structure of the medical information generation system based on model processing provided by the present invention; Figure 4 A schematic diagram of the hardware structure of a possible electronic device provided by the present invention; Figure 5 A schematic diagram of the hardware structure of a possible computer-readable storage medium provided by the present invention. DETAILED DESCRIPTION
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0021] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the specified features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.
[0022] In the description of the present invention, the term "for example" is used to mean "used as an example, illustration or illustration". Any embodiment of the present invention described as "for example" is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any person skilled in the art to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed herein.
[0023] See also Figure 1 , Figure 1 This is a scene diagram of the medical information generation method based on model processing provided by the present invention. Figure 1As shown, the terminal and server are connected via a network, such as a wired or wireless network. Terminals may include, but are not limited to, portable devices such as smartwatches, mobile phones, and tablets with network platform applications installed, as well as fixed devices such as computers, kiosks, and advertising machines. The server provides various services to users, including service push servers and user recommendation servers.
[0024] It should be noted that Figure 1 The scenario diagram of the medical information generation method based on model processing shown is only an example. The terminal, server and application scenario described in the embodiment of the present invention are for more clearly illustrating the technical solution of the embodiment of the present invention, and do not generate limitations on the technical solution provided by the embodiment of the present invention. Ordinary technicians in this field know that with the evolution of the system and the emergence of new business scenarios, the technical solution provided by the embodiment of the present invention is also applicable to similar technical problems.
[0025] The terminal may be a smartwatch, on which an application is installed, which can be used to: The smart watch detects voice signals emitted from the outside; the voice signals are voice signals emitted by different speakers; determining a sound source direction according to the voice signal; Controlling the camera of the smartwatch to face the direction of the sound source to collect video information in the direction of the sound source; Inputting the video information into a text analysis model to obtain a text analysis result; Reanalyzing the text analysis result to obtain a reanalysis result; The reanalysis results are transmitted to a computer.
[0026] See also Figure 2 , provides a flowchart of the medical information generation method based on model processing of the present invention, comprising the following steps: 201. Detect an external voice signal.
[0027] Specifically, a plurality of microphones may be provided on the smart watch to form a microphone array, and the microphone array is used to detect voice signals emitted from the outside, where the voice signals are voice signals emitted by different speakers.
[0028] 202. Determine a sound source direction according to the voice signal.
[0029] In one embodiment, because sound propagates from different directions to each microphone, the time it takes for the sound to reach each microphone will vary slightly depending on the distance. By calculating these time differences, the direction of the sound's source can be determined. For example, if a microphone array has three microphones, the sound first reaches the closest microphone, followed by the other two. By analyzing the time differences between each microphone receiving the sound, the angle of the sound source can be calculated.
[0030] In yet another embodiment, by performing weighted processing on the signals from multiple microphones, the sound signal from a specific direction is enhanced and the noise from other directions is suppressed, so that the direction of the sound source can be determined more accurately.
[0031] 203. Control the camera of the smart watch to face the direction of the sound source to collect video information in the direction of the sound source.
[0032] Among them, the smart watch can map the direction of the sound source with the angle of the camera through the embedded algorithm, and the system in the watch needs to be able to control the rotation or adjustment of the camera in real time.
[0033] In one embodiment, the watch's tilt angle and direction are detected by digital gyroscope and accelerometer sensors. Once the direction of the sound source is calculated, the feedback from these sensors can be used to drive the servo drive system on the smartwatch to physically rotate the camera, which can include pitch angle and horizontal rotation angle. Once the camera is facing the direction of the sound source and is stable, the watch will begin recording video in real time. To achieve efficient video capture, the camera's sensor needs to have a high resolution and autofocus function to ensure that video content near the sound source can be clearly captured even in fast-moving or low-light environments.
[0034] 204. Input the video information into a text analysis model to obtain a text analysis result.
[0035] In one embodiment, step 204 may include the following steps: a. Split the video information and create one or more sub-videos.
[0036] Specifically, voiceprint features of different speakers in the video information are extracted; audio information and facial features corresponding to each voiceprint feature are extracted; the audio information of different speakers and their facial features are synchronized; and sub-video information of each speaker is generated based on the synchronized audio information and facial features.
[0037] b. Delete the sub-videos whose time length is less than the time threshold in the sub-videos to obtain valid sub-videos.
[0038] c. Extracting text data from each of the sub-videos; the text data corresponds one-to-one to the sub-videos.
[0039] Speech recognition is performed on each sub-video, and text data is extracted. Each speech signal is converted into text, and each text segment is associated with a corresponding facial feature frame. The goal of this process is to accurately match speech and facial information, ensuring that each speaker's speech content is consistent with their facial expressions.
[0040] d. Input the text data into a text analysis model to obtain text analysis results.
[0041] In one embodiment, the text analysis model includes a contrastive analysis model and a contextual reasoning model. The contrastive analysis model is a pre-trained model (Roberta) that analyzes the expressions of different speakers and identifies texts that may contain logical conflicts. The contextual reasoning model is a deep bidirectional language model (ELMo) that infers the context of texts with logical conflicts, understands the causal relationship between the expressions of different speakers, and determines the reasonable text content. Thus, the text data is input into the contrastive analysis model to obtain conflicting texts and a first credible text. The conflicting text is then input into the contextual reasoning model to obtain a second credible text. The first credible text and the second credible text are used as the text analysis results.
[0042] It should be noted that the specific implementation process of the Roberta model and the ELMo model belongs to the existing technology and will not be described here.
[0043] 205. Reanalyze the text analysis result to obtain a reanalysis result.
[0044] In one embodiment, the format information of each field of the medical record document is obtained; the text analysis result is matched with the field to obtain a matching result; the text analysis result is corrected based on the matching result to obtain a re-analysis result; the re-analysis result is text data that conforms to the format information.
[0045] 206. Transmit the reanalysis result to a computer.
[0046] The reanalysis results can be transmitted to a computer via Wi-Fi, Bluetooth or a dedicated medical protocol (such as HL7). This method simplifies the data transmission process, reduces manual operations, and improves processing efficiency, especially in emergency and real-time monitoring.
[0047] Before step 201, the method further includes: receiving a user's wake-up signal via a smart watch; and in response to the wake-up signal, detecting an external voice signal via the smart watch. The wake-up signal includes a doctor's touch operation or voice command.
[0048] As a result, smart watches can transcribe document contents through real-time voice, greatly improving the real-time nature of document entry, allowing doctors and nurses to complete document records more quickly and improve work efficiency.
[0049] See also Figure 3 , Figure 3 This is a structural diagram of the medical information generation system based on model processing provided by the present invention.
[0050] like Figure 3 As shown, a medical information generation system based on model processing proposed in an embodiment of the present invention is integrated into a smart watch, and the system includes: The video acquisition module 301 is configured to detect external voice signals, wherein the voice signals are voice signals emitted by different speakers; determine the direction of the sound source based on the voice signals; control the camera of the smart watch to face the direction of the sound source and acquire video information in the direction of the sound source; The text analysis module 302 is configured to input the video information into a text analysis model to obtain a text analysis result; and reanalyze the text analysis result to obtain a reanalysis result. The data transmission module 303 is used to transmit the reanalysis results to a computer.
[0051] See also Figure 4 , Figure 4 Schematic diagram of an embodiment of an electronic device provided by an embodiment of the present invention. Figure 4 As shown, an embodiment of the present invention provides an electronic device 400, including a memory 410, a processor 420, and a computer program 411 stored in the memory 410 and executable on the processor 420. When the processor 420 executes the computer program 411, the following steps are implemented: An application is installed on the electronic device 400, and the application allows the smart watch to detect voice signals emitted from the outside; the voice signals are voice signals emitted by different speakers; the direction of the sound source is determined based on the voice signals; the camera of the smart watch is controlled to face the direction of the sound source to collect video information in the direction of the sound source; the video information is input into a text analysis model to obtain a text analysis result; the text analysis result is re-analyzed to obtain a re-analysis result; and the re-analysis result is transmitted to a computer.
[0052] See also Figure 5 , Figure 5Schematic diagram of an embodiment of a computer-readable storage medium provided in an embodiment of the present invention. Figure 5 As shown, this embodiment provides a computer-readable storage medium 500 on which a computer program 411 is stored. When the computer program 411 is executed by a processor, the following steps are implemented: The computer-readable storage medium 500 includes an application, wherein the smart watch detects voice signals emitted from the outside; the voice signals are voice signals emitted by different speakers; the direction of the sound source is determined based on the voice signals; the camera of the smart watch is controlled to face the direction of the sound source to collect video information in the direction of the sound source; the video information is input into a text analysis model to obtain a text analysis result; the text analysis result is re-analyzed to obtain a re-analysis result; and the re-analysis result is transmitted to a computer.
[0053] It should be noted that, in the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0054] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0055] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A system that specifies the functions of a box or boxes.
[0056] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction system that is implemented in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0057] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0058] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0059] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. A method for generating medical information based on model processing, characterized in that: Including a smart watch; The smartwatch detects voice signals emitted from the outside; the voice signals are voice signals emitted by different speakers; determining a sound source direction according to the voice signal; Controlling the camera of the smartwatch to face the direction of the sound source to collect video information in the direction of the sound source; Inputting the video information into a text analysis model to obtain a text analysis result; Reanalyzing the text analysis result to obtain a reanalysis result; The reanalysis results are transmitted to a computer.
2. The medical information generation method based on model processing according to claim 1, characterized in that: Before the smart watch detects the voice signal emitted from the outside, the smart watch also receives a wake-up signal from the user; in response to the wake-up signal, the smart watch detects the voice signal emitted from the outside.
3. The medical information generation method based on model processing according to claim 2, characterized in that: The wake-up signal includes a touch operation or a voice command.
4. The method for generating medical information based on model processing according to claim 1, characterized in that: The step of inputting the video information into a text analysis model to obtain a text analysis result includes: Splitting the video information to create one or more sub-videos; Extracting text data from each of the sub-videos; wherein the text data corresponds to the sub-videos in a one-to-one manner; The text data is input into a text analysis model to obtain a text analysis result.
5. The method for generating medical information based on model processing according to claim 4, characterized in that: The step of splitting the video information to create one or more sub-videos includes: extracting voiceprint features of different speakers in the video information; Extracting audio information and facial features corresponding to each of the voiceprint features; Synchronize the audio information of different speakers with their facial features; Based on the synchronized audio information and facial features, sub-video information of each speaker is generated.
6. The method for generating medical information based on model processing according to claim 5, characterized in that: After splitting the video information and creating one or more sub-videos, the method further includes: The sub-videos whose time length is less than a preset time threshold are deleted to obtain valid sub-videos.
7. The method for generating medical information based on model processing according to claim 4, characterized in that: The text analysis model includes: a contrast analysis model and a contextual reasoning model.
8. The method for generating medical information based on model processing according to claim 7, characterized in that: Inputting the text data into a text analysis model to obtain a text analysis result includes: Inputting the text data into a comparative analysis model to obtain a conflicting text and a first credible text; Inputting the conflicting text into the contextual reasoning model to obtain a second credible text; The first credible text and the second credible text are used as text analysis results.
9. The method for generating medical information based on model processing according to claim 1, characterized in that: The reanalyzing the text analysis result to obtain a reanalysis result includes: Get the format information of each field of medical record documents; Matching the text analysis result with the field to obtain a matching result; The text analysis result is modified based on the matching result to obtain a re-analysis result; the re-analysis result is text data that conforms to the format information.
10. A medical information generation system based on model processing, characterized in that: Integrated into a smart watch; the system includes: A video acquisition module is configured to detect external voice signals, wherein the voice signals are voice signals emitted by different speakers; determine the direction of the sound source based on the voice signals; and control the camera of the smart watch to face the direction of the sound source to acquire video information in the direction of the sound source; A text analysis module, configured to input the video information into a text analysis model to obtain a text analysis result; and reanalyze the text analysis result to obtain a reanalysis result; The data transmission module is used to transmit the reanalysis results to the computer.
Citation Information
Patent Citations
Video interaction control method and device
CN106888361A
Method and device for generating form based on speech recognition, equipment and medium
CN113380234A
Voice separation enhancement method and system in multi-sound-source and noise environment
CN117238311A
Method for generating medical record report based on doctor-patient dialogue
CN119964717A
Electronic device and method for providing personalized audio information
WO2022216059A1