A medical information entry method, device, electronic device and readable storage medium

By taking medical document photos in the ward and correcting them with audio recording and video data, the problem of time-consuming and error-prone input of traditional medical information is solved, and efficient and accurate information entry and storage is achieved.

CN119625762BActive Publication Date: 2025-07-04PEKING UNION MEDICAL COLLEGE HOSPITAL +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510157747.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-07-04
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

Traditional medical information entry methods are time-consuming and error-prone, especially when doctors hold medical documents in hand to write in wards, OCR recognition is not effective.

Method used

Use portable cameras or fixed cameras in the ward to take photos of medical documents, combine OCR technology, recording data and video data for correction, use voice recognition and body movement analysis for information correction, and finally enter the database.

Benefits of technology

It improves the accuracy and efficiency of information entry, reduces the labor intensity of manual entry, reduces the error rate, and realizes electronic data storage and query.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625762B_ABST
    Figure CN119625762B_ABST
Patent Text Reader

Abstract

The present application provides a medical information entry method, apparatus, electronic device, and readable storage medium. Among them, the method includes: obtaining a medical document photo obtained by photographing the text in a medical document through a camera located in a ward; calling the recorded audio data and video data when the medical document was generated according to the shooting time information and shooting position information of the medical document photo; using OCR character recognition technology to extract first document information from the medical document photo; correcting the first document information based on the recorded audio data and video data; and entering the corrected first document information into a database. By this method, the accuracy of information entry is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical technology, and in particular, to a medical information entry method, device, electronic device, and readable storage medium. Background Art

[0002] In the process of clinical medical treatment, data entry is an important step that is both time-consuming and error-prone. The more traditional method is to use manual entry to fill in case data in a CRF (Case Report Form) form. This method not only has a large workload, but also is prone to errors during data entry, affecting data accuracy.

[0003] In order to improve the efficiency and accuracy of data entry, some technical teams have developed and started to use OCR (Optical Character Recognition) technology for data entry. They obtain the image of the document by scanning or photographing the medical document, and then use OCR technology to perform character recognition on the image to obtain the recognition result. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a medical information entry method, device, electronic device, and readable storage medium to improve the accuracy of information entry.

[0005] In a first aspect, an embodiment of this application provides a medical information entry method, including:

[0006] Obtain a medical document photo obtained by photographing the text in a medical document through a camera located in the ward;

[0007] According to the shooting time information and shooting position information of the medical document photo, call the recording data and video data when the medical document was generated;

[0008] Use OCR character recognition technology to extract first document information from the medical document photo;

[0009] Correct the first document information based on the recording data and the video data;

[0010] Enter the corrected first document information into the database.

[0011] Combined with the first aspect, an embodiment of this application provides a first possible implementation manner of the first aspect, where the camera is a portable camera set on a doctor; the obtaining of the medical document photo obtained by photographing the text in the medical document through a camera located in the ward includes:

[0012] Take pictures at regular intervals to obtain candidate photos;

[0013] Foreground extraction is performed on the candidate photo to extract the current foreground image in the candidate photo; the current foreground image includes a handheld medical document;

[0014] Calculate the foreground image similarity between the current foreground image and the previously obtained historical foreground image;

[0015] If the foreground image similarity is less than the preset similarity, the current foreground image is used as the medical document photo.

[0016] Combined with the first aspect, the embodiment of the present application provides a second possible implementation manner of the first aspect, wherein, the calling the recording data and video data when the medical document is generated according to the shooting time information and shooting location information of the medical document photo includes:

[0017] Determine the candidate room where the medical document photo is taken according to the shooting location information of the medical document photo;

[0018] Adopt a visual positioning method to calculate the score of the reference picture of each candidate room according to the medical document photo;

[0019] Determine the candidate room with a score exceeding the preset value as the target room;

[0020] According to the shooting time information of the medical document photo, select the target recording segment from the video data of the target room as the recording data when the medical document is generated, and select the target video segment as the video data when the medical document is generated.

[0021] Combined with the first aspect, the embodiment of the present application provides a third possible implementation manner of the first aspect, wherein, the correcting the first document information based on the recording data and the video data includes:

[0022] Perform speech recognition on the recording data to obtain speech recognition text;

[0023] Perform analysis on the human body limb movements in the video data to obtain limb movement feature data;

[0024] Extract the reference second document information corresponding to the speech recognition text from the pre-established historical speech recognition text-reference second document information comparison table based on the speech recognition text;

[0025] Extract the reference third document information corresponding to the limb movement feature data from the pre-established historical limb movement feature data-reference third document information comparison table based on the limb movement feature data;

[0026] Based on the reference second document information and the reference third document information, correct the first document information.

[0027] Combined with the second possible implementation manner of the first aspect, the embodiment of the present application provides a fourth possible implementation manner of the first aspect, wherein, the step of inputting the corrected first document information into the database includes:

[0028] If there are multiple pieces of corrected first document information, retrieve the medical information of the patient corresponding to the medical document photo;

[0029] Display the medical information and the multiple pieces of corrected first document information on the doctor terminal, so as to select one of the multiple pieces of corrected first document information as the final first document information according to the selection instruction received by the doctor terminal;

[0030] Store the final first document information in the database, and update the historical speech recognition text-reference second document information comparison table and the historical limb movement feature data-reference third document information comparison table based on the corrected final first document information.

[0031] In a second aspect, the embodiment of the present application further provides a medical information input device, including:

[0032] An acquisition module, configured to acquire a medical document photo obtained by photographing the text in a medical document through a camera located in a ward;

[0033] A call module, configured to call the recorded audio data and video data when the medical document is generated according to the shooting time information and shooting position information of the medical document photo;

[0034] An extraction module, configured to extract first document information from the medical document photo using OCR character recognition technology;

[0035] A correction module, configured to correct the first document information based on the recorded audio data and the video data;

[0036] An input module, configured to input the corrected first document information into the database.

[0037] In a third aspect, the embodiment of the present application further provides an electronic device, including: a processor, a memory, and a bus, where the memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus, and when the machine-readable instructions are executed by the processor, the steps in any one of the possible implementation manners in the first aspect described above are executed.

[0038] Fourthly, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps in any possible implementation manner of the first aspect described above are executed.

[0039] The medical information entry method provided by the embodiment of the present application corrects the first document information in text form extracted from the medical document photo through the audio data and video data generated when taking the medical document photo, so as to further improve the accuracy of the text information originally directly entered into the database and ensure the accuracy of the information.

[0040] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] To more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without creative efforts.

[0042] Figure 1 The flowchart of a medical information entry method provided by the embodiment of the present application is shown;

[0043] Figure 2 The flowchart of another medical information entry method provided by the embodiment of the present application is shown;

[0044] Figure 3 The structural schematic diagram of a medical information entry device provided by the embodiment of the present application is shown;

[0045] Figure 4 The structural schematic diagram of an electronic device provided by the embodiment of the present application is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part rather than all of the embodiments of this application. Components of the embodiments of this application described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative efforts fall within the scope of protection of this application.

[0047] In the related art, when doctors inquire about inpatients, they usually casually record the information from the inquiries, such as the physical condition of the patients, the recovery situation, the way of adjusting medications, and so on. After obtaining this information, the OCR technology can be used to input the information written by the doctors into the system for subsequent viewing. However, when using OCR for text input, misrecognition often occurs. There are two main reasons. One is that when doctors inquire about patients in the ward, they do not record at a desk. Usually, they use the pad in their hands as a support and then write on the pad, so the standardization of writing is much worse. The other is that doctors write relatively fast and have many connected strokes, so the written text is relatively arbitrary, and it is very difficult to accurately complete the recognition using general OCR recognition technology.

[0048] In view of the above situation, this application provides a medical information input method, as Figure 1 shown, including the following steps:

[0049] S101: Obtain a medical document photo obtained by photographing the text in a medical document through a camera located in the ward;

[0050] S102: According to the shooting time information and shooting location information of the medical document photo, call the recorded audio data and video data when the medical document was generated;

[0051] S103: Use OCR text recognition technology to extract the first document information from the medical document photo;

[0052] S104: Correct the first document information based on the recorded audio data and video data;

[0053] S105: Input the corrected first document information into the database.

[0054] The application scenario of the solution provided by this application is the scenario where doctors make rounds in various wards. In this scenario, doctors need to write text on medical documents at any time and place, so the standardization of text writing is relatively low.

[0055] In step S101, the camera can be carried by the doctor or installed at a fixed position in the ward. If a portable camera carried by the doctor is used for shooting, the problem of shooting image quality and the remaining battery power need to be considered. Therefore, when choosing a portable camera, it is necessary to give priority to using a camera that can capture a picture sufficient for subsequent image analysis, and the remaining battery power alarm function should be considered. Correspondingly, for the fixed camera installed at a fixed position in the ward, since it is equipped with a power transmission line, there is no need to consider the problem of power endurance, and only the clarity problem needs to be considered.

[0056] Generally speaking, the portable camera can be carried in the doctor's upper pocket, because when doctors record medical documents, they always place the medical documents in front of them for recording. In this way, the portable camera can easily complete the acquisition of images. Correspondingly, if the camera is a fixed camera installed at a fixed position in the ward, the shooting angle should be considered (doctors or other objects may block the camera from taking pictures of medical documents). Therefore, generally, if the fixed camera method is adopted, it is necessary to consider setting up multiple cameras in the room to achieve this (usually, one camera can be set at each of the four corners of the room ceiling). At the same time, the resolution of the camera should also be considered, mainly because the cameras at the corners of the roof are far from the medical documents. If the shooting clarity is not high enough, it will affect the subsequent OCR recognition.

[0057] During the shooting process, the system can use computer vision technology to perform real-time image quality analysis, combined with a deep learning-based image quality evaluation model, and feedback the image quality information to the user in real time to guide the user to take high-quality images. The captured pictures can be preprocessed first, including operations such as denoising, binarization, dilation, and erosion, to improve the text readability of the images. After the OCR recognition is completed, post-processing can also be performed on the recognition results, mainly including using a deep learning-based language model for semantic correction to improve the accuracy of the recognition results.

[0058] Generally speaking, the shooting of medical document pictures can be carried out at a fixed time or triggered. Specifically, the doctor can manually trigger the controller of the camera to control the camera to take pictures (the controller can be set on the camera or taken pictures through a remote controller), or take pictures at intervals of a certain period of time (usually a few seconds).

[0059] Specifically, as Figure 2 shown, step S101 can be implemented in the following manner:

[0060] S1011: Take a photo at regular intervals to obtain candidate photos;

[0061] S1012: Extract the foreground from the candidate photos to extract the current foreground image in the candidate photos; the current foreground image includes a handheld medical document;

[0062] S1013: Calculate the foreground image similarity between the current foreground image and the previously obtained historical foreground image;

[0063] S1014: If the foreground image similarity is less than the preset similarity, then use the current foreground image as the medical document photo.

[0064] In step S1011, the camera takes photos at regular intervals to obtain candidate photos. In this step, the camera is a portable camera set on the doctor. The candidate photos should include photos of the medical document held by the doctor.

[0065] Furthermore, in step S1012, foreground extraction needs to be performed on the candidate photos. Since it is certain that the object of foreground extraction is the medical document, therefore, in order to improve the efficiency and accuracy of foreground extraction, the following two methods can be selected:

[0066] Method 1: Add an identification mark to the medical document. The identification mark can be a graphic, image or other mark that is not common in the hospital scene. For example, the mark can be a black star, which is located in the upper right corner of the medical document. Then, it can be determined whether the mark exists in the candidate photos through image recognition. If it exists, a predetermined area below the left of the mark is used as the foreground image. Specifically, when implemented, the mark can be designed as a mark with a direction indication. In this way, it can be ensured that after the mark faces a specific direction (such as upward), the area below the left of it or a certain relative position of it is the foreground image.

[0067] That is, step S1012 can be implemented in the following manner:

[0068] Detect whether there is an identification mark in the current candidate photo;

[0069] If it exists, use the position information and shape information of the identification mark to extract a sampling image from the current candidate photo.

[0070] Specifically, the position information and shape information can reflect the size and pointing direction of the recognition identifier. According to the size of the recognition identifier, the size of the foreground image can be reflected (or rather, the size ratio of the foreground image to the recognition identifier is fixed, so the size of the recognition identifier can reflect the size of the foreground image), and thus it can assist in the positioning of the foreground image. The pointing direction reflects the relative position between the recognition identifier and the foreground image.

[0071] In Method 2, a deep learning algorithm can be used for foreground extraction. This deep learning algorithm is mainly composed of structures such as a segmentation model of a convolutional neural network (CNN), an automatic feature extraction mechanism, and a network model combining an encoder-decoder structure. The reason why this method can be used for foreground extraction is mainly that there is a relatively standardized format structure in medical documents. Specifically, the content in medical documents is not simply a blank piece of paper, but is composed of multiple standard-format block diagrams and texts designed according to the rules of doctor inquiries. Specifically, prompt texts and block diagrams / tables are respectively set at different positions in medical documents, and these prompt texts and block diagrams / tables are used to standardize the doctor's inquiry records. Therefore, under a certain degree of standardization, the above deep learning method can be used for foreground extraction (when performing foreground extraction, through these standardized contents, it can play a positioning role). At the same time, since it is very difficult to encounter other contents similar to this standardized content in a hospital scenario, it will not cause misrecognition of the deep learning algorithm.

[0072] In specific implementation, the above two methods can be reused, that is, first use Method 1 for positioning and sampling frame extraction, and then use Method 2 for specific extraction, which avoids the problem that there are too many sampling frames in Method 2 and excessive computing power consumption is required.

[0073] In specific implementation, step S1012 can be implemented in the following manner:

[0074] Step 10121, detect whether there is a recognition identifier in the current candidate photo;

[0075] Step 10122, if there is, use the position information and shape information of the recognition identifier to determine the sampling image of the current candidate photo;

[0076] Step 10123, input the sampling image into the pre-trained foreground extraction model to obtain the current foreground image.

[0077] Since the sampled image is located based on the recognition identifier, the main function of the sampled image is to narrow the search range when the foreground extraction model is processing. Since there are many images generated by shooting, it is preferably to perform a preliminary screening of the foreground image in the manner of step 10121 - step 10123. In some cases, the sampling frame of the foreground extraction model in step 10123 can be similar to the size of the sampled image extracted in step 10122, which can further reduce the sampling range. Specifically, the size and step length of the sampling frame can be determined according to the current remaining computing power. The remaining computing power and the size of the sampling frame are negatively correlated, and the remaining computing power and the step length are negatively correlated.

[0078] In step S1013, it is mainly to calculate the similarity between the current foreground image and the historical foreground image. Generally speaking, there are two types of historical foreground images. One is a completely blank (not handwritten by the doctor) medical document, and there are only standardized prompt texts and block diagrams / tables in this medical document. The other is a medical document taken some time ago, generally the foreground image extracted from the photo of the medical document taken 5 - 10 seconds ago.

[0079] When the historical foreground image is a completely blank medical document, by calculating the similarity, it can be known whether the doctor has handwritten on the medical document (if there are handwritten words in the medical document, the similarity will decrease, and the more words there are, the lower the similarity). However, this method has certain defects. Mainly, as the doctor's handwritten content increases, the similarity will become lower and lower. That is, during the doctor's handwriting process, the situation of continuously triggering step S1014 may occur, resulting in a decrease in accuracy. To avoid this problem, the method of calculating the similarity using multiple consecutive photos can be adopted, that is, a candidate photo is obtained at intervals of a period of time (N seconds, such as 3 seconds), and then the similarity is calculated using the candidate photo. Since the doctor has been handwritten on the medical document, the similarity of the candidate photos obtained later will continue to decrease. When the similarity of the candidate photos obtained within a period of time (M seconds, such as 10 seconds, or 3 - 5 times the time of N seconds) no longer decreases, or no foreground image can be extracted, then in step S1014, the current foreground image with the minimum similarity is used as the medical document photo.

[0080] When the historical foreground image is a medical document taken some time ago, when calculating the similarity, it mainly depends on the similarity of the foreground images of the candidate photos obtained from two adjacent shootings, which reflects whether the doctor has performed a handwriting behavior during this period of time. If the doctor has written a lot of content during this period of time, step S1014 will be triggered. To avoid misrecognition, when no foreground image can be recognized in the current candidate photo, the foreground image in the previous obtained candidate photo should be used as the medical document photo.

[0081] Although there may be people and other things in the background part of the medical document photo, and there may be standardized tables and texts in the foreground image, when extracting the foreground image, it is still possible to extract these parts that are not doctor's abbreviated texts for subsequent OCR recognition.

[0082] Specifically, step S102 can be implemented in the following way:

[0083] S1021: Determine the candidate room where the medical document photo was taken according to the shooting position information of the medical document photo;

[0084] S1022: In the way of visual positioning, calculate the scores of the reference pictures of each candidate room according to the room identifiers recorded in the background image of the medical document photo;

[0085] S1023: Determine the candidate room with a score exceeding the preset value as the target room;

[0086] S1024: According to the shooting time information of the medical document photo, select the target audio segment from the video data of the target room as the audio data when the medical document was generated, and select the target video segment as the video data when the medical document was generated.

[0087] In step S1021, the GPS positioning or Beidou positioning method can be used to determine the candidate room where the medical document photo was taken. Using these two positioning methods, the floor where the doctor is located cannot be determined. Therefore, rooms with the same longitude and latitude coordinates (such as the leftmost room on each floor) may be identified as candidate rooms. Of course, in addition to the above two positioning methods, other positioning technologies can also be used. Specifically, such as wifi positioning, etc. However, it should be noted that there are also deviations in wifi positioning. Even if wifi positioning is combined with GPS / Beidou positioning, there may still be errors. Although there are other positioning technologies in the current technical environment, such as strapdown inertial navigation, infrared and other technologies, the practical value of these technologies in the medical scenario is not high. They can only be realized theoretically, with high actual costs and difficult to be actually used. Therefore, in this solution, usually only GPS or Beidou positioning technology is used. Furthermore, all rooms corresponding to the coordinates can be regarded as candidate rooms.

[0088] In step S1022, visual positioning can be used. Since in this solution, it is certain to obtain a medical document photo. When obtaining the medical document photo, it is generally impossible that the entire image is a medical document, and there will also be some background images. Therefore, auxiliary positioning can be carried out by setting different room identifiers on different floors. Then, when the room identifier is captured, the floor where the medical document photo is located can be determined (the room identifier will be captured in the background image). Of course, more complex methods can also be used for processing, that is, using the visual positioning technology used on the automatic robot. Through direct image coordinate conversion and indirect image coordinate conversion of the captured photo, and mapping the converted coordinates to a pre-determined virtual three-dimensional space (this three-dimensional space is established according to the characteristics of the hospital building), accurate positioning can be carried out. However, this positioning method is not practical, mainly because the computing power consumption is too large. For portable devices, excessive computing power occupancy will cause the captured photos to not be analyzed in time. Of course, it is not denied that this method can achieve the technical purpose.

[0089] After identifying the room identifier through visual positioning, the reference score of the reference picture of each candidate room can be determined (the reference score mainly reflects the probability of the user being in each candidate room when taking the medical document photo). Here, the reference picture of the candidate room (the picture containing the room identifier) is obtained in advance. It can be that each room saves a reference picture (the room identifiers of each room are different), or it can be that the rooms on each floor share a reference picture (the room identifiers of the rooms on the same floor are the same, and the room identifiers of the rooms on different floors are different). After determining the score, in step S1023, the candidate room with the highest score can be used as the target room.

[0090] By pre-setting cameras in each room (generally set at the corners of the room ceiling), in step S1024, the audio data segment whose shooting time conforms to the shooting time information of the medical document photo in the content obtained by the camera can be used as the recording data when the medical document is generated, and the video segment whose shooting time conforms to the shooting time information of the medical document photo can be used as the video data when the medical document is generated. The recording data and video data when the medical document is generated respectively reflect the language and body movements of the doctor when recording the medical document (possibly not the language and body movements of the doctor who writes in the medical document).

[0091] Specifically, step S104 can be implemented in the following way:

[0092] S1041: Perform speech recognition on the recording data to obtain the speech recognition text;

[0093] S1042: Analyze the human body limb movement in the video data to obtain limb movement feature data;

[0094] S1043: Extract the corresponding reference second document information of the speech recognition text from the pre-established historical speech recognition text-reference second document information comparison table based on the speech recognition text;

[0095] S1044: Extract the corresponding reference third document information of the limb movement feature data from the pre-established historical limb movement feature data-reference third document information comparison table based on the limb movement feature data;

[0096] S1045: Correct the first document information based on the reference second document information and the reference third document information.

[0097] In step S1041, it is mainly necessary to perform the recognition of the conventional speech-to-text (the recognition of Mandarin is sufficient). In some cases, if the region where the method is implemented can be determined, the dialect speech recognition technology can be considered for addition. Specifically, this text conversion process mainly depends on the acoustic model and the language model. Among them, the acoustic model is responsible for converting the speech signal into a phoneme or word-level representation, while the language model converts these representations into the final text output. In the training stage, the system needs to establish the acoustic model and the language model by analyzing a large amount of speech data; in the recognition stage, the system automatically recognizes the real-time speech and generates the text result.

[0098] In step S1042, the recognition of the limb movement is mainly performed, and specifically, the human body key point detection, deep learning model, etc. can be used to achieve it. Considering that the medical scenario is indoors and the light is sufficient, and there are few interferences, general technologies can be used. Since the result of the limb movement recognition (limb movement feature data) is directly used to retrieve in the table in the subsequent steps, it is not necessarily necessary to assign a clear meaning to the result of each limb movement recognition, as long as the association can be made. For the sake of clarity, only two examples are listed here for illustration. The meaning of the result of the limb movement recognition can be nodding with hands behind the back, which may indicate that the patient is in a stage that meets the treatment expectation; another example is that the meaning of the result of the limb movement recognition can be a specific limb stretching movement (such as horizontally stretching the arm when the body leans forward, raising the thigh, etc.), and at this time, it means that the doctor is guiding the patient to perform a certain verification rehabilitation movement.

[0099] The content of steps S1043 and S1044 is similar, both are queries in the established table. For example, the content in the historical speech recognition text-reference second document information comparison table and the historical limb movement feature data-reference third document information comparison table is based on the information that has been verified by doctors as correct in history, so it has good reference value.

[0100] It can be considered that in the historical speech recognition text - reference second document information comparison table, the corresponding historical speech recognition text and reference second document information are respectively the text recognition results obtained through speech recognition in step S1041 and the text after the confirmation / modification of the text recognition results. Correspondingly, in the historical body movement feature data - reference third document information comparison table, the corresponding historical body movement feature data and reference third document information are respectively the body movement feature data recognized in the past through step S1042 and the text content obtained after the doctor recognizes the body movement feature data or watches the video data.

[0101] It should be noted that the reference second document information and reference third document information usually do not directly replace the first document information, but the reference second document information and reference third document information can reflect which first document information is the closest to the real information. Specifically, since the first document information extracted from the medical document photo using the OCR text recognition technology is not unique (the model outputs each text arrangement and the corresponding probability), in step S103, the model can be selected to output 3 - 5 results with the highest probability (candidate first document information), and then, using the relevance between the reference second document information and reference third document information and these 3 - 5 results, the one that is logically (usually language logic) most similar to the reference second document information and reference third document information is selected as the final result from these 3 - 5 results.

[0102] For example, if the meaning of the result of body movement recognition is a specific stretching body movement, then among multiple results, the probability of the result describing the rehabilitation advice can be increased as the final result. In some cases, due to the different document attributes, the content recorded by the doctor in the document may not include rehabilitation movements (such as the content recorded in the document is for the doctor's own view and there is no need to record rehabilitation movements). Therefore, the semantic recognition method should be used to calculate the semantic relevance between different candidate first document information and the second document information and the third document information respectively, and then, based on the semantic relevance, it is determined which candidate first document information is the most accurate. For example, if the content recorded in a certain candidate first document information is related to the medication requirements during the rehabilitation stage, although only body movements appear in the third document information and there is no medication movement, it should also be considered that there is a strong correlation between the candidate first document information and the third document information because the objects described by both are related to rehabilitation, and the probability of the candidate first document information as the final result should be increased.

[0103] Similarly, the processing methods of the second document information and the third document information are the same. It is not required that the text contents be completely relevant, but only that the relevance between the two is relatively strong, such as the described objects being related. Since the types of communication contents that occur in the hospital are not many, this method can only be adopted in the hospital environment.

[0104] Specifically, in addition to selecting a certain candidate first document information as the final result when using the second document information and the third document information, the second document information and the third document information can also be used to modify the final output result. However, since this part of the modification is not a substantial change but only a formal fine-tuning, no in-depth description will be given.

[0105] That is, step S1045 can be implemented in the following manner:

[0106] Step 10451, calculate the first semantic relevance between each first document information (candidate first document information) and the reference second document information;

[0107] Step 10452, calculate the second semantic relevance between each first document information (candidate first document information) and the reference third document information;

[0108] Step 10453, select the target first document information from multiple first document information according to the first semantic relevance and the second semantic relevance;

[0109] Step 10454, correct the target first document information by using the reference second document information and the reference third document information.

[0110] In step S105, the first document information entered into the database is the target first document information.

[0111] Of course, when performing OCR recognition, an open-source OCR engine such as Tesseract, or an OCR model based on deep learning such as CRNN or Attention OCR, can be considered to perform character recognition on the preprocessed image.

[0112] Specifically, it can be to calculate the semantic relevance between each result and the reference second document information, and the semantic relevance between each result and the reference third document information respectively. Then, the weighted average method is adopted to calculate the semantic similarity parameter of each result. After that, the result with the largest semantic similarity parameter is used as the correction result of step S104 (the first document information in step S104). Finally, this correction result is input into the database.

[0113] However, in actual processing, there may be a situation where the difference in the semantic similarity parameters between the candidate first document information with the highest semantic similarity parameter and the candidate first document information with the second and third highest semantic similarity parameters is not significant. In this case, the following method can be considered for processing, that is, step S105 can be implemented as follows:

[0114] S1051: If there are multiple corrected first document information, retrieve the medical information of the patient corresponding to the medical document photo;

[0115] S1052: Display the medical information and the multiple corrected first document information on the doctor's terminal, so as to select one from the multiple corrected target first document information as the final first document information according to the selection instruction received by the doctor's terminal;

[0116] S1053: Store the final first document information in the database, and update the historical speech recognition text - reference second document information comparison table and the historical limb movement feature data - reference third document information comparison table based on the corrected final first document information.

[0117] That is, when the top several candidate first document information with relatively high semantic similarity parameters cannot be distinguished by referring to the second document information and the third document information, the historical medical information of the patient can be used for reference. The medical information can be, for example, the patient's personal information (height, age, weight, etc.), historical medical treatment information (what diseases the patient has had in the past, the treatment situation, etc.), and the current diagnosis and treatment information, etc.

[0118] After obtaining the medical information, these information and the corrected first document information can be displayed on the doctor's terminal for the doctor to manually select / modify a certain first document information as the final first document information. Then, the first document information selected / modified by the doctor can be stored in the database. At the same time, the tables in the database can also be updated to make the next recognition more accurate. Of course, from a practical perspective, the doctor's terminal should have certain permissions to view the tables in the database to correct error information in a timely manner.

[0119] The text finally stored in the database is preferably converted into a structured data format, such as JSON, for subsequent processing and application.

[0120] Overall, the method provided by this application has the following advantages:

[0121] 1. High efficiency: The technology provided by this application, through OCR technology and big data model technology, enables users to complete form filling by uploading pictures (such as inspection forms, test reports, etc.), greatly saving the time for users to fill out forms. Compared with the traditional manual filling method, our method can reduce the time for filling out forms by more than 70%, greatly improving the efficiency of data entry.

[0122] 2. Improved accuracy: Since this invention uses voice recognition results and body recognition results for auxiliary recognition, and also adopts a series of technical means such as image quality improvement, image preprocessing and postprocessing, OCR recognition, and data structuring, it can effectively improve the accuracy of recognition results. According to our experimental data, this invention significantly improves the accuracy of data entry compared with traditional OCR technology.

[0123] 3. User-friendly: The technology provided by this application also provides data storage and display functions, presenting the results of OCR in the form of charts or forms to users, enabling users to intuitively view the results of OCR, greatly improving the user experience. At the same time, this invention can directly store the results of OCR in the database, facilitating users' subsequent data query and use.

[0124] 4. Environmental protection and reduced labor intensity: Traditional data entry methods require a large amount of paper and manpower, not only affecting the environment but also having a high labor intensity. The method of this invention is completely electronic, which is not only environmentally friendly but also greatly reduces the labor intensity.

[0125] Based on the same technical concept, this application also provides a medical information entry device, as Figure 3 shown, the device includes:

[0126] An acquisition module 301, configured to acquire a medical document photo obtained by photographing the text in a medical document through a camera located in the ward;

[0127] A calling module 302, configured to call the recorded audio data and video data when the medical document was generated according to the shooting time information and shooting location information of the medical document photo;

[0128] An extraction module 303, configured to extract first document information from the medical document photo using OCR character recognition technology;

[0129] A correction module 304, configured to correct the first document information based on the recorded audio data and the video data;

[0130] An entry module 305, configured to enter the corrected first document information into the database.

[0131] Optionally, the camera is a portable camera set on the doctor; when the obtaining module 301 is used to obtain a medical document photo obtained by photographing the text in the medical document through a camera located in the ward, it specifically is used for:

[0132] Taking a photo at a predetermined time interval to obtain candidate photos;

[0133] Performing foreground extraction on the candidate photos to extract the current foreground image in the candidate photos; the current foreground image includes a handheld medical document;

[0134] Calculating the foreground image similarity between the current foreground image and a previously obtained historical foreground image;

[0135] If the foreground image similarity is less than a preset similarity, using the current foreground image as the medical document photo.

[0136] Optionally, when the calling module 302 is used to call the recording data and video data generated when the medical document is generated according to the shooting time information and shooting location information of the medical document photo, it specifically is used for:

[0137] Determining a candidate room where the medical document photo is taken according to the shooting location information of the medical document photo;

[0138] Adopting a visual positioning method to calculate the score of the reference picture of each candidate room according to the medical document photo;

[0139] Determining the candidate room with a score exceeding a preset value as the target room;

[0140] According to the shooting time information of the medical document photo, selecting a target recording segment from the video data of the target room as the recording data generated when the medical document is generated, and selecting a target video segment as the video data generated when the medical document is generated.

[0141] Optionally, when the correction module 304 is used to correct the first document information based on the recording data and the video data, it specifically is used for:

[0142] Performing speech recognition on the recording data to obtain speech recognition text;

[0143] Performing human body limb movement analysis on the video data to obtain limb movement feature data;

[0144] Extracting the reference second document information corresponding to the speech recognition text from a pre-established historical speech recognition text-reference second document information comparison table based on the speech recognition text;

[0145] Extract the reference third document information corresponding to the limb movement feature data from the pre-established historical limb movement feature data-reference third document information comparison table based on the limb movement feature data;

[0146] Based on the reference second document information and the reference third document information, correct the first document information.

[0147] Optionally, when the input module 305 is used to input the corrected first document information into the database, it is specifically used for:

[0148] If there are multiple corrected first document information, retrieve the medical information of the patient corresponding to the medical document photo;

[0149] Display the medical information and the multiple corrected first document information on the doctor terminal, so as to select one of the multiple corrected first document information as the final first document information according to the selection instruction received by the doctor terminal;

[0150] Store the final first document information in the database, and update the historical speech recognition text-reference second document information comparison table and the historical limb movement feature data-reference third document information comparison table based on the corrected final first document information.

[0151] Figure 4 A schematic structural diagram of an electronic device provided by an embodiment of the present application, including: a processor 401, a memory 402, and a bus 403. The memory 402 stores machine-readable instructions executable by the processor 401. When the electronic device runs the above information processing method, the processor 401 communicates with the memory 402 through the bus 403, and the processor 401 executes the machine-readable instructions to execute the method steps described in Embodiment 1.

[0152] An embodiment of the present application also provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, it executes the method steps described in Embodiment 1.

[0153] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described device, electronic device, and computer-readable storage medium can refer to the corresponding processes in the foregoing method embodiments, and will not be described in detail here.

[0154] In several embodiments provided in this application, it should be understood that the disclosed methods, devices, electronic devices, and computer-readable storage media can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0155] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0156] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0157] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., which can store program codes.

[0158] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present application, which are used to illustrate the technical solutions of the present application, rather than limiting it. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present application can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims described.

Claims

1. A medical information entry method, characterized in that, Including: Obtaining a medical document photo by taking a picture of the text in the medical document through a camera located in the ward; According to the shooting time information and shooting position information of the medical document photo, calling the audio data and video data when the medical document was generated; Using OCR character recognition technology to extract the first document information from the medical document photo; Correcting the first document information based on the audio data and the video data; Entering the corrected first document information into the database; The step of calling the audio data and video data when the medical document was generated according to the shooting time information and shooting position information of the medical document photo includes: According to the shooting position information of the medical document photo, determining the candidate room where the medical document photo was taken; Adopting a visual positioning method to calculate the score of the reference picture of each candidate room according to the medical document photo; Determining the candidate room with a score exceeding the preset value as the target room; According to the shooting time information of the medical document photo, selecting the target audio segment from the video data of the target room as the audio data when the medical document was generated, and selecting the target video segment as the video data when the medical document was generated; The step of correcting the first document information based on the audio data and the video data includes: Performing speech recognition on the audio data to obtain speech recognition text; Performing analysis on the human body limb movement of the video data to obtain limb movement feature data; Extracting the reference second document information corresponding to the speech recognition text from the pre-established historical speech recognition text-reference second document information comparison table based on the speech recognition text; Extracting the reference third document information corresponding to the limb movement feature data from the pre-established historical limb movement feature data-reference third document information comparison table based on the limb movement feature data; Correcting the first document information based on the reference second document information and the reference third document information; The step of entering the corrected first document information into the database includes: If there are multiple corrected first document information, retrieving the medical information of the patient corresponding to the medical document photo; Displaying the medical information and the multiple corrected first document information on the doctor terminal, so as to select one of the multiple corrected first document information as the final first document information according to the selection instruction received by the doctor terminal; Storing the final first document information in the database, and updating the historical speech recognition text-reference second document information comparison table and the historical limb movement feature data-reference third document information comparison table based on the corrected final first document information.

2. The method according to claim 1, characterized in that, The camera is a portable camera set on the doctor; the step of obtaining a medical document photo by taking a picture of the text in the medical document through a camera located in the ward includes: Taking pictures at regular intervals to obtain candidate photos; Performing foreground extraction on the candidate photos to extract the current foreground image in the candidate photos; the current foreground image includes a handheld medical document. Calculate the foreground image similarity between the current foreground image and the previously obtained historical foreground image; If the foreground image similarity is less than the preset similarity, use the current foreground image as the medical document photo.

3. The method according to claim 2, characterized in that, The step of foreground extraction for the candidate photo to extract the current foreground image in the candidate photo includes: Detect whether there is an identification mark in the current candidate photo; If it exists, use the position information and shape information of the identification mark to extract a sampling image from the current candidate photo; Input the sampling image into the pre-trained foreground extraction model to obtain the current foreground image.

4. The method according to claim 1, wherein The step of correcting the first document information based on the reference second document information and the reference third document information includes: Calculate the first semantic correlation degree between each first document information and the reference second document information; Calculate the second semantic correlation degree between each first document information and the reference third document information; Select the target first document information from multiple first document information according to the first semantic correlation degree and the second semantic correlation degree; Use the reference second document information and the reference third document information to correct the target first document information.

5. A medical information entry device, characterized in that, It includes: An acquisition module for acquiring medical document photos obtained by photographing the text in a medical document through a camera located in the ward; A calling module for calling the recording data and video data at the time of generating the medical document according to the shooting time information and shooting position information of the medical document photo; An extraction module for extracting the first document information from the medical document photo using OCR character recognition technology; A correction module for correcting the first document information based on the recording data and the video data; An input module for inputting the corrected first document information into the database; When the calling module is used to call the recording data and video data at the time of generating the medical document according to the shooting time information and shooting position information of the medical document photo, it is specifically used for: Determine the candidate room where the medical document photo was taken according to the shooting position information of the medical document photo; Adopt a visual positioning method to calculate the score of the reference picture of each candidate room according to the medical document photo; Determine the candidate room with a score exceeding the preset value as the target room; According to the shooting time information of the medical document photo, select the target recording segment from the video data of the target room as the recording data at the time of generating the medical document, and select the target video segment as the video data at the time of generating the medical document; When the correction module is used to correct the first document information based on the recording data and the video data, it is specifically used for: Perform speech recognition on the recording data to obtain speech recognition text; Perform analysis on the body movement characteristics of the person in the video data to obtain body movement characteristic data; Extract the reference second document information corresponding to the speech recognition text from the pre-established historical speech recognition text-reference second document information comparison table based on the speech recognition text; Extract the reference third document information corresponding to the limb movement feature data from the pre-established historical limb movement feature data-reference third document information comparison table based on the limb movement feature data; Correct the first document information based on the reference second document information and the reference third document information; When the input module is used to input the corrected first document information into the database, it is specifically used for: If there are multiple corrected first document information, retrieve the medical information of the patient corresponding to the medical document photo; Display the medical information and the multiple corrected first document information on the doctor terminal, so as to select one from the multiple corrected first document information as the final first document information according to the selection instruction received by the doctor terminal; Store the final first document information in the database, and update the historical speech recognition text-reference second document information comparison table and the historical limb movement feature data-reference third document information comparison table based on the corrected final first document information.

6. An electronic device, characterized in that, Including: A processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 4 are executed.

7. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is run by the processor, the steps of the method according to any one of claims 1 to 4 are executed.

Citation Information

Patent Citations

  • Audio and text combination method and device, electronic equipment and storage medium

    CN115396690A

  • Video subtitle extraction method and device and electronic equipment

    CN119299770A

  • Document data verification method and document data verification support system

    JP2009187352A