Mobile terminal document editor control method and system
By using a document editing control model based on preset encoder and self-attention layer in the mobile electronic medical record system, the problem of low quality of document editing in the existing system is solved, and higher document quality and efficiency are achieved.
Patent Information
- Application Number
- CN202510011573.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-04
- Publication Date
- 2025-05-06
AI Technical Summary
When editing documents in the existing electronic medical record system, there is a problem of poor quality of electronic medical documents edited and stored.
A mobile document editor control method and system is provided. By acquiring electronic medical documents and inputting them into a document editing control model trained based on a language model constructed by a preset encoder and a self-attention layer, text errors are identified and corresponding operation information is generated. Users can perform correction operations based on this information to obtain target electronic medical documents.
Through automated error detection and correction mechanisms, the quality of electronic medical documents is improved, the workload of manual proofreading is reduced, and the accuracy and standardization of medical documents is ensured.
Smart Images

Figure CN119940308A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of medical information technology, and in particular to a mobile document editor control method and system. Background Art
[0002] As an important part of hospital informatization, the electronic medical record system plays a vital role in the medical industry. With the continuous development of medical informatization, the electronic medical record system is not only used to record the patient's condition, but also becomes an important tool to assist medical staff in diagnosis and treatment decision-making. As the core component of the electronic medical record system, the editor is responsible for managing and editing various medical documents, covering everything from anesthesia records to health examination forms, as well as other medical-related records and reports.
[0003] In existing electronic medical record systems, the document editing method is usually as follows: medical staff log in to the system through a computer or mobile device and enter the patient's medical information through the electronic medical record editor; the system will use preset templates to organize and format the medical record content; during the information entry process, medical staff need to select the corresponding diagnosis and treatment options from the drop-down menu, or manually enter specific medical terms; after editing is completed, these electronic medical documents will be stored in a database, usually in a semi-structured format.
[0004] However, when using the existing electronic medical record system to edit documents, there is a problem that the quality of the edited and stored electronic medical documents is poor. Summary of the invention
[0005] The present application provides a mobile document editor control method and system to solve the problem of poor quality of edited and stored electronic medical documents when editing documents in existing electronic medical record systems.
[0006] In a first aspect, an embodiment of the present application provides a method for controlling a mobile document editor, which is applied to a server in a mobile document editor control system, and the method includes:
[0007] Obtain electronic medical documents, which are documents obtained by users through text editing in the mobile document editor control system;
[0008] The electronic medical document is input into the document editing control model to obtain text error operation information corresponding to the electronic medical document. The document editing control model is a model trained based on a language model built on a preset encoder and a self-attention layer to identify text errors in the electronic medical document and generate corresponding operation information. The text error operation information includes error detection information and error correction information.
[0009] Receiving user correction instructions sent by the user according to the text error operation information;
[0010] If the user correction instruction indicates to correct the text of the electronic medical document, a correction operation is performed on the electronic medical document according to the text error operation information to obtain a target electronic medical document.
[0011] In a possible implementation, the document editing control model includes a recognition module and a classification module. The electronic medical document is input into the document editing control model to obtain text error operation information corresponding to the electronic medical document, including:
[0012] Input the electronic medical document into the recognition module to extract the hidden state vector of each word in the electronic medical document;
[0013] The hidden state vector of each word in the electronic medical document is input into the classification module. In the classification module, the error detection information of the electronic medical document is determined, and error correction information corresponding to the error detection information is generated, wherein the text error operation information includes error detection information and error correction information.
[0014] In a possible implementation, the recognition module includes an input layer, an embedding layer, a preset encoder and a self-attention layer, the output of the input layer is used as the input of the connected embedding layer, the output of the embedding layer is used as the input of the connected preset encoder, and the output of the preset encoder is used as the input of the connected self-attention layer; the classification module includes a classifier and an output layer; wherein:
[0015] The input layer is used to receive electronic medical documents;
[0016] The embedding layer is used to convert each word of the electronic medical document in the input layer into a word vector;
[0017] A preset encoder is used to semantically encode the word vector corresponding to each word in the embedding layer based on a multi-layer Transformer encoder to obtain a hidden state vector for each word;
[0018] The self-attention layer is used to perform self-attention processing on the hidden state vector of each word in the preset encoder to obtain the target hidden state vector of each word;
[0019] A classifier is used to perform error detection classification on the target hidden state vector of each word in the self-attention layer to obtain a detection label sequence corresponding to the electronic medical document, and to perform error correction classification on the target hidden state vector of each word to obtain a correction label sequence corresponding to the electronic medical document;
[0020] The output layer is used to generate text error operation information corresponding to the electronic medical document based on the detection label sequence and correction label sequence output by the connected classifier.
[0021] In a possible implementation, the classifier includes an error detection classifier and an error correction classifier; wherein:
[0022] An error detection classifier is used to receive the target hidden state vector of each word in the self-attention layer, and predict the error state of each word using the maximization layer according to the target hidden state vector of each word to obtain a detection label sequence, where each label in the detection label sequence is used to indicate whether the corresponding word has an error;
[0023] The error correction classifier is used to perform operation prediction on the target hidden state vector of each word to obtain a correction label sequence. Each label in the correction label sequence is used to indicate the specific correction operation of the corresponding word.
[0024] In a possible implementation, the classification module further includes a specific layer, and the output of the output layer serves as the input of the connected specific layer; wherein:
[0025] A specific layer is used to update the text error operation information in the output layer to obtain updated text error operation information. The information update processing is the processing of the text error operation information in the output layer based on the behavioral characteristics of the target object. The target object is the object corresponding to the electronic medical document input into the document editing control model.
[0026] In a possible implementation manner, before inputting the electronic medical document into the document editing control model to obtain text error operation information corresponding to the electronic medical document, the method further includes:
[0027] Acquire a historical medical document set, where the historical medical document set includes electronic medical documents stored by multiple objects in a target hospital within a preset historical time period;
[0028] Based on the preset text error annotation requirements, the historical medical document set is annotated to obtain a sample set, each sample in the sample set includes: an electronic medical document and an edit tag sequence corresponding to the electronic medical document, the edit tag sequence is used to indicate that the erroneous text in the electronic medical document is corrected to the correct text;
[0029] Based on the sample set, the initial document editing control model is trained until the output information of the initial document editing control model meets the preset training requirements, thereby obtaining the document editing control model.
[0030] In a possible implementation, based on preset text error annotation requirements, a historical medical document set is annotated to obtain a sample set, including:
[0031] Based on the preset text error annotation requirements, the historical medical document set is annotated to obtain the text error position corresponding to each electronic medical document in the historical medical document set and the correct text after the error text is corrected;
[0032] For each electronic medical document, filtering the words of each electronic medical document according to a preset vocabulary list to be filtered to obtain a filtered electronic medical document;
[0033] According to the text error position and the correct text corresponding to the filtered electronic medical document, the filtered electronic medical document is processed with the minimum edit distance to obtain the corrected target text, the text error type and the error correction operation of the filtered electronic medical document;
[0034] Determine an editing tag sequence corresponding to the electronic medical document according to a text error position, a correct text, a text error type, an error correction operation, and a target text corresponding to the electronic medical document;
[0035] A sample set is constructed based on all electronic medical documents and the edit tag sequence corresponding to each electronic medical document.
[0036] In a possible implementation, after the initial document editing control model is trained based on the sample set until the output information of the initial document editing control model meets the preset training requirements and the document editing control model is obtained, the method further includes:
[0037] Obtaining a test medical document set, where the test medical document set and the historical medical document set have different preset historical time periods;
[0038] Determine a test sample set according to the test medical document set, each test sample in the test sample set includes: an electronic medical document and an edit tag sequence corresponding to the electronic medical document;
[0039] Input the test sample set into the document editing control model to obtain the text error operation test information corresponding to each test sample;
[0040] Determine the model evaluation index according to the edit label sequence and the corresponding text error operation test information in each test sample. The model evaluation index includes the accuracy and recall rate of the model prediction, and the weight index between the accuracy and the recall rate.
[0041] The performance information of the document editing control model is determined according to the model evaluation index, and the performance information is used to update the document editing control model.
[0042] In a possible implementation, before obtaining the electronic medical document, the method further includes:
[0043] Obtaining identity authentication information of the user, the identity authentication information including at least one of ID credential information, biometric information, and target factor authentication information;
[0044] Based on the identity authentication information, the user authentication result is determined, and the user authentication result is used to determine the user's operational authority to edit and process electronic medical documents in the mobile document editor control system.
[0045] In a second aspect, an embodiment of the present application provides a mobile document editor control system, including: a user terminal, a server, and a database;
[0046] A user terminal is used to provide an editing interface to the user and generate corresponding electronic medical documents when the user edits the text;
[0047] A server, used to obtain electronic medical documents in a user terminal;
[0048] The server is further used to input the electronic medical document into the document editing control model to obtain text error operation information corresponding to the electronic medical document. The document editing control model is a model obtained by training a language model built based on a preset encoder and a self-attention layer to identify text errors in the electronic medical document and generate corresponding operation information. The text error operation information includes error detection information and error correction information.
[0049] The user terminal is also used to display text error operation information to the user;
[0050] The server is further used to receive a user correction instruction sent by the user according to the text error operation information;
[0051] The server is further configured to perform a correction operation on the electronic medical document according to the text error operation information to obtain a target electronic medical document if the user correction instruction indicates to correct the text of the electronic medical document;
[0052] The database is used to store electronic medical documents, text error operation information, user correction instructions, target electronic medical documents and document editing control models.
[0053] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the first aspect above and / or various possible implementation methods of the first aspect.
[0054] In a fourth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above first aspect and / or various possible implementations of the first aspect.
[0055] The present application provides a mobile document editor control method and system, which obtains an electronic medical document, which is a document obtained by a user performing text editing in a mobile document editor control system; inputs the electronic medical document into a document editing control model to obtain text error operation information corresponding to the electronic medical document, wherein the document editing control model is a model obtained by training a language model constructed based on a preset encoder and a self-attention layer for identifying text errors in electronic medical documents and generating corresponding operation information, and the text error operation information includes error detection information and error correction information; receives a user correction instruction sent by the user according to the text error operation information; if the user correction instruction indicates to perform text correction on the electronic medical document, then according to the text error operation information, a correction operation is performed on the electronic medical document to obtain a means of obtaining a target electronic medical document, and uses the document editing control model to perform text error recognition on the electronic medical document edited by the user and provide error detection results and correction suggestions, and the user can send correction instructions according to these suggestions, so that the system can automatically perform text correction operations, thereby, the method can realize automatic error detection and correction of electronic medical documents, improve the accuracy and efficiency of document editing, reduce the workload of manual proofreading, and thus improve the overall quality of medical documents. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0057] Figure 1 A schematic diagram of a mobile document editor control system provided in an embodiment of the present application;
[0058] Figure 2 A schematic diagram of a user terminal in a mobile document editor control system provided in an embodiment of the present application;
[0059] Figure 3 A flowchart of a method for controlling a mobile document editor provided in an embodiment of the present application;
[0060] Figure 4 A training flow chart of a document editing control model provided in an embodiment of the present application;
[0061] Figure 5 A schematic diagram of the structure of a mobile document editor control device provided in an embodiment of the present application;
[0062] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0063] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0064] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0065] It should be noted that the identity information (including but not limited to the user's name, ID number, other identity authentication information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0066] In the prior art, when medical staff edit medical documents, whether using a computer terminal or a mobile device, they must first log in to the electronic medical record system platform; during the editing process, faced with a variety of medical-related documents such as anesthesia records, health examination forms, follow-up records, imaging reports, examination reports, etc., the system will guide the filling of the content based on a preset fixed template, and medical staff need to select appropriate diagnosis categories, treatment methods and other information from the drop-down menu provided by the system, or supplement specific medical terms and detailed descriptions by manual input; after completing the editing, these medical documents are stored in a semi-structured format in the database for subsequent query, call and management, but this storage method limits the further use and analysis efficiency of the data to a certain extent. At the same time, in actual operations, medical staff may edit medical documents with quality problems due to improper template selection or incorrect information entry, such as incomplete information, non-standard format, inaccurate use of terminology, etc. This not only affects the accuracy and readability of medical records, but also brings many inconveniences to subsequent medical diagnosis, data analysis, and medical quality assessment. It hinders the full realization of the advantages of the electronic medical record system and makes it difficult to meet the needs of modern medicine for high-quality and refined document management.
[0067] In order to solve the above problems, the embodiment of the present application provides a mobile document editor control method and system. Based on the trend of modern medical digitization and mobile office, the electronic medical documents edited and generated by users in the mobile document editor control system are used as the starting point. A document editing control model constructed and trained by a preset encoder and a self-attention layer is designed. The model has the ability to accurately identify various text errors in electronic medical documents and generate detailed and practical operation information. At the same time, a user interaction feedback mechanism is designed, and the user's judgment and professional knowledge are referred to, which can avoid the misjudgment problem caused by relying solely on the automatic correction of the model. Therefore, through the accurate identification and correction suggestions of text errors by the document editing control model, as well as the user's confirmation and system execution, various problems such as typos, grammatical errors, improper use of terms, etc. in electronic medical documents can be effectively reduced, and the quality of electronic medical documents can be improved, so that they are more in line with the strict standards and requirements of the medical industry, and provide a more reliable basis for medical decision-making.
[0068] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0069] Figure 1 A schematic diagram of a mobile document editor control system provided in an embodiment of the present application, such as Figure 1As shown, the mobile document editor control system provided by the present application includes a user terminal, a server and a database. Among them, the user terminal is used to provide an editing interface to the user and generate a corresponding electronic medical document when the user performs text editing; the server is used to obtain the electronic medical document in the user terminal; the server is also used to input the electronic medical document into the document editing control model to obtain the text error operation information corresponding to the electronic medical document. The document editing control model is a model obtained by training a language model built based on a preset encoder and a self-attention layer to identify text errors in electronic medical documents and generate corresponding operation information. The text error operation information includes error detection information and error correction information; the user terminal is also used to display the text error operation information to the user; the server is also used to receive the user correction instruction sent by the user according to the text error operation information; the server is also used to correct the electronic medical document according to the text error operation information if the user correction instruction indicates to correct the text of the electronic medical document; the database is used to store the electronic medical document, the text error operation information, the user correction instruction, the target electronic medical document and the document editing control model. In one example, the user terminal can be connected to a touch screen, keyboard, voice key, or image scanner of a mobile device to provide editing and display functions for electronic medical documents, as well as display typo detection and correction information.
[0070] Optionally, Figure 2 A schematic diagram of a user terminal in a mobile document editor control system provided in an embodiment of the present application, such as Figure 1As shown, the user terminal may include a responsive interface module, a device adaptation middleware, a micro front-end module, a communication and data synchronization module, a multi-language support module, and a multi-language text processing engine. Among them, the responsive interface module can adopt a responsive interface design method, combined with an adaptive layout framework to build the front-end interface of the system, and automatically detect the size, resolution, and pixel density of the device screen to adjust the front-end interface layout and element size in the user terminal to adapt to the operation requirements of different devices; the device adaptation middleware is used to create device adaptation middleware, monitor the access device type in real time, perform adaptation operations for different operating system characteristics, convert touch gestures into editing instructions, and coordinate hardware resource calls to ensure the consistency and fluency of the system on various devices. The micro front-end module is used to divide multiple independent front-end modules based on micro front-end technology, develop adaptation layers for different platforms, and realize seamless integration and functional consistency of the system on platforms such as WeChat, DingTalk, and browsers; the communication and data synchronization module is used to establish a unified communication protocol and data interaction interface to ensure data synchronization and update between different platforms and support seamless operation experience across platforms. In order to provide multi-language support functions, the multi-language support module is used to build a language resource library, dynamically load the corresponding language resources according to the user's device language settings or manually selected language preferences, and realize the multi-language display of the interactive interface and operation guide in the user terminal; in order to ensure the text language type of real-time text processing input, it is necessary to develop a multi-language text processing engine with built-in grammar rules and vocabulary libraries of multiple languages to make judgments through intelligent language recognition algorithms. Through the collaborative work of these modules, the flexibility and scalability of the mobile document editor control system are improved, ensuring that the system can provide a consistent and smooth user experience on different devices and platforms. The modular design not only simplifies the maintenance and management of the system, but also promotes collaborative development, shortens the product iteration cycle, and enables the system to quickly adapt to user needs and industry changes.
[0071] Among them, the mobile document editor control method provided in the embodiment of the present application is applied to the server in the mobile document editor control system, wherein the server can be a mobile phone, a computer, a tablet and other devices. This embodiment does not impose any special restrictions on the selection type of the server, as long as the electronic medical document can be obtained, the electronic medical document is a document obtained by the user performing text editing in the mobile document editor control system; the electronic medical document is input into the document editing control model to obtain text error operation information corresponding to the electronic medical document, the document editing control model is a model obtained by training a language model constructed based on a preset encoder and a self-attention layer for identifying text errors in electronic medical documents and generating corresponding operation information, the text error operation information includes error detection information and error correction information; receiving a user correction instruction sent by the user according to the text error operation information; if the user correction instruction indicates to correct the text of the electronic medical document, then according to the text error operation information, the electronic medical document is corrected to obtain the target electronic medical document.
[0072] Figure 3 A flowchart of a method for controlling a mobile document editor provided in an embodiment of the present application is shown in FIG. Figure 3 As shown, the method may include:
[0073] S201. Obtain electronic medical documents. Electronic medical documents are documents obtained by users through text editing in a mobile document editor control system.
[0074] Among them, electronic medical documents can refer to various written records related to the medical process, including but not limited to electronic medical records, anesthesia records, health examination forms, follow-up records, imaging reports, examination reports, etc. After editing, these documents are presented in electronic form on the mobile document editor control system, recording the patient's condition information, diagnosis process, treatment measures, and medical staff's observations and analysis, etc., and are an important carrier for the digitization of medical information.
[0075] In this step, users can be understood as individuals who use the system to create, edit, and manage medical documents, such as medical staff responsible for recording patient diagnosis, treatment plans, and progress information, patients who are allowed to access or input certain medical information, or their authorized agents.
[0076] For example, in their daily medical work, medical staff use mobile devices (such as using mobile phones during ward rounds) to open the mobile document editor control system, and use the input method provided by the system to input the patient's basic information (such as name, age, gender), condition description (such as symptoms, onset time), preliminary diagnosis results, and treatment measures taken (such as medication, examination item arrangements), etc. word by word. The system records and saves the input text information in real time to form an initial electronic medical document for subsequent viewing and organization.
[0077] S202. Input the electronic medical document into the document editing control model to obtain text error operation information corresponding to the electronic medical document. The document editing control model is a model trained based on a language model constructed based on a preset encoder and a self-attention layer to identify text errors in electronic medical documents and generate corresponding operation information. The text error operation information includes error detection information and error correction information.
[0078] In this step, it should be noted that the preset encoder and self-attention layer are key components in the document editing control model structure, but they are not necessarily the entire model. The preset encoder can be responsible for converting the text information in the input electronic medical document into a vector representation that the computer can understand and process, thereby extracting the features of the text, such as the RoBERTa encoder, which uses the RoBERTa pre-trained model to deeply encode the embedded vector to extract higher-level semantic information; the self-attention layer enables the model to automatically pay attention to the semantic associations between different parts when processing the text, so that the model can better understand the contextual meaning of the text, and enable the model to learn common typos, grammatical structures, semantic logic, and various possible text error patterns in medical texts, so that it can accurately identify errors in electronic medical documents and generate corresponding operation information to help correct these errors.
[0079] In one example, after obtaining the electronic medical document, the system automatically transfers it to the document editing control model. The model first performs word segmentation on the text in the document, splits the text into multiple word units, and then converts these words into vector form through a preset encoder and sends them to the self-attention layer for semantic analysis. For example, the model determines that "chest" in "patients have no chest shortness of breath" is a typo and should be corrected to "chest tightness", and generates error detection information to point out the location of the wrong word "chest", and error correction information provides the correct word "chest tightness".
[0080] S203: Receive a user correction instruction sent by the user according to the text error operation information.
[0081] Among them, the user correction instruction can be an instruction made by the user on whether to correct the detected erroneous text and how to correct it based on the text error operation information provided by the model, combined with his or her own professional knowledge, understanding of the patient's actual situation and other information.
[0082] For example, after seeing the text error operation information fed back by the model on the user terminal, the medical staff carefully read the error detection information to confirm whether the error pointed out by the model actually exists, and decide whether to accept the error correction information provided by the model based on their own judgment. If the medical staff believes that the model's judgment is correct, they can send a confirmation correction instruction to the system by clicking the "Correct" button on the user terminal interface or other similar operations; if they believe that the model's judgment is wrong, the medical staff can choose to ignore the error prompt, or manually modify the error and then send the instruction to inform the system that the editing of this part of the content has been completed.
[0083] S204. If the user correction instruction indicates to correct the text of the electronic medical document, a correction operation is performed on the electronic medical document according to the text error operation information to obtain a target electronic medical document.
[0084] Among them, the target electronic medical document is the final version obtained after error correction and optimization of the original electronic medical document through the above series of steps. It has higher accuracy, standardization and readability, and can better meet various needs in medical work.
[0085] The mobile document editor control method provided in the embodiment of the present application uses a document editing control model to quickly and accurately detect various text errors that may exist in electronic medical documents, including typos, grammatical errors, improper use of terms, etc., thereby improving the quality and accuracy of documents and reducing misunderstandings of medical information and decision-making errors caused by document errors. At the same time, users can flexibly perform corrective operations based on system feedback to ensure that the final document meets medical standards and requirements and avoid possible misjudgments by the model. As a result, this method not only improves the quality and accuracy of documents, but also reduces the time and effort of manual proofreading, thereby improving the overall efficiency of medical work.
[0086] On the basis of the above embodiment, the document editing control model includes a recognition module and a classification module. The electronic medical document is input into the document editing control model to obtain text error operation information corresponding to the electronic medical document, including: inputting the electronic medical document into the recognition module to extract the hidden state vector of each word in the electronic medical document; inputting the hidden state vector of each word in the electronic medical document into the classification module, and in the classification module, judging the error detection information of the electronic medical document, and generating error correction information corresponding to the error detection information, wherein the text error operation information includes error detection information and error correction information.
[0087] Furthermore, the recognition module can be used to perform in-depth feature extraction on the input electronic medical documents. When the electronic medical documents are input into the recognition module, it will use the preset natural language processing technology to disassemble and analyze the text information in the documents, and extract the hidden state vector corresponding to each word through complex algorithm calculations. These hidden state vectors contain the semantic, grammatical and positional feature information of the word in the entire text context, providing basic data for the accurate judgment of the subsequent classification module.
[0088] The classification module can use the rich information carried by these hidden state vectors based on the pre-trained classification model to comprehensively scan and judge the electronic medical documents. First, it will accurately determine whether there are text errors in the electronic medical documents based on the built-in medical text rule library and the error pattern library formed by a large amount of medical document sample data, and determine the error detection information. For example, it determines that "Yinyin" in "The patient's lungs have Yinyin, suspected of having pneumonia" is a typo (corresponding to error detection information); then, based on the knowledge system and error correction algorithm learned through training, it generates correction suggestions for the error, that is, "Yinyin" should be corrected to "Shadow" (corresponding to error correction information).
[0089] The embodiment of the present application divides the document editing control model into the design of the recognition module and the classification module, making the entire process of obtaining text error operation information clearer, more efficient and accurate. Through the accurate extraction of word vectors by the recognition module, it can provide high-quality input data for the classification module, making the classification module more targeted and accurate in judging and generating error correction information, better coping with the complexity and professionalism of medical texts, and effectively improving the recognition rate and correction accuracy of various text errors in electronic medical documents, thereby further improving the quality of electronic medical documents.
[0090] On the basis of the above embodiment, the recognition module includes an input layer, an embedding layer, a preset encoder and a self-attention layer, the output of the input layer is used as the input of the connected embedding layer, the output of the embedding layer is used as the input of the connected preset encoder, and the output of the preset encoder is used as the input of the connected self-attention layer; the classification module includes a classifier and an output layer; wherein: the input layer is used to receive the electronic medical document; the embedding layer is used to convert each word of the electronic medical document in the input layer into a word vector; the preset encoder is used to semantically encode the word vector corresponding to each word in the embedding layer based on a multi-layer Transformer encoder to obtain a hidden state vector of each word; the self-attention layer is used to perform self-attention processing on the hidden state vector of each word in the preset encoder to obtain a target hidden state vector of each word; the classifier is used to perform error detection classification on the target hidden state vector of each word in the self-attention layer to obtain a detection label sequence corresponding to the electronic medical document, and perform error correction classification on the target hidden state vector of each word to obtain a correction label sequence corresponding to the electronic medical document; the output layer is used to generate text error operation information corresponding to the electronic medical document according to the detection label sequence and correction label sequence output by the connected classifier.
[0091] In this embodiment, it can be understood that the input layer, as the starting part of the entire recognition module, is the entrance for receiving electronic medical documents and provides original text information for subsequent processing procedures. For example, when an electronic medical record containing a patient's symptom description, diagnosis results, and treatment process is input into the system, the input layer will receive the text content of the medical record in its entirety and pass it to the embedding layer.
[0092] The embedding layer can use specific word vector conversion technology to convert each word in the document into a word vector with a certain dimension and semantic representation, that is, to convert it into a numerical vector form that the system model can better understand and process for subsequent semantic encoding and analysis. For example, for the words "patient", "temperature", "increase" in the sentence "the patient's body temperature rises", the embedding layer will map them to corresponding word vectors respectively. These word vectors not only contain the basic semantic information of the words, but also may cover their specific meanings and contextual relationships in the medical field. Optionally, the embedding layer can be a combined embedding layer formed by word embedding, paragraph embedding, and position embedding.
[0093] The preset encoder can be based on the architecture of a multi-layer Transformer encoder, which performs deep semantic encoding operations on each word vector output by the embedding layer. The Transformer encoder can capture the complex relationship between words and generate a hidden state vector for each word. For example, when processing a text such as "the patient developed symptoms of coughing and sputum, which lasted for three days", the preset encoder can capture the semantic associations between words such as "cough", "sputum", and "lasting" as well as their semantic roles in the entire sentence, thereby generating a hidden state vector that can accurately reflect these semantic features, providing high-quality data support for subsequent self-attention processing. For example, the encoder can use a standard Transformer encoder or a Transformer model that has been fine-tuned in the medical field.
[0094] The self-attention layer can further perform self-attention processing on the hidden state vector of each word output by the preset encoder. The self-attention mechanism allows the model to dynamically assign attention weights according to the semantic relationship between each word and other words when processing text, so as to more accurately capture the long-distance dependencies and semantic information in the text. For example, when analyzing the sentence "The patient's condition gradually improved after antibiotic treatment, but still needs to be observed", the self-attention layer can pay more attention to the relationship between words such as "condition", "improvement", and "continue observation" according to the contextual semantics, thereby obtaining the optimized target hidden state vector for each word. These vectors can better reflect the importance and semantic characteristics of words in the entire text context, and provide more accurate input information for the classification module.
[0095] Optionally, the self-attention layer can be designed to include a combination of multiple attention mechanisms to enhance the expressiveness and adaptability of the model. For example, a self-attention layer is composed of a combination of a self-attention network, a graph attention network, and an interactive attention network, wherein the combination method can be a series combination (different types of attention layers are connected in series so that the output of each layer becomes the input of the next layer to gradually extract and integrate information at different levels), a parallel combination (multiple attention mechanisms are applied to the input data in parallel, and then their outputs are merged to capture different types of relationships and features at the same time) or a hybrid combination (combining series and parallel methods). By combining multiple attention mechanisms, the model can handle complex data structures more flexibly and capture richer features and relationships. Compared with a single attention mechanism, the combined multi-layer attention layer provides stronger expressiveness and adaptability, and can better meet the needs of different application scenarios.
[0096] The classifier is the core component of the classification module, which is used to classify the target hidden state vector output by the self-attention layer. It can use a simple fully connected neural network as a classifier, or a more complex deep learning model. For example, if the document contains "the patient's white blood cell count is high", the classifier determines that "white blood cell" should be standardized as "leukocyte" and marks the word "white blood cell" as an error in the detection label sequence; then for the error of "white blood cell", the correction label sequence may be "leukocyte" and other possible similar correct words (if any) and their corresponding probability values to indicate the credibility of each correction suggestion.
[0097] The output layer integrates and processes the detection label sequence and correction label sequence output by the classifier to generate the final text error operation information corresponding to the electronic medical document. This includes organizing and formatting the detected error location, error type, and corresponding correction suggestions so that it can be clearly and accurately fed back to the user or other subsequent processing modules. For example, the output layer may organize the error information of "white blood cells" into text error operation information such as "In row X, column Y, there is a terminology error in 'white blood cells', and it is recommended to correct it to 'white blood cells'", thereby providing clear guidance and basis for error correction of electronic medical documents.
[0098] The embodiment of the present application can fully extract and understand the semantic information of the text through the multi-layer processing of the recognition module, providing a solid foundation for the accurate judgment of the classification module; and the dual classification mechanism of the classification module can comprehensively and meticulously detect errors in the text and provide reasonable correction suggestions, thereby improving the ability to recognize and correct various text errors in electronic medical documents, helping to improve the quality and standardization of electronic medical documents, and reducing medical risks and misunderstandings caused by document errors.
[0099] Based on the above embodiment, the classifier includes an error detection classifier and an error correction classifier; wherein: the error detection classifier is used to receive the target hidden state vector of each word in the self-attention layer, and according to the target hidden state vector of each word, use the maximization layer to predict the error state of each word to obtain a detection label sequence, and each label in the detection label sequence is used to indicate whether the corresponding word has an error; the error correction classifier is used to perform operation prediction on the target hidden state vector of each word to obtain a correction label sequence, and each label in the correction label sequence is used to indicate the specific correction operation of the corresponding word.
[0100] In this embodiment, the main task of the error detection classifier is to determine whether there is an error in each word in the text. For example, by inputting these target hidden state vectors into a pre-built neural network structure, the classifier uses its learned language patterns and text feature knowledge in the medical field to predict the error state of each word. This prediction process is based on the principle of maximization layer. By comparing the probability values of different possible states (wrong or correct), the state with the highest probability is selected as the prediction result, thereby generating a detection label sequence for the entire electronic medical document to quickly locate words that may have errors. Among them, the detection label sequence can be binary, such as "1" (no error) and "0" (error); it can also be multivariate, indicating different types of errors (such as spelling errors, grammatical errors, etc.).
[0101] The error correction classifier also receives the target hidden state vector for each word, but the focus is on predicting specific correction operations for these erroneous words. Optionally, using a large medical text corpus for training, the error correction classifier learns common correction methods and possible correct word replacements for different error types, and for each word in the text that is judged to be erroneous, the error correction classifier can provide one or more possible correction labels. Among them, the correction label sequence may include "1" (keep the current word unchanged), "2" (delete the current word), "3(n)" (add word n after the current word), "4(n)" (add word n before the current word) and "5(n)" (replace the current word with word n).
[0102] The embodiment of the present application divides the classifier into two parts: error detection and error correction, which can handle errors in electronic medical documents more finely. The error detection classifier first identifies words that may be erroneous, and then the error correction classifier provides specific correction suggestions, which can improve the accuracy of error identification and make the correction process more efficient and accurate.
[0103] On the basis of the above embodiment, the classification module also includes a specific layer, and the output of the output layer serves as the input of the connected specific layer; wherein: the specific layer is used to perform information update processing on the text error operation information in the output layer to obtain updated text error operation information, and the information update processing is the processing of the text error operation information in the output layer based on the behavioral characteristics of the target object, and the target object is the object corresponding to the electronic medical document input into the document editing control model.
[0104] In this embodiment, it can be understood that the specific layer plays a personalized customization role in the entire document editing control model. After the output layer generates preliminary text error operation information, this information will be input into the specific layer. The core task of the specific layer is to update and process this information based on the behavioral characteristics of the target object. The target object here has multiple dimensions of definition, which can be a specific doctor, a specific department or department, or even different types of medical documents, such as electronic medical records, various medical-related records and reports, etc., wherein the target object can be manually selected by the user when editing the document, or it can be determined by the identity verification information of the user logging into the system.
[0105] Furthermore, from the perspective of a specific doctor, different doctors usually have their own unique language habits and expressions when writing medical documents. For example, some doctors may prefer to use certain specific medical term abbreviations, or have their own preferred sentence structures when describing the condition. The specific layer can identify these personalized feature patterns by analyzing and learning a large number of medical documents written by the doctor in the past. When processing new electronic medical documents written by the doctor, even if some situations that may be misjudged as errors by the general model (i.e., the document editing control model without the specific layer) are encountered (in fact, the doctor's habitual expression), the specific layer can adjust and update the preliminary text error operation information based on the learned behavioral characteristics of the doctor, avoid unnecessary error prompts, thereby improving the accuracy and adaptability of the doctor's document correction and reducing interference with his normal writing habits.
[0106] For specific departments or divisions, different departments have obvious differences in the writing standards and terminology of medical documents due to differences in their professional fields. For example, the medical records of the Department of Cardiology may frequently involve descriptions of professional examination indicators and treatment methods related to the heart, while the Department of Dermatology will focus more on terms related to skin symptoms, pathological diagnosis, etc. The specific layer mines and trains a large amount of historical medical document data of a specific department or department to master the vocabulary, grammar and semantic rules unique to its professional field. When processing the electronic medical documents of the department or department, it can more accurately identify real text errors and provide more practical error correction suggestions based on its professional characteristics, improve the pertinence and practicality of text error operation information, and meet the personalized needs of different departments in medical document management.
[0107] From the perspective of different types of medical documents, electronic medical records differ from other medical records, reports, etc. in terms of format, content focus, and language style. Electronic medical records need to record in detail the patient's condition evolution, treatment process, medication status, and other information, and the language is relatively detailed and standardized; while some examination reports may focus more on presenting examination results and key data in a concise and clear manner. The specific layer is trained according to the characteristics of different types of documents, and can optimize text error operation information according to the type characteristics of the document. For example, for an imaging examination report, the specific layer will pay more attention to the accuracy of terms related to imaging descriptions and diagnostic conclusions, while for electronic medical records, the text quality of various aspects such as condition description, diagnostic basis, and treatment plan will be comprehensively considered to ensure that different types of medical documents can receive error correction and optimization that best suits their characteristics, and improve the overall quality and readability of various types of medical documents.
[0108] Based on the consideration of individual differences and differences in document types, the embodiments of the present application can achieve the following effects by introducing a specific layer: First, the personalized behavioral characteristics of the target object can be learned to accurately distinguish between real text errors and personalized normal expressions, avoiding a large number of misjudgments caused by the one-size-fits-all approach of the general model, so that errors in electronic medical documents can be more accurately discovered and corrected. Among them, the adaptability and flexibility of the system are enhanced. Whether facing the diverse writing habits of different doctors, or the professional characteristics of different departments and departments, and the unique requirements of different types of medical documents, the specific layer can perform targeted processing based on its pre-learned features, so that the system can better adapt to various actual scenarios in medical work, and improve the usability and practicality of the system in different environments. Finally, the user experience and work efficiency are improved. For medical staff, unnecessary error correction operations and interference with normal writing content are reduced, so that they can focus more on the diagnosis and treatment of patients, saving time and energy.
[0109] On the basis of the above embodiment, before inputting the electronic medical document into the document editing control model to obtain the text error operation information corresponding to the electronic medical document, the method may also include: obtaining a historical medical document set, the historical medical document set includes electronic medical documents stored by multiple objects in the target hospital within a preset historical time period; based on preset text error annotation requirements, annotating the historical medical document set to obtain a sample set, each sample in the sample set includes: an electronic medical document, and an editing label sequence corresponding to the electronic medical document, the editing label sequence is used to indicate that the erroneous text in the electronic medical document is corrected to the correct text; based on the sample set, training the initial document editing control model until the output information of the initial document editing control model meets the preset training requirements, and obtaining the document editing control model.
[0110] Among them, obtaining a historical medical document set can be to widely collect electronic medical documents generated and stored by multiple objects (such as different departments, different medical staff and other medical staff, different document types) in a specific preset historical time period from the information system of the target hospital. These documents cover various types of daily medical activities in the hospital, such as electronic medical records, surgical records, test reports, medical orders, etc., forming a rich and diverse set of original data. The purpose of this step is to ensure the comprehensiveness and representativeness of the data, so that it can reflect the use of texts in different scenarios in the actual medical work of the hospital, and provide sufficient and real material basis for subsequent model training.
[0111] Furthermore, annotation based on preset text error annotation requirements can refer to organizing professional medical personnel and language experts to perform detailed annotation work on each document in the historical medical document collection based on a set of pre-formulated text error annotation requirements. The annotation process should not only accurately identify various errors in the text, such as typos, grammatical errors, improper use of medical terms, etc., but also clearly indicate the correct text content that each error should be corrected to, thereby generating an editing label sequence corresponding to each electronic medical document. This precise annotation work provides clear learning objectives and supervision information for model training, allowing the model to understand what kind of text is wrong and how to correct it.
[0112] Furthermore, the annotated sample set is input as training data into the initial document editing control model. During the training process, the model learns the mapping relationship between the text features of the electronic medical documents in the sample and the editing label sequence by continuously adjusting its own parameters. For example, the model will learn the association between certain specific vocabulary combinations, grammatical structures, and contextual features and common error types, as well as the correct correction methods for these errors. The training process will continue until the output information of the model meets the preset training requirements. The preset training requirements can be that the model can accurately identify text errors in the input electronic medical documents and generate error correction information that is highly consistent with manual annotations, or the number of model training iterations reaches a preset value, at which point the model becomes the final usable document editing control model. Optionally, during the training process, strategies such as early stopping are used to avoid overfitting.
[0113] Correspondingly, when the structure of the document editing control model includes a specific layer, after the document editing control model is trained, the sample corresponding to the target object can be selected from the sample set, and the selected sample can be annotated with the behavioral characteristics of the target object (such as the common terms or abbreviations of a specific doctor / department) and stored in the editing label sequence to obtain an updated sample set. Then, the document editing control model is trained using the updated sample set, and the parameters in the specific layer (such as the structure, the fine-tuned learning rate, and the amount of training data) are adjusted until the output information of the document editing control model meets the preset training requirements, and the adjusted document editing control model is obtained to improve the accuracy of error recognition and correction of the target object. Optionally, the parameters close to the output layer in the document editing control model can also be fine-tuned. At this time, the parameters of other layers in the model can be frozen to improve the model training effect. Among them, the electronic medical documents stored by the target object in the target hospital within a preset historical time period can also be obtained, and the updated sample set is prepared according to the above method.
[0114] It should be noted that the sample set can be updated and generated by regularly collecting new data, and the model can be continuously trained based on the updated sample set to maintain the accuracy and effectiveness of the model.
[0115] Furthermore, after obtaining the document editing control model, the model is updated according to a certain period or when specific conditions are met. In one example, a regular inspection mechanism is set in the system, for example, the analysis of newly collected medical document data is automatically started at fixed time periods (such as every week, every month, etc.), to determine whether there is enough new data accumulation and whether the newly emerging text error patterns reach a certain threshold ratio (for example, the newly emerging typos that are not accurately recognized by the model account for one thousandth of the total number of words in the recently added documents, etc.), which is used as a condition to trigger the adaptive update. When such conditions are met, the adaptive update process of the model is started.
[0116] Furthermore, (1) incremental data collection and screening: continuously collect electronic medical document data generated by multiple subjects in a new time period from the target hospital information system to form a new data set to be updated; through the intelligent text mining algorithm (the algorithm is used to quickly identify new typos and incorrect expressions in the text), screen out electronic medical documents containing suspected new typos, new incorrect expressions and text situations that have frequently appeared recently, so as to reduce the amount of data processing and focus on the data parts that may affect the accuracy of the model; (2) real-time collaborative annotation: for the screened electronic medical documents, start the real-time collaborative annotation process, through the online collaborative annotation platform (the platform is integrated into the system based on the existing real-time communication platform technology, supporting real-time communication and collaboration of annotators), invite medical staff, medical editors and language experts from different departments as annotators, and annotate according to the preset text error annotation requirements; during the annotation process, the system automatically records The system will also provide real-time information on the parts that need further annotation based on the annotated data. (3) Small sample incremental learning strategy: The annotated new sample data is integrated into the existing document editing control model using a small sample incremental learning strategy for training. The differences between the new and old samples in terms of text features and error patterns are analyzed, and the relevant parameters in the model are adjusted in a targeted manner (such as adjusting the parts related to the differences in the new samples) to quickly adapt to new text error changes. (4) Online evaluation and feedback adjustment: The reserved new samples are used as online test sets to evaluate the performance of the updated model. The effectiveness of the adaptive update is judged by comparing it with the previously recorded model evaluation indicators. If some indicators do not meet expectations, the system will automatically backtrack and analyze the data annotation or learning algorithm parameter setting issues, and make adjustments again until the model meets the preset updated performance requirements under the new data set.
[0117] The embodiment of the present application obtains the historical document set of the target hospital for targeted training, which can adapt to its unique internal terminology habits, department culture and writing style, improve the accuracy and adaptability of the model, reduce errors, and meet the hospital's strict requirements for document quality; professional personnel mark in detail according to preset requirements to provide high-quality samples for the model, so that it can deeply learn error patterns and correct expressions, effectively handle various errors, and improve the quality and efficiency of the editing process; use updated sample sets to train the model, so that the model evolves with the development of the hospital's business and changes in text habits, maintains high-accuracy processing capabilities, and ensures that the electronic medical record system provides high-quality services in the long term.
[0118] On the basis of the above embodiments, based on preset text error annotation requirements, a method for annotating a historical medical document set to obtain a sample set may include: annotating the historical medical document set based on preset text error annotation requirements to obtain a text error position corresponding to each electronic medical document in the historical medical document set and a correct text after the error text is corrected; for each electronic medical document, filtering each electronic medical document according to a preset vocabulary to be filtered to obtain a filtered electronic medical document; performing minimum edit distance processing on the filtered electronic medical document according to the text error position and correct text corresponding to the filtered electronic medical document to obtain a corrected target text, as well as a text error type and an error correction operation of the filtered electronic medical document; determining an edit tag sequence corresponding to the electronic medical document according to the text error position, correct text, text error type, error correction operation and target text corresponding to the electronic medical document; and constructing a sample set according to all electronic medical documents and the edit tag sequence corresponding to each electronic medical document.
[0119] In this step, it can be understood that, based on the preset text error annotation requirements, professionals will carefully review each electronic medical document in the historical medical document collection, accurately find the text error position therein, and clarify the correct text after the corresponding error text is corrected. Optionally, a detailed annotation guide is formulated for reference by the annotation personnel, or an automated tool is used for preliminary annotation, and then manually reviewed to improve efficiency and accuracy. Next, for each electronic medical document, the words are filtered according to the preset vocabulary to be filtered, and unnecessary words (such as stop words) are removed to simplify subsequent processing. Then, based on the determined text error position and correct text corresponding to the filtered electronic medical document, the minimum edit distance processing method is used to calculate the minimum number of editing operations and specific operations (such as replacement, deletion, insertion of characters, etc.) required to transform from the error text to the correct text, thereby obtaining the corrected target text, and clarifying the text error type (such as typos, non-standard terms, etc.) and specific error correction operations. Afterwards, the text error position, correct text, text error type, error correction operation, and target text corresponding to the electronic medical document are integrated to determine the complete edit label sequence, which is used to guide the model to learn how to generate correct text from incorrect text. Finally, all electronic medical documents and their corresponding edit label sequences are summarized to construct a sample set, which can be stored in a structured data format, such as JSON or CSV, for subsequent processing and model input.
[0120] The embodiment of the present application generates a high-quality sample set through a systematic annotation and processing process, which significantly improves the effectiveness of model training. Compared with the prior art, it emphasizes the refined annotation of text errors and the clarification of correction operations, especially through the calculation of the minimum edit distance, which realizes the accurate modeling of the correction process.
[0121] On the basis of the above embodiment, the initial document editing control model is trained based on the sample set until the output information of the initial document editing control model meets the preset training requirements and the document editing control model is obtained. The method may also include: obtaining a test medical document set, wherein the test medical document set is different from the preset historical time period of the historical medical document set; determining a test sample set based on the test medical document set, wherein each test sample in the test sample set includes: an electronic medical document, and an editing label sequence corresponding to the electronic medical document; inputting the test sample set into the document editing control model to obtain text error operation test information corresponding to each test sample; determining a model evaluation index based on the editing label sequence in each test sample and the corresponding text error operation test information, wherein the model evaluation index includes the accuracy of model prediction, the recall rate, and a weight index between the accuracy and the recall rate; determining performance information of the document editing control model based on the model evaluation index, wherein the performance information is used to update the document editing control model.
[0122] In this step, after completing the training of the initial document editing control model based on the historical medical document set, a test medical document set of a different time period from the historical medical document set used for training is further selected to help evaluate the generalization ability of the model under different data distributions. Then, a test sample set is determined from the test medical document set in the same way, each test sample contains an electronic medical document and its corresponding editing label sequence. Subsequently, the test sample set is input into the trained document editing control model, and the model outputs the text error operation test information corresponding to each test sample. Then, by comparing the editing label sequence in the test sample (i.e., the actual error situation and correction information) with the text error operation test information generated by the model, the model evaluation index is calculated, where the precision measures the accuracy of the model prediction, the recall rate reflects the ability of the model to find all actual errors, and the weight index balances the relationship between the two and comprehensively reflects the model performance. Finally, the performance information of the document editing control model is determined based on these model evaluation indicators, and the subsequent update of the model is guided based on this performance information, such as adjusting the model parameters, optimizing the algorithm structure, etc., so that the model continues to evolve to better adapt to the error correction task of electronic medical documents.
[0123] The embodiment of the present application selects test medical document sets from different time periods for independent testing, and uses a comprehensive model evaluation indicator system including precision, recall rate and weight indicators to comprehensively and accurately quantify the performance of the model in practical applications, providing a reliable basis for model updates, so that the model can continuously and specifically improve its ability to handle errors in electronic medical documents, thereby ensuring the high quality and stability of electronic medical document editing.
[0124] Based on the above embodiment, before obtaining the electronic medical document, the method may also include: obtaining the user's identity authentication information, the identity authentication information including at least one of ID credential information, biometric information, and target factor authentication information; determining the user verification result based on the identity authentication information, and the user verification result is used to determine the user's operating authority to edit and process electronic medical documents in the mobile document editor control system.
[0125] Before the step of obtaining the electronic medical document, the system server may provide the user with an identity authentication interface based on the user terminal to obtain the user's identity authentication information.
[0126] Optionally, the identity authentication information here covers multiple dimensions. When obtaining the user's identity authentication information, one or more dimensions of identity authentication information can be obtained according to actual security considerations to perform user identity authentication. Among them, ID credential information can be the exclusive account number, work number, and other identification information that can indicate the uniqueness of the identity of the medical staff registered in the hospital information system. With this information, it can be matched and checked with the personnel information pre-stored in the hospital backend database; biometric information includes highly individual unique biometric data such as fingerprints, facial features, and irises. The identity is confirmed by comparing the collected data with the entered biometric samples through corresponding biometric technology equipment such as fingerprint scanners and cameras; target factor authentication information can be related to some elements related to the department where the medical staff is located, the job level, and the specific business authorization, such as whether a doctor in a department has the authority to edit a certain type of special medical documents, and other additional authentication factors.
[0127] Furthermore, based on the identity authentication information obtained, the system uses the corresponding verification algorithm and process to strictly compare with the standard information stored in the background to determine the user verification result. This result is directly related to the operation permissions that the user can have when editing and processing electronic medical documents in the mobile document editor control system. For example, if the verification is passed and the person is a high-level physician, he or she has full permissions to modify and review all types of electronic medical documents; if he or she is an intern, he or she only has the permission to perform simple document entry operations.
[0128] Furthermore, in order to ensure that user information will not be leaked and that medical records are stored and shared securely, when users authenticate their identities, not only are they required to provide their authentication information, but the authentication requirements also need to be dynamically adjusted based on the user's historical behavior patterns and current environment. For example, the system can decide whether to add additional verification steps, such as a temporary one-time password (OTP) or security questions, based on factors such as the user's login time, geographic location, and device type. Another example is using blockchain technology to record the user's operation log. Every time a user operates in the system, a log record is generated. These records are stored in the form of blockchain to ensure that they cannot be tampered with and are traceable. While logging, the system uses machine learning algorithms to analyze the user's operation log in real time and detect abnormal behavior. If a user frequently tries different login methods in a short period of time, the system can automatically mark it as an abnormality and trigger a security alert or restrict account access.
[0129] Optionally, this user verification result can also be used to determine the target object, that is, when there is a specific layer in the model, users with different identities correspond to different target objects, and the user verification result can be used to determine that the target object is a specific doctor, a specific department, or a specific type of electronic medical document.
[0130] The embodiment of the present application significantly improves the security and compliance of electronic medical document editing and processing by introducing a multi-level identity authentication mechanism. It not only relies on traditional ID credentials, but also combines biometrics and target factor authentication to provide a more comprehensive and flexible identity authentication solution, which can effectively prevent unauthorized access, ensure the security and privacy protection of sensitive medical information, and enhance the overall security of the system and user trust.
[0131] Figure 4 A training flow chart of a document editing control model provided in an embodiment of the present application, such as Figure 4 As shown, in this embodiment Figure 3 Based on the embodiment, the training process of the document editing control model is described in detail, including:
[0132] S301: Data collection.
[0133] Furthermore, the electronic medical record data for data collection can be sourced from the information systems of multiple departments in a certain hospital, including outpatient prescriptions, inpatient medical orders, inspection reports, etc. Through a professional annotation team, annotators with medical knowledge and an understanding of medical record writing manually annotate the collected electronic medical record data to identify and mark the typos in it. The accuracy of the annotation results is ensured through methods such as cross-validation and expert review. The annotated data is organized into a usable dataset, including information such as the original medical record text, the positions of typos, and the correct words. The distribution of different types of text data may be unbalanced, which can lead to a lower prediction accuracy for the less numerous categories during model training. Therefore, to improve the generalization ability of the model, 1000 sentences from the original dataset are selected through an appropriate random sampling method. The sentences vary in length, and data augmentation techniques (here, replacing with synonyms or data back-translation can be used. Data augmentation is a way to improve the model) are used to convert these data into 10,000. They are divided into a training set and a test set in an 8:2 ratio, where the proportions of incorrect sentences in the training set and the test set are 77.2% and 89.3% respectively.
[0134] S302. Data preprocessing.
[0135] This step may include:
[0136] (1) Text cleaning: Remove irrelevant characters in the electronic medical records, such as meaningless symbols, numbers, etc., and handle missing values and outliers to ensure the cleanliness and integrity of the sentences.
[0137] (2) Word segmentation: Use the tokenizer tool for word segmentation and stop word filtering, that is, remove common words without actual meaning, such as "de", "he", "shi", etc., to improve the processing efficiency and accuracy of the text.
[0138] (3) Edit label extraction: The label extraction algorithm based on the minimum edit distance algorithm converts the parallel sentence pairs in the text into an edit label sequence.
[0139] Among them, the minimum edit distance refers to the minimum number of edit operations required to convert one string into another. Given an input Chinese sentence X = {x1, x2,..., x n}, x i refers to the i-th character of sentence X, and the grammar error correction model outputs a sentence Y = {y1, y2,..., y m}, y jRefers to the jth word in sentence Y. The lengths of X and Y can be equal or unequal. If the lengths are equal, then the Chinese sentence has no grammatical errors or only substitution errors. Define D(i,j) as the minimum edit distance between the strings X[1...i] and Y[1...j]. Then D(i,j) can be calculated by the following recursive formula:
[0140]
[0141] Among them, D(i-1,j) is a deletion operation; D(i,j-1) is an insertion operation; and D(i-1,j-1) is a replacement operation.
[0142] S303: Model training.
[0143] In this step, model construction is also included before model training, and the model framework is set to the input layer, embedding layer, RoBERTa encoder, self-attention layer, classifier (error detection layer, error correction layer), and editing operation output layer connected in sequence.
[0144] Furthermore, the model parameters are set as follows: the word embedding dimension, the maximum position embedding dimension and the maximum sentence length are all 100, the hidden layer unit number dimension is 768, the encoder layer number is 12, the multi-head attention head number is 12, the fully connected layer size is 768, the number of layers is 2, the feedforward network dimension is 3072, the model learning rate is 1e-05, all Dropouts in the model are set to 0.1, the activation function is GELU, the optimizer is ADAM with a learning rate of 0.002, the β value is (0.9, 0.998), the WARMUP step number is 8000, and the batch size is 128.
[0145] Furthermore, the preprocessed data is input through the model input layer. The following takes the sentence "The patient complained of chest pain, no chest tightness and shortness of breath" as an example to illustrate the model training process.
[0146] 1. Input layer: Input the preprocessed sentences into the model, that is:
[0147] X={ <s> The patient complained of chest pain but no shortness of breath.< / s>},in <s> and< / s> Used to indicate the beginning and end of text.
[0148] 2. Embedding layer: The text is first embedded as a vector to capture the semantic features of the vocabulary. Furthermore, since RoBERTa's word vector model has excellent representation effect, the text is input into the Token Embedding layer to obtain the vector representation (100 dimensions) of each word; Segment Embedding is to distinguish the identifiers of different sentences; and Position Embedding is to allow the model to learn the sequential properties of the input. Among them, the final word vector output of RoBERTa is the sum of the three Embedding output layers.
[0149] 3. RoBERTa encoder: Use the RoBERTa pre-trained model to deeply encode the embedded vector and extract higher-level semantic information. This embodiment selects Grammarly's open source GECToR as the baseline model. The model is mainly composed of a bidirectional Transformer-based encoding part and two linear layers. The first linear layer detects whether a specific word is wrong, and the second linear layer selects the correction label to be performed on the word. The error detection label includes "1" (no error) and "2" (error); the error correction label includes "1" (keep the current word unchanged), "2" (delete the current word), "3 (n)" (add word n after the current word), "4 (n)" (add word n before the current word) and "5 (n)" (replace the current word with word n).
[0150] After the encoder is processed, each word x is obtained i The feature embedding representation F(x i ), and then combine it with the last hidden state vector h output by the encoding part of the RoBERTa model i L Perform concatenation (add up) to get the final hidden state vector X i express.
[0151] ①The RoBERTa model contains a total of 12 layers of Transformer encoders. The main structure in the Transformer encoder layer is the multi-head self-attention mechanism, which maps 1 query and 1 set of key-value pairs to 1 output; the output is a weighted sum of 1 numerical value. Each multi-head self-attention calculation module consists of 12 self-attention calculation modules, and there are 768 feature dimensions in each layer. In this set, the specified weight is calculated by using a query with a key value. Multiple attention mechanisms enable the model to focus on information from different subdomains at different locations. The formula of the multi-head self-attention mechanism is as follows:
[0152]
[0153] Where Q is the query matrix; K is the key matrix; V is the value matrix; W i Q ,W i k ,W i V are the weight matrices of matrices Q, K, and V in the i-th attention head; n is the number of heads in the multi-head self-attention; d k is the size of the query value or key value; s represents the softmax function, where the attention score is During the calculation process, the attention function is calculated for the same set of vectors at the same time and then sent to the matrix Q; the key and value are also sent to the matrix K and V respectively. The function A is used to output the attention calculation, the function s is used to compare the input weights, and h i The function is used to calculate the i-th attention head, the function of the M function is used to calculate multi-head self-attention. The C function is the folding function.
[0154] ② Feedforward Neural Network:
[0155] Each attention output is input into a feed-forward neural network, a two-layer fully connected network, which has the same structure at all positions in the model.
[0156] FFN(x)=ReLu(Wx+b);
[0157] Among them, W is the weight matrix and b is the bias term. The output of the last layer of feedforward neural network is:
[0158] FFNoutput=W2·h+b2,h=ReLu(W1·A+b1);
[0159] Among them, W1, b1 are the weight and bias of the first layer respectively; W2, b2 are the weight and bias of the second layer respectively; h is the hidden layer output calculated in the previous step.
[0160] ③Residual connection layer and BatchNomm batch normalization:
[0161] Each sublayer in the RoBERTa model has a residual connection. To prevent the gradient vanishing and exploding problems, the output of the sublayer is added to its input, and the result is layer normalized to ensure that it has a stable distribution, i.e.: Residual = FFNoutput + A, Output = LayerNorm (Residual).
[0162] 4. Self-attention layer: Add a self-attention layer or multiple attention layers between the last model network layer and the classification layer to make the semantic information deeper.
[0163] The steps to build a self-attention layer are as follows:
[0164] Step 1: Input sequence (the input of the self-attention mechanism layer consists of the hidden state vector output by the RoBERTa network layer);
[0165] Step 2: Create three vectors: query vector (Query), matching vector (Key), and target vector (Value). The query, key, and value of self-attention are obtained from the same input sequence X through linear transformation;
[0166] Step 3: Define the encoder input sequence as X and the training weight as W Q ,W k ,W V , which are randomly generated weight matrices. Multiply each input vector by the weight matrix to get Q = XW Q ,K=XW k ,V=XW V ;
[0167] Step 4: Calculate the dot product between the query vector Q and the matching vector K. To prevent the dot product from being too large, divide the result by the square root of the dimension of the matching vector K.
[0168] Step 5: Normalize the dot product result using the Softmax function to get the attention score;
[0169] Step 6: The target vector V is weighted and summed with the attention score, which is the output of self-attention, namely:
[0170] It should be noted that when adding multiple attention layers, different types of attention layers can be used, such as combining self-attention network, interactive attention network, graph attention network, etc. to construct the final attention layer. Among them, the graph attention network can obtain the grammatical information of the text during editing; the interactive attention network can fuse the semantic information extracted by the self-attention network and the grammatical information extracted by the graph attention network, so as to process and obtain the final semantic and grammatical information.
[0171] For example, the graph attention network can use the parsing tool LAL-Parser to generate the phrase structure tree and dependency tree structure diagram corresponding to the input text. The input here is the same as the input of the self-attention layer, because this layer and the self-attention layer run in parallel.
[0172] Construct a grammar graph: The resulting graph has 7 nodes (because the preprocessed sentence has seven phrases), and each layer can be represented by a different adjacency matrix CA l , the words in the same phrase can be considered to form a small fully connected graph:
[0173]
[0174] in, represents the adjacency matrix constructed at the lth level of the grammar structure tree; w i and w j Represents different words in a sentence. The dependency tree structure can also be represented as an adjacency matrix DA containing 7 nodes ij , the nodes in the graph represent the words in the sentence, and the dependencies between the words are connected by edges, that is:
[0175]
[0176] Next, the two types of grammatical information are fused and the edges are added by position, thereby obtaining an adjacency matrix containing the two types of grammatical information: FA=CA+DA.
[0177] The graph attention network adds an attention mechanism to the graph neural network:
[0178] h p =GAT(H,FA:θ);
[0179] Among them, H represents the node features (the features here are obtained by the graph neural network), FA represents the adjacency matrix, and θ represents the parameter set; h p Represents the grammatical structure features of the obtained text.
[0180] The interactive attention network can be the semantic information h output by the self-attention layer b And the grammatical information h output by the graph attention layer p Input into the interactive attention module to get the attention score a for each word b and a p ,Right now Next, combine the two pieces of information: h' b = Dropout(a p h b +h b ),h' p = Dropout(a b h p +h p ). Finally, we get the hidden state vector O = h' b +h' p .
[0181] 5. Classifier: After the self-attention layer, the final hidden state vector is sent to the sofmax layer to predict the label of each word. That is, it is used for two different classification tasks: error detection and error correction. Among them:
[0182] Error detection classifier: It predicts a binary label sequence S = s1, s2, ..., s N , where each s i Indicates whether the corresponding word has a grammatical error, and the loss function is:
[0183] Error Correction Classifier: It predicts the sequence of edit action labels T = t1, t2, …, t N , where each label is the specific editing operation that should be performed on the corresponding word, and the loss function is:
[0184] Among them, p(s i |X e ,θ) represents the embedding X given the input e The probability of correctly predicting the error state of the i-th word under the conditions of and model parameters θ; p(t i |X e ,θ) represents the probability of correctly predicting the editing action of the i-th word under the same conditions, where N = 12. The goal is to minimize the sum of these two loss functions L = L d +L e , while optimizing the model’s ability to detect grammatical errors and correct them.
[0185] Therefore, the classification label prediction results for the sentence "The patient complained of chest pain, but no chest tightness or shortness of breath" was mistakenly written as "The patient complained of chest pain, but no chest shortness of breath" are shown in Table 1:
[0186] Table 1 Classification label prediction results
[0187] Error Detection Layer 1 1 1 0 1 1 1 1 0 1 1 1 Error Correction Layer 1 1 1 2 1 1 1 1 4 (Boring) 1 1 1
[0188] 6. Output layer: The model predicts a set of edit sequences for each word in the Chinese sentence.
[0189] Optionally, after the output layer, model training can also include:
[0190] 7. Post-processing stage: Perform corresponding operations on the text according to the predicted labels, apply the editing operations generated by the model to the input erroneous text sequence, and obtain the corrected text sequence. That is, apply the edit sequence predicted by the model to the input sentence to obtain the output sentence, and then re-input it into the model as input.
[0191] 8. Iterative training: Repeat this process until the maximum number of iterations is reached or the sentence output by the model is consistent with the input sentence. During the training process, you can use strategies such as early stopping to avoid overfitting. Because when making predictions, sometimes you need to perform multiple corrections on an edit position and call the correction algorithm multiple times so that the algorithm iterates until there are no errors. However, in fact, an upper limit on the number of iterations is set. Basically, after two rounds of selection, the model effect is basically not improved. There may even be a situation where the model effect decreases due to too many iterations. Therefore, the upper limit of the number of iterations is set to 4.
[0192] Optionally, use the PyTorch framework for model training so that the model can run directly on mobile devices and deploy the trained model to mobile applications.
[0193] S304: Model evaluation.
[0194] Based on step S303, the trained model is evaluated. The evaluation indicators include precision, recall and F 0.5 value.
[0195] Furthermore, the accuracy satisfies:
[0196]
[0197] The recall rate meets:
[0198]
[0199] F 0.5 The value satisfies:
[0200]
[0201] Among them, {s1,s2,…,s i} is the correction set proposed by the system model; {g1,g2,...,g i} is the manually annotated correction set; |s i ∩g i | represents the number of matches between the model's correction set for sentence i and the manually annotated correction set.
[0202] Therefore, based on the above-mentioned document editing control model training, the mobile terminal document editor control method can be:
[0203] ① Text input: Users input medical record text on mobile devices through touch screen, keyboard, voice key, and image scanning (paper version can be converted into electronic version through scanning).
[0204] ② Typo detection: Use the trained attention-based sequence-to-edit generation model to perform real-time supervision on the input medical record text and automatically detect typos.
[0205] ③ Correction of wrong characters: Based on the wrong characters detected, the system will provide correction suggestions, and the user can choose to accept or reject them. If the user accepts, the system will automatically replace the wrong characters with the correct words.
[0206] ④ Medical record saving: After the user confirms that it is correct, the edited medical record text will be saved to the database.
[0207] ⑤ Data privacy: When processing electronic medical record data, relevant laws and regulations should be strictly followed to ensure that user information will not be leaked and that medical record storage and sharing are secure. When entering the editing page, multiple verification methods such as username and password, face recognition, fingerprint verification or voice lock are required, and the information of authorized operators should be recorded when filling in, modifying or deleting the patient's electronic medical record content, so that responsibilities are assigned to individuals.
[0208] ⑥Model update: With the development of medical informatization and the emergence of new typos, models and data sets need to be updated regularly to maintain the accuracy and effectiveness of the models.
[0209] It should be noted that this embodiment is not limited to the execution steps of the above steps. The specific steps can be adjusted according to the actual medical document editing and storage requirements. For example, data privacy can be verified by operator identity verification before text input, or at regular intervals during the editing process, or when the document is saved.
[0210] The mobile document editor control method provided in the embodiment of the present application can detect and correct typos in real time during the editing process, thereby solving the problem that when editing medical records on mobile devices, doctors or other medical workers are prone to entering typos due to limitations such as screen size and input methods, and use artificial intelligence for automatic review to improve the quality of medical records.
[0211] Figure 5 The structural diagram of the mobile terminal document editor control device provided in this application is applied to the server in the mobile terminal document editor control system, such as Figure 5 As shown, the mobile document editor control device 40 provided in this embodiment includes:
[0212] An acquisition module 401 is used to acquire an electronic medical document, where the electronic medical document is a document obtained by a user through text editing in a mobile document editor control system;
[0213] An input module 402 is used to input the electronic medical document into the document editing control model to obtain text error operation information corresponding to the electronic medical document. The document editing control model is a model trained based on a language model constructed based on a preset encoder and a self-attention layer to identify text errors in the electronic medical document and generate corresponding operation information. The text error operation information includes error detection information and error correction information.
[0214] The receiving module 403 is used to receive a user correction instruction sent by the user according to the text error operation information;
[0215] The processing module 404 is used to perform a correction operation on the electronic medical document according to the text error operation information to obtain a target electronic medical document if the user correction instruction indicates to correct the text of the electronic medical document.
[0216] In a possible implementation, the input module 402 includes a recognition module and a classification module. The input module 402 can be specifically used for:
[0217] Input the electronic medical document into the recognition module to extract the hidden state vector of each word in the electronic medical document;
[0218] The hidden state vector of each word in the electronic medical document is input into the classification module. In the classification module, the error detection information of the electronic medical document is determined, and error correction information corresponding to the error detection information is generated, wherein the text error operation information includes error detection information and error correction information.
[0219] In a possible implementation, the recognition module includes an input layer, an embedding layer, a preset encoder and a self-attention layer, the output of the input layer is used as the input of the connected embedding layer, the output of the embedding layer is used as the input of the connected preset encoder, and the output of the preset encoder is used as the input of the connected self-attention layer; the classification module includes a classifier and an output layer; wherein:
[0220] The input layer is used to receive electronic medical documents;
[0221] The embedding layer is used to convert each word of the electronic medical document in the input layer into a word vector;
[0222] A preset encoder is used to semantically encode the word vector corresponding to each word in the embedding layer based on a multi-layer Transformer encoder to obtain a hidden state vector for each word;
[0223] The self-attention layer is used to perform self-attention processing on the hidden state vector of each word in the preset encoder to obtain the target hidden state vector of each word;
[0224] A classifier is used to perform error detection classification on the target hidden state vector of each word in the self-attention layer to obtain a detection label sequence corresponding to the electronic medical document, and to perform error correction classification on the target hidden state vector of each word to obtain a correction label sequence corresponding to the electronic medical document;
[0225] The output layer is used to generate text error operation information corresponding to the electronic medical document based on the detection label sequence and correction label sequence output by the connected classifier.
[0226] In a possible implementation, the classifier includes an error detection classifier and an error correction classifier; wherein:
[0227] An error detection classifier is used to receive the target hidden state vector of each word in the self-attention layer, and predict the error state of each word using the maximization layer according to the target hidden state vector of each word to obtain a detection label sequence, where each label in the detection label sequence is used to indicate whether the corresponding word has an error;
[0228] The error correction classifier is used to perform operation prediction on the target hidden state vector of each word to obtain a correction label sequence. Each label in the correction label sequence is used to indicate the specific correction operation of the corresponding word.
[0229] In a possible implementation, the classification module further includes a specific layer, and the output of the output layer serves as the input of the connected specific layer; wherein:
[0230] A specific layer is used to update the text error operation information in the output layer to obtain updated text error operation information. The information update processing is the processing of the text error operation information in the output layer based on the behavioral characteristics of the target object. The target object is the object corresponding to the electronic medical document input into the document editing control model.
[0231] In a possible implementation, the processing module 404 may also be used to:
[0232] Acquire a historical medical document set, where the historical medical document set includes electronic medical documents stored by multiple objects in a target hospital within a preset historical time period;
[0233] Based on the preset text error annotation requirements, the historical medical document set is annotated to obtain a sample set, each sample in the sample set includes: an electronic medical document and an edit tag sequence corresponding to the electronic medical document, the edit tag sequence is used to indicate that the erroneous text in the electronic medical document is corrected to the correct text;
[0234] Based on the sample set, the initial document editing control model is trained until the output information of the initial document editing control model meets the preset training requirements, thereby obtaining the document editing control model.
[0235] In a possible implementation, the processing module 404 may also be used to:
[0236] Based on the preset text error annotation requirements, the historical medical document set is annotated to obtain the text error position corresponding to each electronic medical document in the historical medical document set and the correct text after the error text is corrected;
[0237] For each electronic medical document, filtering the words of each electronic medical document according to a preset vocabulary list to be filtered to obtain a filtered electronic medical document;
[0238] According to the text error position and the correct text corresponding to the filtered electronic medical document, the filtered electronic medical document is processed with the minimum edit distance to obtain the corrected target text, the text error type and the error correction operation of the filtered electronic medical document;
[0239] Determine an editing tag sequence corresponding to the electronic medical document according to a text error position, a correct text, a text error type, an error correction operation, and a target text corresponding to the electronic medical document;
[0240] A sample set is constructed based on all electronic medical documents and the edit tag sequence corresponding to each electronic medical document.
[0241] In a possible implementation, the processing module 404 may also be used to:
[0242] Obtaining a test medical document set, where the test medical document set and the historical medical document set have different preset historical time periods;
[0243] Determine a test sample set according to the test medical document set, each test sample in the test sample set includes: an electronic medical document and an edit tag sequence corresponding to the electronic medical document;
[0244] Input the test sample set into the document editing control model to obtain the text error operation test information corresponding to each test sample;
[0245] Determine the model evaluation index according to the edit label sequence and the corresponding text error operation test information in each test sample. The model evaluation index includes the accuracy and recall rate of the model prediction, and the weight index between the accuracy and the recall rate.
[0246] According to the model evaluation index, the performance information of the document editing control model is determined, and the performance information is used to update the document editing control model.
[0247] In a possible implementation, the acquisition module 401 may also be used to:
[0248] Obtaining identity authentication information of the user, the identity authentication information including at least one of ID credential information, biometric information, and target factor authentication information;
[0249] Based on the identity authentication information, the user authentication result is determined, and the user authentication result is used to determine the user's operational authority to edit and process electronic medical documents in the mobile document editor control system.
[0250] The mobile document editor control device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effects are similar, and this embodiment will not be repeated here.
[0251] Figure 6 This is a schematic diagram of the structure of the electronic device provided in this application. Figure 6 As shown, the electronic device 50 provided in this embodiment includes: at least one processor 501 and a memory 502. Optionally, the device 50 also includes a communication component 503. The processor 501, the memory 502 and the communication component 503 are connected via a bus 504.
[0252] In a specific implementation process, at least one processor 501 executes the computer-executable instructions stored in the memory 502, so that at least one processor 501 executes the above method.
[0253] The specific implementation process of the processor 501 can be found in the above method embodiment, and its implementation principle and technical effect are similar, so this embodiment will not be repeated here.
[0254] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the invention may be directly implemented as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.
[0255] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.
[0256] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of the present application is not limited to only one bus or one type of bus.
[0257] The present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.
[0258] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above method is implemented.
[0259] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special-purpose computer.
[0260] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (Application Specific Integrated Circuits, referred to as: ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.
[0261] The division of units is only a logical function division, and there may be other divisions in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0262] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0263] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0264] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0265] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.
[0266] Finally, it should be noted that those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses or adaptations of the present invention, which follow the general principles of the present invention and include common knowledge or customary technical means in the art not disclosed by the present invention, are not limited to the precise structure described above and shown in the drawings, and may be modified and changed in various ways without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.
Claims
1. A method for controlling a mobile document editor, characterized in that: Applied to a server in a mobile document editor control system, the method comprises: Acquire an electronic medical document, wherein the electronic medical document is a document obtained by a user through text editing in the mobile document editor control system; Inputting the electronic medical document into a document editing control model to obtain text error operation information corresponding to the electronic medical document, wherein the document editing control model is a model trained based on a language model constructed based on a preset encoder and a self-attention layer for identifying text errors in electronic medical documents and generating corresponding operation information, wherein the text error operation information includes error detection information and error correction information; Receiving a user correction instruction sent by the user according to the text error operation information; If the user correction instruction indicates to perform text correction on the electronic medical document, a correction operation is performed on the electronic medical document according to the text error operation information to obtain a target electronic medical document.
2. The method according to claim 1, characterized in that The document editing control model includes a recognition module and a classification module. The electronic medical document is input into the document editing control model to obtain text error operation information corresponding to the electronic medical document, including: Inputting the electronic medical document into the recognition module, extracting the hidden state vector of each word in the electronic medical document; The hidden state vector of each word in the electronic medical document is input into the classification module, in which the error detection information of the electronic medical document is determined, and error correction information corresponding to the error detection information is generated, wherein the text error operation information includes error detection information and error correction information.
3. The method according to claim 2, characterized in that The recognition module includes an input layer, an embedding layer, a preset encoder and a self-attention layer, the output of the input layer is used as the input of the connected embedding layer, the output of the embedding layer is used as the input of the connected preset encoder, and the output of the preset encoder is used as the input of the connected self-attention layer; the classification module includes a classifier and an output layer; wherein: The input layer is used to receive electronic medical documents; The embedding layer is used to convert each word of the electronic medical document in the input layer into a word vector; The preset encoder is used to perform semantic encoding on the word vector corresponding to each word in the embedding layer based on a multi-layer Transformer encoder to obtain a hidden state vector of each word; The self-attention layer is used to perform self-attention processing on the hidden state vector of each word in the preset encoder to obtain a target hidden state vector of each word; The classifier is used to perform error detection classification on the target hidden state vector of each word in the self-attention layer to obtain a detection label sequence corresponding to the electronic medical document, and perform error correction classification on the target hidden state vector of each word to obtain a correction label sequence corresponding to the electronic medical document; The output layer is used to generate text error operation information corresponding to the electronic medical document according to the detection label sequence and the correction label sequence output by the connected classifier.
4. The method according to claim 3, characterized in that The classifier includes an error detection classifier and an error correction classifier; wherein: The error detection classifier is used to receive the target hidden state vector of each word in the self-attention layer, and predict the error state of each word using the maximization layer according to the target hidden state vector of each word to obtain a detection label sequence, wherein each label in the detection label sequence is used to indicate whether the corresponding word has an error; The error correction classifier is used to perform operation prediction on the target hidden state vector of each word to obtain a correction label sequence, and each label in the correction label sequence is used to indicate a specific correction operation of the corresponding word.
5. The method according to claim 3, characterized in that: The classification module further includes a specific layer, and the output of the output layer serves as the input of the connected specific layer; wherein: The specific layer is used to perform information update processing on the text error operation information in the output layer to obtain updated text error operation information. The information update processing is the processing of the text error operation information in the output layer based on the behavioral characteristics of the target object. The target object is the object corresponding to the electronic medical document input into the document editing control model.
6. The method according to any one of claims 1 to 5, characterized in that Before inputting the electronic medical document into the document editing control model to obtain text error operation information corresponding to the electronic medical document, the method further includes: Acquire a historical medical document set, wherein the historical medical document set includes electronic medical documents stored by multiple objects in a target hospital within a preset historical time period; Based on the preset text error marking requirements, the historical medical document set is marked to obtain a sample set, each sample in the sample set includes: an electronic medical document and an edit tag sequence corresponding to the electronic medical document, the edit tag sequence is used to indicate that the erroneous text in the electronic medical document is corrected to the correct text; Based on the sample set, the initial document editing control model is trained until the output information of the initial document editing control model meets the preset training requirements, thereby obtaining the document editing control model.
7. The method according to claim 6, characterized in that The historical medical document set is annotated based on the preset text error annotation requirements to obtain a sample set, including: Based on the preset text error marking requirements, the historical medical document set is marked to obtain the text error position corresponding to each electronic medical document in the historical medical document set and the correct text after the error text is corrected; For each electronic medical document, filtering the words of each electronic medical document according to a preset vocabulary list to be filtered to obtain a filtered electronic medical document; According to the text error position and the correct text corresponding to the filtered electronic medical document, the filtered electronic medical document is subjected to minimum edit distance processing to obtain a corrected target text, as well as a text error type and an error correction operation of the filtered electronic medical document; Determining an editing tag sequence corresponding to the electronic medical document according to a text error position, a correct text, a text error type, an error correction operation, and a target text corresponding to the electronic medical document; A sample set is constructed based on all electronic medical documents and the edit tag sequence corresponding to each electronic medical document.
8. The method according to claim 6, characterized in that After the initial document editing control model is trained based on the sample set until the output information of the initial document editing control model meets the preset training requirements and the document editing control model is obtained, the method further includes: Acquire a test medical document set, wherein the test medical document set and the historical medical document set have a different preset historical time period; Determine a test sample set according to the test medical document set, each test sample in the test sample set includes: an electronic medical document and an edit tag sequence corresponding to the electronic medical document; Inputting the test sample set into the document editing control model to obtain text error operation test information corresponding to each test sample; Determine a model evaluation index according to the edit label sequence and the corresponding text error operation test information in each test sample, wherein the model evaluation index includes the accuracy and recall rate of model prediction, and a weight index between the accuracy and the recall rate; The performance information of the document editing control model is determined according to the model evaluation index, and the performance information is used to update the document editing control model.
9. The method according to claim 1, characterized in that: Before obtaining the electronic medical document, the method further includes: Obtaining identity authentication information of the user, the identity authentication information including at least one of ID credential information, biometric information, and target factor authentication information; A user verification result is determined based on the identity authentication information, and the user verification result is used to determine the user's operational authority to edit and process electronic medical documents in the mobile document editor control system.
10. A mobile document editor control system, characterized in that: include: User terminals, servers, and databases; The user terminal is used to provide an editing interface to the user and generate a corresponding electronic medical document when the user edits the text; The server is used to obtain the electronic medical documents in the user terminal; The server is further used to input the electronic medical document into a document editing control model to obtain text error operation information corresponding to the electronic medical document, wherein the document editing control model is a model obtained by training a language model constructed based on a preset encoder and a self-attention layer for identifying text errors in electronic medical documents and generating corresponding operation information, wherein the text error operation information includes error detection information and error correction information; The user terminal is further used to display the text error operation information to the user; The server is further configured to receive a user correction instruction sent by the user according to the text error operation information; The server is further configured to perform a correction operation on the electronic medical document according to the text error operation information to obtain a target electronic medical document if the user correction instruction indicates to perform text correction on the electronic medical document; The database is used to store the electronic medical document, the text error operation information, the user correction instruction, the target electronic medical document and the document editing control model.
Citation Information
Patent Citations
Method for forming personalized error correcting model and input method system of personalized error correcting
CN101350004A
Text error correction method and device, equipment and storage medium
CN114611494A
Multi-granularity Chinese text error correction method and device
CN116127952A
Chinese spelling error correction method and device based on comparative learning and medium
CN116127953A
Method and apparatus of NER-oriented chinese clinical text data augmentation
US20240013000A1