Electronic medical record text analysis system and method
By designing an electronic medical record text analysis system and using the collaborative work of multiple modules, the problem of difficulty in the existing technology to quickly extract key information in medical record text and discover medical laws is solved, efficient data analysis and potential laws are realized, and the utilization value of medical data is significantly improved.
Patent Information
- Application Number
- CN202510648344.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When existing electronic medical record systems process large-scale text data, it is difficult to quickly extract critical information and discover potential medical rules, resulting in a large amount of valuable data not being fully utilized.
An electronic medical record text analysis system was designed, including a case database, keyword sensitive neural module, similar item search module, statistical mining module and pathological law calculation model. Through the coordinated work of these modules, key information in medical record text can be efficiently extracted and analyzed and potential medical laws can be explored.
It has achieved efficient extraction of key information in medical record text, quickly explored potential medical laws, supported disease distribution statistics, association rule mining and trend analysis, significantly improved the utilization value of medical data, and ensured the accuracy and reliability of analysis results through multi-dimensional authenticity verification.
Smart Images

Figure CN120183735A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic medical records, and specifically to an electronic medical record text analysis system and method. Background Art
[0002] Medical records are records of the medical activities of medical staff in the process of the occurrence, development, and outcome of patients' diseases, including examinations, diagnoses, treatments, etc. They are also the medical health records of patients compiled by summarizing and organizing the collected data in accordance with the specified format and requirements. Medical records are not only the summary of clinical practice work but also the legal basis for exploring disease laws and handling medical disputes, playing an important role in medical treatment, prevention, teaching, scientific research, hospital management, etc.
[0003] With the wide application of the Electronic Medical Record (EMR) system, the amount of medical data has increased explosively. However, most existing EMR systems focus on data storage and management and lack the ability to analyze unstructured text data. When dealing with large-scale text data, existing technologies are difficult to quickly extract key information and discover potential medical laws, resulting in the underutilization of a large amount of valuable data. Therefore, an electronic medical record text analysis system and method are needed. Summary of the Invention
[0004] Aiming at the deficiencies of the existing technology, the present invention provides an electronic medical record text analysis system and method, which solves the problem that existing technologies are difficult to quickly extract key information and discover potential medical laws when dealing with large-scale text data, resulting in the underutilization of a large amount of valuable data.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: An electronic medical record text analysis system, comprising: A case database for storing and statistically analyzing the input case text data; A keyword-sensitive neural module for collecting and statistically analyzing the correlation data of cases in the case database; A similar item retrieval module for using the data obtained by the keyword-sensitive neural module as a standard to retrieve individuals for review and improve the authenticity of the data; A statistical mining module for analyzing basic information such as disease distribution and patient characteristics, mining the associations between different variables, and discovering potential medical laws; A pathological law calculation model for providing computing power for the statistical mining module.
[0006] Preferably, the case database includes: A data import unit for importing case data, identifying and splitting the keywords in the case data, and performing labeling; A data storage unit for storing case text data and corresponding secondary text editing data; A data sorting unit that divides the mother serial number, sub-serial number, and number according to the imported cases for classification, sorting, and quick retrieval; A paradox error clearing unit for manually clearing and correcting the error diagnosis records stored in the data storage unit, only retaining some features, and the same type of items can be quickly selected during the clearing process.
[0007] Preferably, the keyword sensitive nerve module includes: A keyword general statistics table for manually importing sensitive keywords for retrieving and collecting data; A sensitive reticular nerve unit for browsing a large amount of data in the case database according to sensitive keywords, triggering nerve stimulation according to sensitive keywords, and accumulating sensitive word coefficients; A threshold triggering unit, when the sensitive word coefficients triggered and accumulated by the sensitive reticular nerve unit exceed the threshold, the entered collective disease law characteristics are correspondingly triggered to collect case correlation data.
[0008] Preferably, the similar item retrieval module includes: An individual retrieval unit for analyzing and screening text individuals that meet the correlation data; A special individual discrimination unit for excluding data irrelevant to the correlation data and transferring it to manual review; A verification and comparison unit for randomly selecting text individuals and sequentially judging the authenticity of the correlation data through intelligent automatic judgment and manual judgment.
[0009] Preferably, the statistical mining module includes: A descriptive statistics unit for statistically analyzing basic information such as disease distribution and patient characteristics; An association rule mining unit for mining the associations between different variables and discovering potential medical laws; A trend analysis unit for analyzing disease trends and patient health change trends.
[0010] Preferably, the pathological law calculation model is based on any one of deep learning frameworks such as BERT and LSTM as the training framework, and processes and learns the data obtained from the keyword sensitive nerve module.
[0011] An electronic medical record text analysis method includes the following steps: Step 1: Import case data through the data import unit and identify and split keywords for annotation; Step 2: Collect and statistically analyze the correlation data in the case database through the keyword sensitive nerve module, accumulate sensitive word coefficients, and obtain the correlation data; Step 3: Retrieve text individuals that match the relevant data through the similarity item retrieval module and verify the data authenticity; Step 4: Conduct statistical analysis of basic information such as disease distribution and patient characteristics through the statistical mining module.
[0012] Preferably, the specific steps for accumulating the sensitive word coefficient in Step 2 are as follows: S1: Initialize the sensitive word coefficient to zero; S2: Scan the text data in the case database through the sensitive reticular nerve unit; S3: Whenever the sensitive keyword is triggered, increase the sensitive word coefficient by a preset value; S4: When the sensitive word coefficient exceeds the preset threshold, trigger the threshold trigger unit to record the relevant data.
[0013] Preferably, the specific steps for authenticity verification in Step 3 are as follows: S1: Screen out text individuals that match the relevant data through the individual retrieval unit; S2: Exclude data irrelevant to the relevant data through the special individual discrimination unit; S3: Randomly select text individuals through the verification and comparison unit for intelligent automatic judgment and manual judgment; S4: Modify the relevant data according to the judgment results to ensure the authenticity of the data.
[0014] Preferably, when the keyword retrieval fails in S3, the range will be gradually narrowed down according to the mother serial number and sub - serial number of the case in sequence, identify the case record, sort out the causal relationship and conduct pairing judgment again.
[0015] The present invention provides an electronic medical record text analysis system and method. It has the following beneficial effects: 1. The present invention can efficiently extract key information in the medical record text, such as disease names, symptoms and treatment plans, and quickly mine potential medical laws, support disease distribution statistics, association rule mining and trend analysis, provide a scientific basis for clinical decision - making, reduce manual intervention in the intelligent analysis process, improve data processing efficiency, reduce operation difficulty, and significantly enhance the utilization value of medical data.
[0016] 2. The present invention can conduct multi - dimensional authenticity verification on the extracted information to ensure the accuracy and reliability of the analysis results. The system can identify and exclude irrelevant data, reduce misjudgment and omission, and in the analysis of adverse events and complaint information in the medical record, can timely discover potential medical quality problems, providing data support for the hospital to optimize processes and improve service quality. Description of the Drawings
[0017] Figure 1Schematic diagram of an electronic medical record text analysis system of the present invention; Figure 2 Schematic diagram of the case database system of an electronic medical record text analysis system of the present invention; Figure 3 Schematic diagram of the keyword sensitive neural module system of an electronic medical record text analysis system of the present invention; Figure 4 Schematic diagram of the similar item retrieval module system of an electronic medical record text analysis system of the present invention; Figure 5 Schematic diagram of the statistical mining module system of an electronic medical record text analysis system of the present invention. Detailed implementation manners
[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0019] Embodiment: As one aspect of the present application, please refer to the attached Figure 1 - attached Figure 5 , the embodiment of the present invention provides an electronic medical record text analysis system, including: A case database for storing and statistically analyzing the input case text data, including: A data import unit for importing case data, identifying and splitting keywords in the case data, and performing annotation; A data storage unit for storing case text data and corresponding secondary text editing data; A data sorting unit for dividing the imported cases into mother sequence numbers, sub-sequence numbers, and numbers for classification, sorting, and quick retrieval; A paradox error clearing unit for manually clearing and correcting the error diagnosis records stored in the data storage unit, only retaining some features, and the same type of items can be quickly selected during the clearing process.
[0020] A keyword sensitive neural module for collecting and statistically analyzing the correlation data of cases in the case database, including: A keyword general statistics table for manually importing sensitive keywords for retrieval and data collection; A sensitive reticular neural unit for browsing a large amount of data in the case database according to the sensitive keywords, triggering nerve stimulation according to the sensitive keywords, and accumulating sensitive word coefficients; A threshold trigger unit. When the sensitive word coefficient accumulated by the sensitive reticular nerve unit exceeds the threshold, it correspondingly triggers the entered collective disease pattern features to collect case correlation data.
[0021] Among them, the collective disease pattern features are the disease-related rules and features mined from analyzing a large amount of case data, such as the associations and distribution rules between diseases and symptoms, treatment plans, patient characteristics, etc., as well as the high incidence of specific diseases in specific populations, the frequent occurrence of specific symptom combinations, and the effective events of a certain treatment plan in specific diseases.
[0022] A similar item retrieval module, which is used to take the data obtained by the keyword sensitive nerve module as a standard to retrieve individuals for review and improve the authenticity of the data, including: An individual retrieval unit, which is used to analyze and screen text individuals that meet the correlation data. The texts are all marked with classification labels of different pathological features for retrieval and classification. Through the retrieval and screening of the corresponding classification labels, the corresponding text individual data associated with the classification labels can be obtained; A special individual discrimination unit, which is used to exclude data irrelevant to the correlation data and transfer it to manual review; A verification and comparison unit, which is used to randomly select text individuals and sequentially verify the authenticity of the correlation data through intelligent automatic judgment and manual judgment. The intelligent automatic judgment is based on the system to perform data matching of similarity and convergence to verify whether the similarity result of the data is higher than the minimum similarity value to confirm the true reliability of the data.
[0023] A statistical mining module, which is used to analyze basic information such as disease distribution and patient characteristics, mine the associations between different variables, and discover potential medical rules, including: A descriptive statistics unit, which is used to statistically analyze basic information such as disease distribution and patient characteristics; An association rule mining unit, which is used to mine the associations between different variables and discover potential medical rules; A trend analysis unit, which is used to analyze the disease trend and the trend of patient health changes; Among them, the relevant data statistically analyzed by the descriptive statistics unit will be processed by data visualization and converted into visual digital data. According to various data characteristics, including the feature weight data points of disease states, disease duration, disease development trends, etc., and they will be planned in statistical charts such as line charts and bar charts to summarize the characteristics and analyze the trend of patient health changes; A pathological rule calculation model, which is used to provide a computing power basis for the statistical mining module. The pathological rule calculation model is based on any one of deep learning frameworks such as BERT and LSTM as the training framework, and will process and learn the data obtained by the keyword sensitive nerve module.
[0024] Among them, the potential medical laws are the internal connections and laws discovered by the statistical mining module through the analysis and mining of a large amount of medical record data, such as the associations between diseases and symptoms, treatment plans, patient characteristics, etc., including the associations between diseases and symptoms, disease distributions, treatment plans and disease prognoses, disease development trends, the associations between patient characteristics and diseases, etc., providing scientific basis and data support for clinical decision-making, hospital management, teaching and research, etc.
[0025] Based on the electronic medical record text analysis system provided above, as another aspect of this application, an electronic medical record text analysis method includes the following steps: Step 1: Import case data through the data import unit, and identify and split keywords for annotation; Step 2: Collect and statistically analyze the relevant data in the case database through the keyword sensitive neural module to accumulate the sensitive word coefficients and obtain the relevant data. Among them, the specific steps for accumulating the sensitive word coefficients are: S1: Initialize the sensitive word coefficient to zero; S2: Scan the text data in the case database through the sensitive reticular neural unit; S3: Whenever a sensitive keyword is triggered, increase the sensitive word coefficient by a preset value; S4: When the sensitive word coefficient exceeds the preset threshold, trigger the threshold trigger unit to record the relevant data.
[0026] When the keyword retrieval fails in S3, the range will be gradually narrowed down according to the mother serial number and child serial number of the case in turn, the case record will be identified, and the causal relationship will be sorted out and re-paired for determination.
[0027] Step 3: Retrieve the text individuals that meet the relevant data through the similarity item retrieval module and verify the data authenticity. Among them, the specific steps for authenticity verification are: S1: Screen out the text individuals that meet the relevant data through the individual retrieval unit; S2: Exclude the data irrelevant to the relevant data through the special individual discrimination unit; S3: Randomly select text individuals through the verification and comparison unit for intelligent automatic judgment and manual judgment; S4: Modify the relevant data according to the judgment result to ensure the data authenticity. If the judgment passes, no modification is required. If the judgment fails, repeat Step 2 with the generation record of this relevant data to re-output the relevant data and make a comparison to remove the anomalies caused by the accumulation of sensitive word coefficients in Step 2. If the comparison fails, manually modify the data according to the generation record.
[0028] Step 4: Conduct statistical analysis of basic information such as disease distribution and patient characteristics through the statistical mining module.
[0029] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An electronic medical record text analysis system, characterized in that: include: Case database, used to store and compile statistics of entered case text data; Keyword-sensitive neural module, used to collect correlation data of cases in the statistical case database; The similar item retrieval module is used to use the data obtained by the keyword-sensitive neural module as the standard, retrieve individuals for review, and improve the authenticity of the data; Statistical mining module, which is used to analyze basic information such as disease distribution and patient characteristics, explore the associations between different variables, and discover potential medical laws; Pathological law calculation model, used to provide computing power basis for statistical mining module; The keyword-sensitive neural module includes: Keyword statistics table, used to manually import sensitive keywords for searching and collecting data; The sensitive reticular neural unit is used to browse a large amount of data in the case database according to sensitive keywords, trigger neural stimulation according to sensitive keywords, and accumulate sensitive word coefficients; Threshold trigger unit, when the accumulated sensitive word coefficient triggered by the sensitive reticular neural unit exceeds the threshold, it corresponds to the triggering of the recorded collective disease law characteristics to collect case correlation data.
2. The electronic medical record text analysis system according to claim 1, characterized in that: The case database includes: A data import unit is used to import case data, identify and split the keywords in the case data and mark them; A data storage unit, used for storing case text data and corresponding secondary text editing data; The data sorting unit divides the imported cases into parent numbers, sub-sequence numbers and serial numbers for classification, sorting and quick retrieval; The paradox error clearing unit is used to manually clear and correct the error diagnosis records stored in the data storage unit, retaining only some features, and the clearing process can quickly select similar items.
3. The electronic medical record text analysis system according to claim 1, characterized in that: The similar item retrieval module comprises: Individual retrieval unit, used to analyze and filter text individuals that meet the relevant data; Special individual distinguishing unit, used to exclude data irrelevant to the relevant data and transfer it to manual review; The verification and comparison unit is used to randomly extract text individuals and sequentially conduct intelligent automatic judgment and manual judgment on the authenticity of the associated data.
4. The electronic medical record text analysis system according to claim 1, characterized in that: The statistical mining module includes: Descriptive statistical units are used to collect basic information such as disease distribution and patient characteristics; Association rule mining unit, used to mine the associations between different variables and discover potential medical laws; Trend analysis unit, used to analyze disease trends and patient health change trends.
5. The electronic medical record text analysis system according to claim 1, characterized in that: The pathological law calculation model is based on any one of the deep learning frameworks such as BERT, LSTM, etc. as a training framework, and will process and learn the data obtained by the keyword sensitive neural module.
6. A method for analyzing electronic medical record text, using an electronic medical record text analysis system as claimed in any one of claims 1 to 5, characterized in that: The following steps are involved: Step 1: Import case data through the data import unit, and identify split keywords for annotation; Step 2: Collect and count the correlation data in the case database through the keyword sensitive neural module, accumulate sensitive word coefficients, and obtain correlation data; Step 3: Retrieve text individuals that match the relevant data through a similar item retrieval module and verify the authenticity of the data; Step 4: Perform statistical analysis of basic information such as disease distribution and patient characteristics through the statistical mining module.
7. The electronic medical record text analysis method according to claim 6, characterized in that: The specific steps of accumulating sensitive word coefficients in step 2 are: S1: Initialize the sensitive word coefficient to zero; S2: Scanning text data in the case database through sensitive reticular neural units; S3: Whenever a sensitive keyword is triggered, the sensitive word coefficient increases by a preset value; S4: When the sensitive word coefficient exceeds a preset threshold, the threshold trigger unit is triggered to record the correlation data.
8. The electronic medical record text analysis method according to claim 6, characterized in that: The specific steps of authenticity verification in step 3 are: S1. Filter out text individuals that meet the relevant data through individual retrieval units; S2, exclude data irrelevant to the correlation data through special individual distinguishing units; S3, randomly extracting text individuals through the verification and comparison unit, and performing intelligent automatic judgment and manual judgment; S4. Modify the associated data based on the judgment results to ensure the authenticity of the data.
9. The electronic medical record text analysis method according to claim 7, characterized in that: When the keyword search in S3 fails, the scope will be gradually narrowed down according to the parent serial number and child serial number of the case, the case record will be identified, the causal relationship will be sorted out, and the matching determination will be re-performed.
Citation Information
Patent Citations
Chinese electronic case text analysis method and system
CN108831559A
An ES-based electronic medical record retrieval method
CN109299239A
Epidemiological data integration system and method
CN113362962A