Medical examination report analysis system and method based on big data

The medical laboratory report system, which integrates big data and performs intelligent analysis, solves the problems of low efficiency and poor accuracy in existing technologies. It achieves efficient and accurate analysis of laboratory reports, reduces the risk of missed diagnoses and misdiagnoses, and is suitable for the integration of multi-dimensional laboratory data and clinical auxiliary diagnosis.

CN121747818APending Publication Date: 2026-03-27ZHENXIONG COUNTY PEOPLES HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Current medical test report analysis suffers from problems such as low efficiency, poor accuracy, insufficient correlation analysis, lack of unified standards, difficulty in data exchange, and poor adaptability to individual differences, leading to high risks of diagnostic delays, missed diagnoses, and misdiagnoses.

Method used

The system employs a big data-based medical laboratory report analysis system. Through modules for data collection, preprocessing, storage, indicator correlation analysis, anomaly warning, and user interaction, it integrates multi-source data, utilizes association rule algorithms and machine learning models for automated analysis, generates standardized reports, and provides a visual interactive interface.

Benefits of technology

It improves the efficiency and accuracy of laboratory report analysis, identifies abnormal indicator combinations that are difficult to detect manually, reduces the risk of missed or misdiagnosed diagnoses, supports personalized risk assessment, and adapts to clinical diagnostic needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121747818A_ABST
    Figure CN121747818A_ABST
Patent Text Reader

Abstract

The invention discloses a medical examination report analysis system and method based on big data, and relates to the technical field of medical data processing, the system comprises a data acquisition module, a preprocessing module, a storage module, an index correlation analysis module, an abnormity early warning module, a result generation module and a user interaction module, and the method comprises corresponding data processing and analysis steps. Multi-source data such as inspection reports, historical diagnosis and treatment, clinical cases and the like are collected through a standardized interface, are subjected to cleaning, standardization and structural conversion, and are accessed by means of a distributed storage architecture; an Apriori association rule algorithm is adopted to mine index association, a random forest and a neural network model are combined to realize risk assessment, and a three-level grading early warning mechanism is matched to push accurate suggestions. According to the method, traditional manual interpretation is replaced, the analysis efficiency and accuracy are improved, the problems of insufficient associated information mining, non-uniform diagnosis standards and the like are solved, medical institution scenes at all levels are adapted, interactive optimization is supported, clinical precise diagnosis is assisted, the burden of doctors is relieved, and the method has important popularization value for intelligent application of medical data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical data processing technology, specifically to a medical test report analysis system and method based on big data, which is particularly suitable for the integrated analysis of multi-dimensional test data, early warning of abnormal indicators, and clinical auxiliary diagnosis scenarios. Background Technology

[0002] With the development of medical testing technology, the number of clinical testing items is increasing. A complete medical test report often contains dozens or even hundreds of indicators and data, covering multiple fields such as hematology, biochemistry, immunology, and microbiology.

[0003] Currently, the analysis of medical test reports mainly relies on manual interpretation by physicians, which has the following drawbacks:

[0004] Firstly, manual analysis is inefficient and cannot meet the demand for rapid processing of massive amounts of test data. This is especially true in high-traffic settings such as tertiary hospitals. The laboratory department of a certain provincial tertiary hospital receives more than 5,000 outpatient and inpatient applications per day. If manual entry is relied upon, even a team of 30 people would find it difficult to cope, which could easily lead to diagnostic delays.

[0005] Secondly, the correlation between different test indicators is complex, and manual interpretation is difficult to fully uncover the potential correlation between indicators, which may lead to the omission of key diagnostic information.

[0006] Third, the lack of unified analytical standards and the differences in interpretation experience among different doctors can easily lead to diagnostic errors and affect treatment outcomes.

[0007] Fourth, existing analytical tools are mostly limited to threshold judgment of a single indicator and cannot be combined with multi-source information such as patient historical data, clinical case data, and epidemiological data for comprehensive analysis, resulting in insufficient accuracy and comprehensiveness of the analysis results.

[0008] Furthermore, the current medical testing field also faces prominent challenges in data interoperability and standardization: the laboratory information systems (LIS) of different medical institutions are significantly heterogeneous, with some using the HL7 v2.x protocol and others using the HL7 FHIR protocol. The lack of a unified data interface standard leads to the need for a large amount of additional manpower for format adaptation when integrating data across institutions, which further reduces analysis efficiency.

[0009] Primary healthcare institutions face challenges such as insufficient testing services and non-standard operating procedures, resulting in low acceptance rates for their reports. Patients often require repeated testing during referrals, increasing both financial burden and wasting medical resources. Furthermore, existing analytical systems lack dynamic updates, failing to synchronize medical indicator thresholds with the latest clinical guidelines and failing to adequately consider individual differences such as patient age, disease duration, and underlying conditions, hindering personalized risk assessment. In addition, manual interpretation is susceptible to fatigue and experience levels, and younger physicians lack sensitivity to rare indicator combinations, further exacerbating the risk of missed diagnoses.

[0010] Therefore, there is an urgent need for a technical solution that can integrate multi-source big data, automate the analysis and inspection of reports, mine the correlation of indicators, and provide accurate analysis results to solve the problems of low efficiency, poor accuracy, and insufficient correlation analysis in existing technologies. Summary of the Invention

[0011] The purpose of this invention is to provide a medical test report analysis system and method based on big data. By integrating multi-source data through big data processing technology, it can achieve automated and precise analysis of medical test reports, improve analysis efficiency and accuracy, and provide reliable support for clinical diagnosis.

[0012] To achieve the above objectives, the present invention provides the following technical solution:

[0013] A big data-based medical test report analysis system includes a data acquisition module, a data preprocessing module, a big data storage module, an indicator correlation analysis module, an anomaly warning module, a result generation module, and a user interaction module connected in sequence.

[0014] The data acquisition module is used to collect multi-source data, including raw data from medical test reports, patient history of medical treatment data, clinical case database data, medical indicator standard threshold data, and epidemiological statistics.

[0015] The data preprocessing module is used to clean, standardize, and convert formats of multi-source data.

[0016] The big data storage module adopts a distributed storage architecture, including a relational database for storing structured data and a non-relational database for storing unstructured data and massive historical data;

[0017] The indicator correlation analysis module is used to perform single indicator anomaly detection, multi-indicator correlation mining, and comprehensive risk assessment based on big data algorithms.

[0018] The anomaly warning module is used to provide graded warnings and generate warning information based on the analysis results;

[0019] The result generation module is used to integrate analysis results and early warning information, generate standardized analysis reports, and support export in multiple formats.

[0020] The user interaction module provides a visual operation interface, supporting data uploading, result viewing, parameter adjustment, report export, and feedback input.

[0021] Furthermore, the raw data of the medical test report collected by the data acquisition module includes the index values, test time, and test equipment model information of blood test, biochemical test, immunological test, and microbiological test items.

[0022] Furthermore, the data preprocessing module uses the Raida criterion to remove outlier data, performs data normalization based on the Z-score standardization formula, and employs OCR recognition technology to convert unstructured data into structured data.

[0023] Furthermore, the big data algorithms used in the indicator correlation analysis module include association rule algorithms and machine learning models; the association rule algorithm is the Apriori algorithm, which is used to mine the correlation between multiple indicators; the machine learning models include random forest algorithms and neural network algorithms, which are used for comprehensive risk assessment.

[0024] Furthermore, the anomaly warning module has three warning levels: Level 1, Level 2, and Level 3. The warning information includes the name of the anomaly indicator, the degree of anomaly, related indicators, risk assessment results, and recommended measures.

[0025] A method for analyzing medical test reports based on big data, using the system described in any one of claims 1-5, includes the following steps:

[0026] S1: Data collection, collecting multi-source data, including raw data from medical test reports, patient history of medical treatment data, clinical case database data, medical indicator standard threshold data, and epidemiological statistics data;

[0027] S2: Data preprocessing, which involves cleaning, standardizing, and format conversion of the collected multi-source data;

[0028] S3: Data storage, storing pre-processed structured data in a relational database, and unstructured data and massive historical data in a non-relational database;

[0029] S4: Indicator correlation analysis, based on big data algorithms to identify single indicator anomalies, mine multi-indicator correlations, and conduct comprehensive risk assessment;

[0030] S5: Anomaly warning, which provides graded warnings and generates warning information based on the analysis results;

[0031] S6: Results generation, integrates analysis results and early warning information, generates standardized analysis reports and supports export in multiple formats;

[0032] S7: Interactive feedback, which allows users to view, export reports, and enter feedback through a visual interface. The system optimizes algorithms and analysis standards based on the feedback.

[0033] Furthermore, in step S2, data cleaning includes removing missing values, outliers, and duplicate data; standardization processing includes converting indicator data from different units into a unified standard unit and performing normalization processing; and format conversion includes converting unstructured data into structured data.

[0034] Furthermore, in step S4, single-indicator anomaly judgment is achieved by comparing the indicator value with standard threshold data; multi-indicator correlation mining uses the Apriori algorithm; comprehensive risk assessment combines patients' historical treatment data, clinical case data, and epidemiological data, and is implemented using the random forest algorithm or neural network algorithm.

[0035] Furthermore, in step S5, the first-level warning corresponds to a single indicator with severe abnormality, a combination of high-risk indicators, or a high risk of disease; the second-level warning corresponds to an important indicator with abnormality or a medium-risk situation; and the third-level warning corresponds to a general indicator with abnormality or a low-risk situation.

[0036] Furthermore, in step S6, the standardized analysis report includes a summary table of test indicators, details of abnormal indicators, conclusions of indicator correlation analysis, disease risk assessment results, early warning information, and clinical diagnostic suggestions; the supported export formats include PDF, Word, and Excel.

[0037] Compared with the prior art, the present invention has the following advantages:

[0038] 1. By integrating multi-source big data, it breaks through the limitations of traditional single-indicator analysis and combines multi-dimensional information such as patient historical data and clinical case data, the analysis results are more comprehensive and accurate;

[0039] 2. By adopting automated data processing and analysis algorithms, manual interpretation is replaced, which greatly improves the efficiency of test report analysis, reduces the workload of physicians, and avoids diagnostic delays;

[0040] 3. By mining the potential relationships between indicators through association rule algorithms, it is possible to identify abnormal indicator combinations that are difficult to detect manually, providing new reference for clinical diagnosis;

[0041] 4. Establish a tiered early warning mechanism to promptly remind physicians to pay attention to high-risk situations and reduce the risk of missed or misdiagnosed cases;

[0042] 5. It adopts a distributed storage architecture, which supports the storage of massive amounts of data and fast read and write, meeting the storage needs of big data analysis;

[0043] 6. Provides a visual interactive interface and multi-format report export function, which is easy to operate and adaptable to actual clinical application scenarios. Attached Figure Description

[0044] Figure 1 This is a system framework diagram of the big data-based medical test report analysis system of the present invention;

[0045] Figure 2 This is a flowchart illustrating the implementation of the big data-based medical test report analysis method of this invention. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0047] This invention proposes a medical test report analysis system and method based on big data, characterized in that:

[0048] I. Big Data-Based Medical Laboratory Report Analysis System

[0049] The system includes a data acquisition module, a data preprocessing module, a big data storage module, an indicator correlation analysis module, an anomaly early warning module, a result generation module, and a user interaction module. These modules are connected sequentially and work together to achieve full-process analysis of test reports.

[0050] Data acquisition module:

[0051] It is used to collect multi-source data, including the original data of medical test reports to be analyzed, patients' historical medical data, clinical case database data, medical indicator standard threshold data, and epidemiological statistics. Among them, the original data of medical test reports covers the indicator values, test time, and test equipment model of multiple test items such as blood tests, biochemical tests, immune tests, and microbiological tests.

[0052] Data preprocessing module:

[0053] The collected multi-source data underwent cleaning, standardization, and format conversion. Data cleaning included removing missing values, outliers, and duplicate data, and using the Laida criterion to identify and remove outliers exceeding three times the standard deviation. Standardization involved converting indicator data from different testing equipment and units into a unified standard unit, and normalizing the indicator data based on the Z-score standardization formula. Format conversion involved converting unstructured data (such as scanned copies of paper reports and text description data) into structured data, and using OCR recognition technology to extract indicator values ​​and related information from the scanned documents.

[0054] Big data storage module:

[0055] It adopts a distributed storage architecture, including relational databases and non-relational databases; the relational database is used to store structured data, such as standardized test index data, patient basic information, standard threshold data, etc.; the non-relational database is used to store unstructured data and massive historical data, such as scanned copies of original test reports, clinical case text data, etc., supporting fast data reading and writing and massive storage.

[0056] Indicator Correlation Analysis Module:

[0057] The analysis of preprocessed data based on big data algorithms includes single-indicator anomaly detection, multi-indicator correlation mining, and comprehensive risk assessment. Single-indicator anomaly detection determines whether a single indicator is abnormal by comparing its value with standard threshold data. Multi-indicator correlation mining uses association rule algorithms (such as the Apriori algorithm) to mine the correlation between different test indicators and identify abnormal indicator combinations. Comprehensive risk assessment combines patients' historical medical data, clinical case data, and epidemiological data, and uses machine learning models (such as random forest algorithms and neural network algorithms) to quantitatively assess patients' disease risk.

[0058] Anomaly warning module:

[0059] Based on the analysis results of the indicator correlation analysis module, a graded early warning is issued for severe abnormalities of single indicators, high-risk indicator combinations, and high disease risk. The early warning levels are divided into Level 1 (urgent), Level 2 (important), and Level 3 (general). Early warning information is generated, including the name of the abnormal indicator, the degree of abnormality, the associated indicators, the risk assessment results, and the recommended measures.

[0060] Result generation module:

[0061] The system integrates analysis results and early warning information to generate standardized analysis reports. These reports include a summary table of test indicators, details of abnormal indicators, conclusions of indicator correlation analysis, disease risk assessment results, early warning information, and clinical diagnostic recommendations. The system supports exporting reports in multiple formats, including PDF, Word, and Excel.

[0062] User interaction module: Provides a visual operation interface, supporting users to upload test report data, view analysis results, adjust analysis parameters, and export analysis reports. It also supports the input of user feedback to optimize system algorithms and analysis standards.

[0063] II. Big Data-Based Medical Laboratory Report Analysis Methods

[0064] This method applies the above system and includes the following steps:

[0065] Data collection steps:

[0066] The data acquisition module collects multi-source data, including raw data from medical test reports to be analyzed, patient history of medical treatment data, clinical case database data, medical indicator standard threshold data, and epidemiological statistics.

[0067] Data preprocessing steps:

[0068] The data preprocessing module cleans the collected multi-source data, removing missing values, outliers, and duplicate data; it standardizes indicator data of different formats and units, converting them into a unified standard; and it converts unstructured data into structured data.

[0069] Data storage steps:

[0070] Preprocessed structured data is stored in a relational database, while unstructured data and massive historical data are stored in a non-relational database, enabling data to be categorized, stored, and retrieved quickly.

[0071] Steps for correlation analysis of indicators:

[0072] The indicator correlation analysis module first compares the value of a single indicator with the standard threshold to determine whether the single indicator is abnormal; then it uses an association rule algorithm to mine the correlation between multiple indicators and identify abnormal indicator combinations; finally, it combines multi-source data and uses a machine learning model to conduct a comprehensive risk assessment.

[0073] Abnormal warning steps:

[0074] The anomaly warning module provides tiered warnings based on the analysis results, generating warning information at the corresponding level.

[0075] Result generation steps:

[0076] The results generation module integrates analysis results and early warning information to generate standardized analysis reports, which support export in multiple formats.

[0077] Interactive feedback steps:

[0078] Users can view and export analysis reports and input feedback through the user interaction module. The system optimizes algorithms and analysis standards based on the feedback. Specific implementation examples:

[0080] The present invention will be further described in detail below with reference to the embodiments:

[0081] Example 1

[0082] This embodiment provides a medical test report analysis system based on big data. Its data acquisition module collects the raw data of the medical test reports to be analyzed through the HL7 FHIR protocol interface of the hospital's HIS and LIS systems. It uses RabbitMQ message queue to handle peak concurrent requests and uses a breakpoint resume mechanism to ensure that data is not lost when the network is interrupted. The collected content includes the index values ​​of test items such as patients' blood routine, liver function, kidney function, blood glucose, and blood lipids. At the same time, it collects patients' historical diagnosis and treatment data through the hospital's electronic medical record system, obtains relevant case data from the national clinical case database, and obtains indicator standard threshold data from the medical industry standard database.

[0083] The data preprocessing module uses the Raida criterion to remove outliers in the blood routine data whose white blood cell count exceeds 3 times the standard deviation. It normalizes the data of liver function indicators such as alanine aminotransferase and aspartate aminotransferase using the Z-score standardization formula. For missing values, the K-nearest neighbor interpolation method is used to supplement numerical data. The mode is used to fill categorical data. The deep learning-optimized OCR recognition technology is used to extract the values ​​of indicators such as cholesterol and triglycerides from the scanned paper test reports, with a recognition accuracy of 98.7%, and converts them into structured data.

[0084] The big data storage module uses a MySQL relational database to store standardized indicator data and patient basic information, and a MongoDB non-relational database to store scanned copies of original test reports and clinical case text data, supporting concurrent read and write of 1000+ data entries per second;

[0085] The indicator association analysis module compares alanine aminotransferase (ALT) values ​​with standard thresholds to determine whether they are abnormal. Using the Apriori algorithm with a minimum support of 0.15 and a minimum confidence of 0.8, it uncovers the association between elevated ALT, elevated aspartate aminotransferase (AST), and elevated bilirubin. Combining the patient's historical hepatitis diagnosis and treatment data with data from 100,000 clinical hepatitis cases, it employs a random forest algorithm containing 100 decision trees, inputting 12 features such as liver function indicators, history of viral infection, and history of alcohol consumption, to assess the patient's risk of hepatitis recurrence.

[0086] The abnormal warning module issues a Level 1 warning for cases of significantly elevated alanine aminotransferase (ALT) (exceeding the standard threshold by 2 times). The warning information is pushed through three channels: system pop-up window, mobile APP notification, and SMS. The content includes "Abnormally elevated ALT (280 U / L, standard threshold 0-40 U / L), associated abnormal AST and bilirubin indicators, high risk of hepatitis recurrence (92%), and it is recommended to further check hepatitis virus markers within 24 hours." The warning is based on the requirements of the "Industry Guidelines for the Interpretation of Clinical Laboratory Results".

[0087] The results generation module integrates the above information to generate a PDF analysis report containing a summary table of indicators, anomaly details, correlation analysis conclusions, risk assessment results, and early warning information. Users upload test report data through the interactive interface, view the analysis report, and enter feedback such as "suggest adding correlation analysis of cirrhosis-related indicators." After collecting 100 similar feedback comments, the system adjusts the feature weights of the correlation rule algorithm, increases the correlation mining dimensions of cirrhosis-related indicators such as albumin and prothrombin time, and improves the accuracy of the model's early warning of cirrhosis complications by 3.2%.

[0088] Example 2

[0089] This embodiment provides a method for analyzing medical test reports based on big data, including the following steps:

[0090] Data Acquisition: Raw blood routine and biochemical test data of 65-year-old patient Zhang were collected, including white blood cell count, red blood cell count, hemoglobin, blood glucose (11.2 mmol / L), creatinine, blood urea nitrogen, and other indicators. Data on Zhang's diabetes treatment over the past 3 years (including insulin dosage, blood glucose fluctuation curve, and history of complications) were also collected. Test data of 2,000 elderly diabetic patients were obtained from the clinical case database. Standard thresholds for the above indicators were obtained from the medical standard database. Data synchronization with the HIS-LIS system was achieved within 3 seconds through the HL7FHIR protocol, with a synchronization success rate of 99.98%.

[0091] Data preprocessing: missing values ​​in urea nitrogen data were removed, and blood glucose and creatinine data were normalized using the Z-score standardization formula. Zhang's paper test report was scanned and converted into structured data using OCR recognition. Image enhancement technology was used to optimize blurry text in the scanned document to ensure error-free extraction of key indicators.

[0092] Data storage: Standardized indicator data and Zhang's basic information (including age, gender, course of disease, etc.) are stored in a MySQL database, and original scans and data from 2,000 cases are stored in a MongoDB database. Data sharding technology is used to achieve distributed storage of massive amounts of data.

[0093] Correlation analysis: Comparison revealed that Zhang's blood glucose level (11.2 mmol / L, standard threshold 3.9-6.1 mmol / L) was abnormal. The Apriori algorithm was used to discover a strong correlation between elevated blood glucose, elevated creatinine, and decreased hemoglobin (Lift = 2.8). Combining Zhang's past diabetes treatment data, data from 2000 cases, and regional diabetes epidemiology data, a neural network model with 3 hidden layers (activation function ReLU) was used. With 15 features including age, blood glucose level, creatinine level, disease duration, and blood pressure as input, the risk of diabetic nephropathy was assessed to be 85%.

[0094] Abnormal Warning: A Level 1 warning is issued for high-risk cases, generating the warning message "Abnormally high blood glucose (11.2 mmol / L), associated abnormal creatinine and hemoglobin indicators, high risk of diabetic nephropathy (85%), it is recommended to check urinary microalbumin and kidney ultrasound". The warning level is set according to the risk stratification standard in the "Guidelines for the Diagnosis and Treatment of Diabetic Nephropathy".

[0095] Results generation: Generates a Word format analysis report containing a summary of indicators, details of anomalies, correlation analysis, risk assessment, and early warning information. The report includes line charts of indicator trends and heatmaps of correlation relationships, allowing physicians to intuitively view data changes.

[0096] Interactive feedback: After reviewing the report, the physician provided feedback that "the risk assessment model needs to optimize the weighting of elderly patients." After collecting 50 feedback responses from elderly patients, the system adjusted the age feature weights of the neural network algorithm (from 0.12 to 0.18), supplemented the model with data from 1,000 diabetic patients over 60 years old, and retrained the model. After optimization, the accuracy of risk assessment for elderly patients increased by 4.5%, effectively reducing the rate of missed diagnoses in the elderly population.

[0097] In conclusion,

[0098] This invention discloses a big data-based medical laboratory report analysis system and method, focusing on the core pain points of medical laboratory data processing. It constructs a comprehensive solution encompassing "multi-source data integration, intelligent analysis, tiered early warning, and interactive optimization." The system collects multi-dimensional data, including laboratory reports, historical medical records, and clinical cases, through standardized interfaces. After preprocessing such as cleaning, normalization, and structure transformation, it leverages a distributed storage architecture to achieve efficient storage and retrieval of massive amounts of data. The core technology employs the Apriori association rule algorithm to uncover potential correlations between indicators, combined with machine learning models such as random forests and neural networks to complete personalized risk assessments. Coupled with a three-tiered early warning mechanism, it accurately identifies high-risk situations and pushes targeted recommendations.

[0099] Compared to traditional manual interpretation and single-indicator analysis, this invention significantly improves the efficiency and accuracy of laboratory report analysis, effectively addressing industry challenges such as insufficient correlation information mining, inconsistent diagnostic standards, and poor adaptability to individual differences. Its visual interactive interface and multi-format report export function are adapted to actual clinical needs, and its interactive feedback mechanism supports dynamic algorithm optimization. It can be widely applied in various scenarios, including tertiary hospitals and primary healthcare institutions, helping physicians reduce their workload, lower the risk of missed or misdiagnosed diagnoses, and provide reliable data support for accurate clinical diagnosis. This invention not only promotes the transformation of medical testing from "data collection" to "intelligent interpretation" but also facilitates the optimal allocation of medical resources, possessing significant practical value and promotional potential for improving the quality of medical services and promoting the intelligent application of medical data.

[0100] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.

Claims

1. A medical laboratory report analysis system based on big data, characterized in that: It includes a data acquisition module, a data preprocessing module, a big data storage module, an indicator correlation analysis module, an anomaly early warning module, a result generation module, and a user interaction module, which are connected in sequence. The data acquisition module is used to collect multi-source data, including raw data from medical test reports, patient history of medical treatment data, clinical case database data, medical indicator standard threshold data, and epidemiological statistics. The data preprocessing module is used to clean, standardize, and convert the format of multi-source data. The big data storage module adopts a distributed storage architecture, including a relational database for storing structured data and a non-relational database for storing unstructured data and massive historical data; The indicator correlation analysis module is used to perform single indicator anomaly detection, multi-indicator correlation mining, and comprehensive risk assessment based on big data algorithms. The anomaly warning module is used to provide graded warnings and generate warning information based on the analysis results; The result generation module is used to integrate analysis results and early warning information, generate standardized analysis reports, and support export in multiple formats. The user interaction module provides a visual operation interface, supporting data uploading, result viewing, parameter adjustment, report export, and feedback input.

2. The medical test report analysis system based on big data according to claim 1, characterized in that: The raw data of the medical test reports collected by the data acquisition module includes the index values, test time, and test equipment model information of blood test, biochemical test, immunological test, and microbiological test items.

3. The medical test report analysis system based on big data according to claim 1, characterized in that: The data preprocessing module uses the Laida criterion to remove outlier data, performs data normalization based on the Z-score standardization formula, and uses OCR recognition technology to convert unstructured data into structured data.

4. The medical test report analysis system based on big data according to claim 1, characterized in that: The big data algorithms used in the indicator correlation analysis module include association rule algorithms and machine learning models; the association rule algorithm is the Apriori algorithm, which is used to mine the correlation between multiple indicators; the machine learning models include random forest algorithms and neural network algorithms, which are used for comprehensive risk assessment.

5. The medical test report analysis system based on big data according to claim 1, characterized in that: The anomaly warning module has three warning levels: Level 1, Level 2, and Level 3. The warning information includes the name of the anomaly indicator, the degree of anomaly, related indicators, risk assessment results, and recommended measures.

6. A method for analyzing medical test reports based on big data, characterized in that, Applying the system according to any one of claims 1-5 includes the following steps: S1: Data collection, collecting multi-source data, including raw data from medical test reports, patient history of medical treatment data, clinical case database data, medical indicator standard threshold data, and epidemiological statistics data; S2: Data preprocessing, which involves cleaning, standardizing, and format conversion of the collected multi-source data; S3: Data storage, storing pre-processed structured data in a relational database, and unstructured data and massive historical data in a non-relational database; S4: Indicator correlation analysis, based on big data algorithms to identify single indicator anomalies, mine multi-indicator correlations, and conduct comprehensive risk assessment; S5: Anomaly warning, which provides graded warnings and generates warning information based on the analysis results; S6: Results generation, integrates analysis results and early warning information, generates standardized analysis reports and supports export in multiple formats; S7: Interactive feedback, which allows users to view, export reports, and enter feedback through a visual interface. The system optimizes algorithms and analysis standards based on the feedback.

7. The method for analyzing medical test reports based on big data according to claim 6, characterized in that: In step S2, data cleaning includes removing missing values, outliers, and duplicate data; standardization processing includes converting indicator data from different units into a unified standard unit and performing normalization processing; and format conversion includes converting unstructured data into structured data.

8. The method for analyzing medical test reports based on big data according to claim 6, characterized in that: In step S4, single-indicator anomaly judgment is achieved by comparing the indicator value with the standard threshold data; multi-indicator correlation mining uses the Apriori algorithm; comprehensive risk assessment combines the patient's historical treatment data, clinical case data and epidemiological data, and is implemented using the random forest algorithm or neural network algorithm.

9. The method for analyzing medical test reports based on big data according to claim 6, characterized in that: In step S5, the first-level warning corresponds to a single indicator with severe abnormality, a combination of high-risk indicators, or a high risk of disease; the second-level warning corresponds to an important indicator with abnormality or a medium-risk situation; and the third-level warning corresponds to a general indicator with abnormality or a low-risk situation.

10. The method for analyzing medical test reports based on big data according to claim 6, characterized in that: In step S6, the standardized analysis report includes a summary table of test indicators, details of abnormal indicators, conclusions of indicator correlation analysis, disease risk assessment results, early warning information, and clinical diagnostic recommendations. Supported export formats include PDF, Word, and Excel.