Lung disease data acquisition method based on high-precision medical data platform
Through standardized data collection and processing based on a high-precision medical data platform, combined with relational and NoSQL databases, comprehensive collection and efficient analysis of severe pneumonia data are achieved, solving the problem of incomplete data collection in existing technologies, improving data accuracy and security, and meeting the data access needs of multiple users and multiple hospitals.
Patent Information
- Application Number
- CN202510406562.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-09-12
AI Technical Summary
The existing technology for diagnosing severe pneumonia has problems with early identification and inaccurate pathogen identification. The data collection and analysis system lacks comprehensiveness and fails to effectively utilize patients' treatment plans, medication data, test data, examination data and imaging data, resulting in an insufficient understanding of severe pneumonia.
Based on a high-precision medical data platform, standardized processes are used to collect and process multi-dimensional data, combined with relational and NoSQL databases to achieve secure data storage and management. Data cleaning and standardization are performed through template management, data source management, data verification and manual entry, and machine learning algorithms are used for data analysis and visualization.
It achieves data compatibility and security across hospital systems, ensures data accuracy and consistency, supports multi-user access, meets different research needs, and improves the efficiency of data collection and analysis of severe pneumonia.
Smart Images

Figure CN120636658A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of medical data classification, and in particular to a method for collecting lung disease data based on a high-precision medical data platform. Background Art
[0002] Severe pneumonia develops rapidly, has complex etiologies, and has a high mortality rate. Diagnosis is plagued by difficulties and challenges, including early identification and inaccurate pathogen identification. Common symptoms of coronavirus infection include respiratory symptoms, fever, cough, shortness of breath, and difficulty breathing. Currently, data collection and analysis systems collect public health data such as the name, ID card, address, time of diagnosis, and history of epidemic areas from confirmed patients through epidemic report cards. These data rarely include treatment plans, medication, laboratory tests, examinations, imaging, or epidemiological data, hindering a comprehensive understanding of severe pneumonia. Summary of the Invention
[0003] The purpose of the present invention is to propose a lung disease data collection method based on a high-precision medical data platform to solve one or more technical problems existing in the prior art and at least provide a beneficial option or create conditions.
[0004] The lung disease data collection method, based on a high-precision medical data platform, utilizes standardized processes for data collection and processing across different data types. Data collection covers patient baseline information, laboratory test data, imaging data, interventions, and clinical outcomes. Laboratory test data will be uniformly converted to international standard units, and imaging data will be stored in DICOM format to ensure compatibility across hospital systems. Metadata will be appended to all data collection processes.
[0005] In terms of data processing, all multimodal data will be cleaned and standardized to ensure data consistency, completeness and accuracy. All patients' clinical data will be recorded using standardized forms, including scores on the day of admission, including PSI, APACHE II, CURB-65 and SOFA scores, as well as dynamic changes in the disease. The data cleaning process includes checking missing values, removing duplicate data and processing outliers to ensure the quality of the data finally included in the database is reliable.
[0006] The database architecture will adopt a hybrid model, combining the strengths of relational and NoSQL databases. For structured data, relational databases such as MySQL or PostgreSQL will be used to ensure consistent and scalable data storage. For unstructured data, NoSQL databases such as MongoDB will be used to support parallel processing and distributed storage of massive amounts of data, enhancing flexibility in data expansion and retrieval.
[0007] The database will utilize a distributed storage structure to ensure data security and efficient management, supporting simultaneous access by multiple users and hospitals. The system provides a RESTful API interface to facilitate data access and sharing with external systems, enabling flexible data invocation and integration. Regarding data permission management, the system will set different access levels based on user roles to ensure the privacy of sensitive data in compliance with international privacy regulations. Patients' personal privacy information will be anonymized, and encryption technology will be used to ensure data security during storage and transmission.
[0008] Furthermore, the data acquisition system includes a template management unit, a data source management unit, a data verification unit and a manual entry editing unit; The template management unit includes a mapping construction module and a template splitting module. The mapping construction module is used to construct a data mapping template, and the data mapping template is used to realize the standardized mapping processing of HIS data, LIS data, PACS data and disease prevention and control system data; the template splitting module is used to split the data mapping template into multiple business-related data models, namely, a treatment plan model, a medication data model, a test data model, an examination data model, an imaging data model and an epidemiological data model, and generate a data collection SQL script for each data model; The data source management unit is used to record and store configuration information in different hospital management systems accessed during data collection, including data source driver files, database names, URLs, and login information configurations, while providing connectivity support for existing medical data source configurations for data collection; The data verification unit includes a data execution module and a data verification module. The data execution module is used to implement simulation operation based on the data model acquisition process and display the operation result set to obtain data information that meets the data model standard; the data verification module is used to verify the compliance of data during the data acquisition process and provide abnormal prompts for data that does not meet the standards; The manual entry editing unit edits the missing data items in the data mapping template using a manual entry method.
[0009] Furthermore, the data analysis system includes a data fusion and preprocessing module and a machine learning algorithm module. The data fusion and preprocessing module is used to process missing values and convert categorized data; the machine learning algorithm module is used to perform data analysis and result visualization.
[0010] Furthermore, the data fusion and preprocessing module connects and fuses the disease prevention and control system data with the HIS data, LIS data, and PACS data based on the patient's ID number or patient identification number. For each patient after the fused data, the missing rate of all data items of the patient is calculated, and patients and their corresponding data whose missing rate exceeds the threshold are eliminated, and data features are supplemented for missing features that do not exceed the threshold.
[0011] Furthermore, the machine learning algorithm module presets multiple machine learning algorithms, encapsulates the machine learning algorithms into function form, and the user selects the machine learning algorithm and sets the algorithm parameters. The machine learning algorithm module receives the data table output by the data fusion and preprocessing module, converts the data table into dataframe format data, and uses the dataframe format data and the user-set algorithm parameters as input of the user-selected function to complete data analysis, and finally visualizes the analysis results in the form of charts.
[0012] Furthermore, in the data feature completion process, the following data feature completion algorithm is introduced, specifically as follows: Let X = (x(1), x(2), ..., x(p)), X is an n×p matrix, for a given variable x(s), its missing indicator set is to divide the data into four parts: , represents the observed value of x(s), represents the missing value of x(s), Represents the data missing in the rows of x(s) corresponding to the remaining p-1 variables. Represents the supplementary data for the corresponding rows of the remaining p-1 variables; E(n×p) = ; Where E(n×p) represents the rows in the n×p matrix that do not belong to X; Due to the randomness of missing data, it is not completely known, nor is it completely missing. The specific filling process is as follows: (a) Initially fill in X using mean filling or other simple filling methods; (b) The indicator set of missing columns in X is denoted as M, and the variables (columns) are arranged from small to large according to the missing rate; (c) When the stopping criterion γ is not satisfied, store the existing filling matrix, which is recorded as for , use the random forest algorithm to build y and x models, and after the model is built, use E (n × p) to predict Using the predicted values Update the filling matrix, which is recorded as continuing to fill in the remaining missing variables in s until the stopping criterion γ is met; (d) The final filling matrix is obtained, denoted as ; The stopping criterion γ is: if the difference between the new filling matrix and the previous filling matrix increases, then the loop stops, where the difference of continuous variables is: where, Indicates the value after filling; Indicates the value before filling; The difference between discrete variables is: F=*NA ; Where *NA indicates the number of missing data for discrete variables.
[0013] Furthermore, the data fusion and preprocessing module connects and fuses the epidemiological data with the treatment plan, medication data, test data, examination data, and imaging data based on the patient's ID number or patient identification number.
[0014] Furthermore, the machine learning algorithm module adopts any one of linear regression, logistic regression, support vector machine, and random forest machine learning algorithms, encapsulates the machine learning algorithm into a function form, and the user can select the machine learning algorithm and set the algorithm parameters.
[0015] Compared with other acquisition systems, the advantages of the present invention are: The present invention is constructed based on the characteristics of pneumonia virus and is highly professional. During the data collection and analysis process, epidemiological data is connected with the patient's treatment plan, test data, examination data, medication data, and imaging data according to the patient identification number or ID number, realizing the integration of public health data and hospital internal information system data, which is conducive to a comprehensive understanding of pneumonia virus.
[0016] The pre-processed data is converted into dataframe format data, and combined with multiple preset machine learning algorithms of the present invention to form an efficient data analysis system, which can meet the different research needs of researchers. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The above and other features of the present invention will become more apparent through a detailed description of the embodiments shown in conjunction with the accompanying drawings. In the drawings of the present invention, the same reference numerals represent the same or similar elements. Obviously, the drawings described below are only some embodiments of the present invention. It is possible for a person skilled in the art to derive other drawings based on these drawings without inventive effort. In the drawings: Figure 1 This is a flow chart of the lung disease data collection method based on a high-precision medical data platform. DETAILED DESCRIPTION
[0018] The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0019] In the description of the present invention, "several" means one or more, "many" means more than two, "greater than," "less than," and "exceed" are understood to exclude the number itself, while "above," "below," and "within" are understood to include the number itself. The use of "first" and "second" in the description is solely for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance, implicitly specifying the number of the indicated technical features, or implicitly specifying the order of the indicated technical features.
[0020] The following describes the embodiments of the present disclosure through specific examples, and those skilled in the art can easily understand the advantages and effects of the present disclosure from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure.
[0021] The purpose of the present invention is to propose an animal monitoring method and system based on an ad hoc network infrared camera to solve one or more technical problems existing in the prior art and at least provide a beneficial option or create conditions.
[0022] (1) Data acquisition system The data collection system is used to collect information such as patient treatment plans, medication data, test data, examination data, imaging data, and epidemiological data from HIS systems, LIS systems, PACS systems, and disease prevention and control systems. The collected data is associated with the patient's ID number or patient identification number; The data acquisition system includes a template management unit, a data source management unit, a data verification unit, and a manual entry and editing unit; The template management unit includes a mapping construction module and a template splitting module. The mapping construction module is used to construct a data mapping template, and use the data mapping template to realize the standardized mapping processing of HIS data, LIS data, PACS data, and disease prevention and control system data; the template splitting module is used to split the data mapping template into multiple data models with business associations, and generate data collection SQL scripts for each data model.
[0023] The above-mentioned normalized mapping process is to address the problems of different data types and dimensions in the extracted data. In order to facilitate data analysis, the data is mapped to the same data type and dimension.
[0024] Since the data items in the data mapping template come from different systems, such as treatment plans, medication purposes, and medication plans from the HIS system; white blood cell count, activated partial thromboplastin time, and D-dimer values from the LIS system; imaging features and image files from the PACS system; and epidemiological data from the disease prevention and control system, the data items in the data mapping template are divided into medication plan models, treatment data models, test data models, inspection data models, and epidemiological data models based on the business relationships to which the data items belong, to facilitate data collection from the HIS system, LIS system, PACS system, and disease prevention and control system. A data model is a combination of data items with the same business relationship. For example, testing items such as white blood cell count and activated partial thromboplastin time are combined to form a test data model; data items such as medication purpose and medication regimen are combined to form a medication regimen model; data items such as critical illness assessment and treatment-related adverse reactions are combined to form a treatment data model; data items such as imaging features and imaging examination categories are combined to form an examination data model; and data items such as recent contact with confirmed pneumonia virus cases and the presence of respiratory symptoms are combined to form an epidemiological data model. Based on the data model, relevant data items are collected and irrelevant data items are eliminated, which reduces data collection time and facilitates standardized collection.
[0025] The data source management unit is used to record and store configuration information in different hospital management systems accessed during data collection, including data source driver files, database names, URLs, and login information configurations. It also provides connectivity support for existing medical data source configurations for data collection. Data originates from different hospitals, and to facilitate future data collection, the data source access process and configuration information are recorded and saved. The data verification unit includes a data execution module and a data validation module. The data execution module is used to implement simulation operations based on the data model acquisition process and display the operation result set to obtain data information that meets the data model standards. The data validation module is used to verify the compliance of data during the data acquisition process and to issue exception prompts for data that does not meet the standards. Based on the data items in the data model, SQL statements are written and SQL scripts are run to complete the model data acquisition. During this process, exception prompts are issued for data that does not meet the standards. The manual entry editing unit uses manual entry to edit missing data items in the data mapping template. Certain data items in the data mapping template, such as flight number and travel history to epidemic areas, are not available from HIS, LIS, PACS, or disease prevention and control systems. This patient information must be obtained by telephone or other means. Data entry personnel complete this data entry on a web page using a VPN. Manually entered data can also be modified.
[0026] (2) Data Analysis System The data analysis system is connected to the data acquisition system. The data analysis system is used to process missing values and classify the collected data, and perform data analysis and result visualization to form a highly efficient analysis system. The data analysis system includes a data fusion and preprocessing module and a machine learning algorithm module. The data fusion and preprocessing module handles missing values and converts categorical data. The machine learning algorithm module performs data analysis and result visualization, forming a highly efficient data analysis system that improves the utilization of multi-category and multi-dimensional medical data and meets the diverse research needs of researchers. The data fusion and preprocessing module connects and fuses the disease prevention and control system data with HIS data, LIS data, and PACS data based on the patient's ID number or patient identification number. For each patient after the fused data, the missing rate of all data items of the patient is calculated, and patients and their corresponding data with missing rates exceeding the threshold are eliminated, and missing features that do not exceed the threshold are supplemented.
[0027] Furthermore, in the data feature completion process, the following data feature completion algorithm is introduced, specifically as follows: Let X = (x(1), x(2), ..., x(p)), X is an n×p matrix, for a given variable x(s), its missing indicator set is to divide the data into four parts: , represents the observed value of x(s), represents the missing value of x(s), Represents the data missing in the rows of x(s) corresponding to the remaining p-1 variables. Represents the supplementary data for the corresponding rows of the remaining p-1 variables; E(n×p) = ; Where E(n×p) represents the rows in the n×p matrix that do not belong to X; Due to the randomness of missing data, it is not completely known, nor is it completely missing. The specific filling process is as follows: (a) Initially fill in X using mean filling or other simple filling methods; (b) The indicator set of missing columns in X is denoted as M, and the variables (columns) are arranged from small to large according to the missing rate; (c) When the stopping criterion γ is not satisfied, store the existing filling matrix, which is recorded as for , use the random forest algorithm to build y and x models, and after the model is built, use E (n × p) to predict Using the predicted values Update the filling matrix, which is recorded as continuing to fill in the remaining missing variables in s until the stopping criterion γ is met; (d) The final filling matrix is obtained, denoted as ; The stopping criterion γ is: if the difference between the new filling matrix and the previous filling matrix increases, then the loop stops, where the difference of continuous variables is: where, Indicates the value after filling; Indicates the value before filling; The difference between discrete variables is: F=*NA ; Where *NA represents the number of missing data for a discrete variable. The machine learning algorithm module pre-sets multiple machine learning algorithms and encapsulates them into function form, allowing the user to select a machine learning algorithm and set algorithm parameters. The machine learning algorithm module receives the data table output by the data fusion and preprocessing module, converts the data table into data in dataframe format, and uses the dataframe format and user-set algorithm parameters as input to the user-selected function to complete data analysis and ultimately visualize the analysis results in the form of charts.
[0028] (3) Result display module The result display module is connected to the data analysis system and is used to display the results of the data analysis system.
[0029] In one specific implementation, the data collection system uses a VPN (Virtual Private Network) to collect patient treatment plans and medication data from the HIS, laboratory and examination data from the LIS, imaging data from the PACS, and epidemiological data from the disease prevention and control system. Using a VPN for data transmission provides a stable and secure data transmission channel for the data collection and analysis system, ensuring data security.
[0030] In a specific embodiment, the machine learning algorithm module uses any one of the machine learning algorithms such as linear regression, logistic regression, support vector machine, random forest, etc., encapsulates the machine learning algorithm into a function form, and allows the user to select the machine learning algorithm and set the algorithm parameters; receives the data table output by fusion and preprocessing, converts the data table into dataframe format data, and uses it together with the user-set algorithm parameters as the input of the user-selected function to complete data analysis, and visualizes the analysis results in the form of charts.
[0031] In a specific embodiment, the disease prevention and control system stores pneumonia virus epidemiological data, but does not have comprehensive information such as treatment plans, medication data, test data, examination data and imaging data. According to the patient's ID number or patient identification number, the data fusion and preprocessing module connects and fuses the epidemiological data with the treatment plan, medication data, test data, examination data and imaging data. For each patient after the fused data, the missing rate of all data items of the patient is calculated, and patients and their corresponding data whose missing rate exceeds a threshold are eliminated, and the missing features that do not exceed the threshold are supplemented.
Claims
1. A lung disease data collection method based on a high-precision medical data platform, characterized in that: Standardized processes are used to collect and process data across different data types. Data collection covers patient baseline information, laboratory test data, imaging data, interventions, and clinical outcomes. Laboratory test data will be uniformly converted to international standard units, and imaging data will be stored in DICOM format to ensure compatibility across hospital systems. Metadata will be appended to all data collection processes. In terms of data processing, all multimodal data will be cleaned and standardized to ensure data consistency, completeness and accuracy. All patients' clinical data will be recorded using standardized forms, including scores on the day of admission, including PSI, APACHE II, CURB-65 and SOFA scores, as well as dynamic changes in the disease. The data cleaning process includes checking missing values, removing duplicate data and processing outliers to ensure the quality of the data finally included in the database is reliable. The database architecture will adopt a hybrid model, combining the strengths of relational and NoSQL databases. For structured data, relational databases such as MySQL or PostgreSQL will be used to ensure consistent and scalable data storage. For unstructured data, NoSQL databases such as MongoDB will be used to support parallel processing and distributed storage of massive amounts of data, enhancing flexibility in data expansion and retrieval. The database will utilize a distributed storage structure to ensure data security and efficient management, supporting simultaneous access by multiple users and hospitals. The system provides a RESTful API interface to facilitate data access and sharing with external systems, enabling flexible data invocation and integration. Regarding data permission management, the system will set different access levels based on user roles to ensure the privacy of sensitive data in compliance with international privacy regulations. Patients' personal privacy information will be anonymized, and encryption technology will be used to ensure data security during storage and transmission.
2. The lung disease data collection method based on a high-precision medical data platform according to claim 1 is characterized in that The data acquisition system includes a template management unit, a data source management unit, a data verification unit and a manual entry editing unit; The template management unit includes a mapping construction module and a template splitting module. The mapping construction module is used to construct a data mapping template, and the data mapping template is used to realize the standardized mapping processing of HIS data, LIS data, PACS data and disease prevention and control system data; the template splitting module is used to split the data mapping template into multiple business-related data models, namely, a treatment plan model, a medication data model, a test data model, an examination data model, an imaging data model and an epidemiological data model, and generate a data collection SQL script for each data model; The data source management unit is used to record and store configuration information in different hospital management systems accessed during data collection, including data source driver files, database names, URLs, and login information configurations, while providing connectivity support for existing medical data source configurations for data collection; The data verification unit includes a data execution module and a data verification module. The data execution module is used to implement simulation operation based on the data model acquisition process and display the operation result set to obtain data information that meets the data model standard; The data verification module is used to verify the compliance of data during the data collection process and to provide abnormal prompts for data that does not meet the standards; The manual entry editing unit edits the missing data items in the data mapping template using a manual entry method.
3. The lung disease data collection method based on a high-precision medical data platform according to claim 1 is characterized in that: The data analysis system includes a data fusion and preprocessing module and a machine learning algorithm module. The data fusion and preprocessing module is used to process missing values and convert categorical data; the machine learning algorithm module is used to perform data analysis and result visualization.
4. The lung disease data collection method based on a high-precision medical data platform according to claim 1 is characterized in that: The data fusion and preprocessing module connects and fuses the disease prevention and control system data with the HIS data, LIS data, and PACS data based on the patient's ID number or patient identification number. For each patient after the data is fused, the missing rate of all data items of the patient is calculated, and patients and their corresponding data whose missing rate exceeds the threshold are eliminated, and data features are supplemented for missing features that do not exceed the threshold.
5. The lung disease data collection method based on a high-precision medical data platform according to claim 1 is characterized in that: The machine learning algorithm module presets multiple machine learning algorithms, encapsulates the machine learning algorithms into function form, and allows the user to select the machine learning algorithm and set the algorithm parameters. The machine learning algorithm module receives the data table output by the data fusion and preprocessing module, converts the data table into dataframe format data, and uses the dataframe format data and the user-set algorithm parameters as inputs of the user-selected function to complete data analysis, and finally visualizes the analysis results in the form of charts.
6. The lung disease data collection method based on a high-precision medical data platform according to claim 1 is characterized in that: In the data feature completion process, the following data feature completion algorithm is introduced, specifically as follows: Let X = (x(1), x(2), ..., x(p)), X is an n×p matrix, for a given variable x(s), its missing indicator set is divided into four parts: x(s) a 、x(s) b 、x(s) p 、x(s) q , the x(s) a represents the observed value of x(s), said x(s) b represents the missing value of x(s), the x(s) p Represents the data missing in x(s) for the rows corresponding to the remaining p-1 variables, where x(s) q Represents the supplementary data for the corresponding rows of the remaining p-1 variables; Where E(n×p) represents the rows in the n×p matrix that do not belong to X; Due to the randomness of missing data, it is not completely known, nor is it completely missing. The specific filling process is as follows: (a) Initially fill in X using mean filling or other simple filling methods; (b) The indicator set of missing columns in X is denoted as M, and the variables (columns) are arranged from small to large according to the missing rate; (c) When the stopping criterion γ is not met, store the existing filling matrix, denoted as γ s For s∈M, use the random forest method to build y and x models. After the model is built, use E(n×p) to predict γ s Using the predicted value γ s+1 Update the filling matrix, which is recorded as continuing to fill in the remaining missing variables in s until the stopping criterion γ is met; (d) Get the final filling matrix, denoted as X imp ; The stopping criterion γ is: if the difference between the new filling matrix and the previous filling matrix increases, then the loop stops, where the difference of continuous variables is: where x ijnew Indicates the value after filling; x ijold Indicates the value before filling; The difference between discrete variables is: Where *NA indicates the number of missing data for discrete variables.
7. The lung disease data collection method based on a high-precision medical data platform according to claim 1 is characterized in that: The data fusion and preprocessing module connects and fuses the epidemiological data with the treatment plan, medication data, test data, examination data, and imaging data according to the patient's ID number or patient identification number.
8. The lung disease data collection method based on a high-precision medical data platform according to claim 1 is characterized in that: The machine learning algorithm module adopts any one of linear regression, logistic regression, support vector machine, and random forest machine learning algorithms, encapsulates the machine learning algorithm into a function form, and the user can select the machine learning algorithm and set the algorithm parameters.