Disease library construction method and system based on big data

Through the disease database construction method and system based on big data, clinical data from different medical systems are collected, preprocessed and featured extracted, a category-specific disease database is constructed, and dynamic data information is updated in real time, solving the problem of difficulty in sharing and integrating traditional medical data, and achieving efficient and accurate disease database construction and update.

CN120015342AInactive Publication Date: 2025-05-16SUZHOU MUNICIPAL HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510084224.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The sharing and integration of traditional medical data is limited by problems such as inconsistent information system standards and incompatible data formats, which lead to data silos and hinder cross-institutional and cross-regional scientific research cooperation and clinical decision-making support. At the same time, data quality is uneven and there is a lack of unified data standards and coding systems, which affects data analysis and application.

Method used

A disease database construction method and system based on big data is proposed. By collecting and preprocessing clinical data from different medical systems, features are extracted and data aggregation are carried out, feature information is associated and classified, feature database is constructed, and category feature database is updated in real time through data dynamic information.

Benefits of technology

Cross-system clinical data integration is realized, the availability and value of data is improved, the accuracy and consistency of data is ensured, the understanding of complex data relationships is enhanced, the timeliness and accuracy of the disease database is improved, and the allocation of medical resources and diagnosis and treatment efficiency is optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015342A_ABST
    Figure CN120015342A_ABST
Patent Text Reader

Abstract

The invention provides a disease category database construction method and system based on big data, and relates to the technical field of database construction. Clinical data are collected and preprocessed, feature extraction and data aggregation are carried out, and aggregated feature data are obtained; data association of the data features is carried out, classification is carried out according to preset feature association categories, a medical data view of each preset feature association category is generated, and then feature subsets are obtained; constructing a category feature disease category library, calculating a data dynamic coefficient, analyzing the data dynamic coefficient, triggering an exclusion instruction, and obtaining inclusion and exclusion information; the inclusion information and the exclusion information are compared with a preset inclusion threshold value and a preset exclusion threshold value respectively, a database updating instruction is triggered according to the comparison result, then database increment updating or database decrement updating is carried out, and dynamic processing, updating and optimization of the disease database are achieved through an efficient data processing and analysis method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention proposes a method and system for constructing a disease database based on big data, which relates to the technical field of database construction, and specifically to the technical field of constructing a disease database based on big data. Background Art

[0002] With the acceleration of the global medical informationization process, various medical institutions have accumulated a large amount of clinical data in their daily operations, including but not limited to electronic medical records, examination and test results, medical images, gene sequencing information, etc. These data not only reflect the patient's health status and treatment process, but also contain a wealth of valuable information such as disease characteristics, development patterns and treatment effects. However, traditionally, the mining and utilization of these data are subject to multiple restrictions: due to different information system construction standards and incompatible data formats among different medical institutions, data sharing is difficult, forming data islands. This not only limits the integration and analysis of data, but also hinders cross-institutional and cross-regional scientific research cooperation and clinical decision support. In the process of collecting, recording and processing clinical data, human factors and system differences often lead to uneven data quality, such as missing data, errors, duplications, etc. At the same time, the lack of unified data standards and coding systems makes it difficult to standardize data processing, affecting subsequent data analysis and application. Faced with massive and multi-dimensional clinical data, traditional data processing methods seem to be unable to cope with it. Especially in complex tasks such as disease classification, feature extraction and association, more efficient and intelligent algorithms and technical support are needed. It is difficult to perform sorting processing and dynamic update management within the database. Summary of the invention

[0003] The present invention provides a method and system for constructing a disease database based on big data to solve the above problems:

[0004] The present invention proposes a method and system for constructing a disease database based on big data, the method comprising:

[0005] S1. Collect and preprocess clinical data from different medical systems to obtain clinical processing data from multiple medical systems, perform feature extraction and data aggregation on the clinical processing data to obtain aggregated feature data;

[0006] S2, performing data association of data features on the aggregated feature data and classifying them according to preset feature association categories, obtaining category feature association information, generating a medical data view of each preset feature association category, and then obtaining a feature subset;

[0007] S3. Build a category characteristic disease database, obtain data dynamic information, calculate data dynamic coefficients, analyze the data dynamic coefficients, and then trigger exclusion instructions to include and exclude the category characteristic database to obtain inclusion and exclusion information;

[0008] S4. Compare the inclusion and exclusion information with the preset inclusion threshold and exclusion threshold respectively, trigger a database update instruction according to the comparison result, and then perform a database incremental update or a database decrement update.

[0009] Furthermore, the S1 includes:

[0010] Collect clinical data from each medical system to obtain clinical collection data from each medical system;

[0011] Preprocess each clinical collection data to obtain corresponding clinical processing data;

[0012] Acquire preset feature requirement information, perform feature extraction on clinical processing data of each medical system according to the preset feature requirement information, and obtain system feature extraction data of each medical system;

[0013] The feature extraction data of each system is aggregated to obtain aggregated feature data.

[0014] Further, the S2 includes:

[0015] In the aggregated feature data, data association is performed on the system feature extraction data of each medical system to obtain feature association information;

[0016] Obtaining a preset feature association category, classifying all feature association information according to the preset feature association category, and obtaining category feature association information of multiple categories;

[0017] Generate a medical data view corresponding to a preset feature association category according to each category feature association information;

[0018] Obtain a preset correlation threshold, compare the total number of correlations between each data feature in the medical data view and other features with the preset correlation threshold, and obtain a feature correlation comparison result;

[0019] The feature subsets of the data features of the category feature association information are screened according to the feature association comparison result to obtain multiple feature subsets of each category feature association information.

[0020] Further, the S3 includes:

[0021] The disease database is constructed by aggregating feature data, category feature association information and feature subsets to obtain a category feature disease database;

[0022] Obtain data dynamic information of the category characteristic disease database, based on the data dynamic information;

[0023] Calculate the data dynamic coefficient of the category characteristic disease database based on data dynamic information;

[0024] Comparing the data dynamic coefficient with a preset dynamic threshold to obtain a data dynamic comparison result;

[0025] Triggering a sorting instruction according to the dynamic comparison result of the data;

[0026] Obtaining inclusion conditions and exclusion conditions according to the inclusion and exclusion instructions;

[0027] The category feature database is included and excluded according to the inclusion conditions and exclusion conditions to obtain inclusion and exclusion information.

[0028] Further, the S4 includes:

[0029] The inclusion and exclusion information are compared with the preset inclusion and exclusion thresholds, respectively;

[0030] When the inclusion and exclusion information is greater than the preset inclusion threshold or exclusion threshold, a database update instruction is triggered;

[0031] When the included information is greater than the preset inclusion threshold, the database incremental update instruction is triggered;

[0032] When the exclusion information is greater than the preset exclusion threshold, a database reduction update instruction is triggered.

[0033] Furthermore, the system comprises:

[0034] A data feature aggregation module is used to collect and preprocess clinical data from different medical systems, obtain clinical processing data from multiple medical systems, perform feature extraction and data aggregation on the clinical processing data, and obtain aggregated feature data;

[0035] A category data association extraction module is used to perform data association on the aggregated feature data and classify it according to the preset feature association categories, obtain category feature association information, generate a medical data view of each preset feature association category, and then obtain a feature subset;

[0036] The database construction processing module is used to construct a category characteristic disease database, obtain data dynamic information, calculate data dynamic coefficients, analyze data dynamic coefficients, and then trigger exclusion instructions to include and exclude the category characteristic database and obtain inclusion and exclusion information;

[0037] The regional update module is used to compare the inclusion and exclusion information with the preset inclusion threshold and exclusion threshold respectively, trigger the database update instruction according to the comparison result, and then perform the database incremental update or database decrement update.

[0038] Furthermore, the data feature aggregation module includes:

[0039] Collect clinical data from each medical system to obtain clinical collection data from each medical system;

[0040] Preprocess each clinical collection data to obtain corresponding clinical processing data;

[0041] Acquire preset feature requirement information, perform feature extraction on clinical processing data of each medical system according to the preset feature requirement information, and obtain system feature extraction data of each medical system;

[0042] The feature extraction data of each system is aggregated to obtain aggregated feature data.

[0043] Furthermore, the category data association extraction module includes:

[0044] In the aggregated feature data, data association is performed on the system feature extraction data of each medical system to obtain feature association information;

[0045] Obtaining a preset feature association category, classifying all feature association information according to the preset feature association category, and obtaining category feature association information of multiple categories;

[0046] Generate a medical data view corresponding to a preset feature association category according to each category feature association information;

[0047] Obtain a preset correlation threshold, compare the total number of correlations between each data feature in the medical data view and other features with the preset correlation threshold, and obtain a feature correlation comparison result;

[0048] The feature subsets of the data features of the category feature association information are screened according to the feature association comparison result to obtain multiple feature subsets of each category feature association information.

[0049] Furthermore, the database construction processing module includes:

[0050] The disease database is constructed by aggregating feature data, category feature association information and feature subsets to obtain a category feature disease database;

[0051] Obtain data dynamic information of the category characteristic disease database, based on the data dynamic information;

[0052] Calculate the data dynamic coefficient of the category characteristic disease database based on data dynamic information;

[0053] Comparing the data dynamic coefficient with a preset dynamic threshold to obtain a data dynamic comparison result;

[0054] Triggering a sorting instruction according to the dynamic comparison result of the data;

[0055] Obtaining inclusion conditions and exclusion conditions according to the inclusion and exclusion instructions;

[0056] The category feature database is included and excluded according to the inclusion conditions and exclusion conditions to obtain inclusion and exclusion information.

[0057] Furthermore, the area update module includes:

[0058] The inclusion and exclusion information are compared with the preset inclusion and exclusion thresholds, respectively;

[0059] When the inclusion and exclusion information is greater than the preset inclusion threshold or exclusion threshold, a database update instruction is triggered;

[0060] When the included information is greater than the preset inclusion threshold, the database incremental update instruction is triggered;

[0061] When the exclusion information is greater than the preset exclusion threshold, a database reduction update instruction is triggered.

[0062] Beneficial effects of the invention: The accuracy and consistency of the data used to construct the disease library are ensured through data preprocessing and feature extraction. Data feature association and classification can reveal the potential relationship between different features, providing more comprehensive information for the construction of the disease library. By continuously monitoring data dynamics and triggering exclusion instructions, the disease library can reflect the latest clinical data and disease characteristics in real time, improving the timeliness and accuracy of the disease library. The construction of a disease library based on big data can provide disease distribution data and patient demand data, thereby optimizing resource allocation and improving the efficiency of medical services. The disease library provides doctors with rich disease information and feature descriptions, which can help doctors diagnose diseases, formulate treatment plans and evaluate prognosis more accurately. The disease library construction method and system based on big data realize the dynamic update and optimization of the disease library through efficient data processing and analysis technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 This is a schematic diagram of a method for constructing a disease database based on big data. DETAILED DESCRIPTION

[0064] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0065] In one embodiment of the present invention, a method and system for constructing a disease database based on big data is proposed by the present invention, and the method comprises:

[0066] S1. Collect and preprocess clinical data from different medical systems to obtain clinical processing data from multiple medical systems, perform feature extraction and data aggregation on the clinical processing data to obtain aggregated feature data;

[0067] S2, performing data association of data features on the aggregated feature data and classifying them according to preset feature association categories, obtaining category feature association information, generating a medical data view of each preset feature association category, and then obtaining a feature subset;

[0068] S3. Build a category characteristic disease database, obtain data dynamic information, calculate data dynamic coefficients, analyze the data dynamic coefficients, and then trigger exclusion instructions to include and exclude the category characteristic database to obtain inclusion and exclusion information;

[0069] S4, comparing the inclusion and exclusion information with the preset inclusion threshold and exclusion threshold respectively, triggering a database update instruction according to the comparison result, and then performing an incremental update of the database or a decremental update of the database, such as Figure 1 shown.

[0070] The working principle of the above technical solution is as follows: clinical data is collected from multiple medical systems, which may include multiple electronic medical record systems, imaging systems, laboratory information systems, etc. of multiple hospitals. The collected data is cleaned, formatted, and missing values ​​are processed to ensure data quality and consistency. Key features such as patient age, gender, diagnosis information, treatment process, etc. are extracted from the preprocessed data, and these features are aggregated to form aggregated feature data. The relationship between different features in the aggregated feature data is analyzed, such as the relationship between disease and symptoms, and between disease and treatment methods. The associated information is classified according to the preset feature association categories (such as disease type, treatment method classification, etc.), and medical data views are generated, which help understand and analyze data features. Specific feature subsets are extracted from the classified data for subsequent construction of the disease library. The initial disease library is constructed based on the feature subset, including feature descriptions of different diseases. New clinical data is continuously monitored and acquired, and the data dynamic coefficient is calculated, which reflects the update frequency and degree of change of the data. The data dynamic coefficient is analyzed, and when the data change reaches the preset conditions, the inclusion or exclusion instructions are triggered to update the disease library. The inclusion and exclusion information is compared with the preset inclusion and exclusion thresholds. Based on the comparison results, if the data meets the inclusion criteria, the database is incrementally updated (adding new data or disease types); if the exclusion criteria are met, the database is decrementally updated (deleting data or disease types that no longer meet the criteria).

[0071] Assistance steps include:

[0072] Collect and integrate clinical data from different medical systems, including medical records, examination and test reports, imaging data, etc.

[0073] Clean, deduplicate and standardize the data to ensure its quality and consistency.

[0074] Use natural language processing technology to parse unstructured data (such as medical record text) and extract key information.

[0075] Based on machine learning algorithms (such as cluster analysis, decision tree, support vector machine, etc.), disease characteristics are extracted from the preprocessed data.

[0076] Build a disease classification model based on the characteristics to achieve automatic identification and classification of diseases.

[0077] Continuously optimize algorithm parameters to improve the accuracy and stability of disease classification.

[0078] Integrate the classified disease information into a disease database, including disease name, characteristic description, diagnostic criteria, treatment plan, etc.

[0079] Establish a dynamic update mechanism for the disease database to continuously update and improve disease information based on new data and new knowledge.

[0080] It provides query, retrieval and analysis functions for the disease database, supporting application scenarios such as clinical decision-making and disease research.

[0081] During data processing and storage, encryption, desensitization and other technologies are used to protect patient privacy and data security.

[0082] The technical effects of the above technical scheme are as follows: through data preprocessing and feature extraction, the accuracy and consistency of the data used to construct the disease library are ensured. Data feature association and classification can reveal the potential relationship between different features and provide more comprehensive information for the construction of the disease library. By continuously monitoring data dynamics and triggering exclusion instructions, the disease library can reflect the latest clinical data and disease characteristics in real time, improving the timeliness and accuracy of the disease library. The construction of a disease library based on big data can provide disease distribution data and patient demand data, thereby optimizing resource allocation and improving medical service efficiency. The disease library provides doctors with rich disease information and feature descriptions, which can help doctors diagnose diseases, formulate treatment plans and evaluate prognosis more accurately. The disease library construction method and system based on big data realize the dynamic update and optimization of the disease library through efficient data processing and analysis technology.

[0083] Through big data technology, automatic identification and classification of diseases can be achieved, the accuracy and efficiency of classification can be improved, and human intervention and errors can be reduced.

[0084] Establish a unified and standard disease classification system, break down data barriers between different medical institutions, promote the sharing and integration of medical data, and provide richer and more comprehensive data support for disease research and clinical decision-making.

[0085] Based on the information provided by the disease database, doctors can formulate diagnosis and treatment plans more scientifically and accurately, thereby improving medical quality and patient satisfaction.

[0086] Through data analysis of the disease database, we can understand the incidence and treatment needs of different diseases, provide a scientific basis for the optimal allocation of medical resources, and improve the efficiency and fairness of medical services.

[0087] The present invention aims to achieve in-depth mining and efficient use of medical data through a disease database construction method based on big data, so as to provide strong support for clinical decision-making, disease research, allocation of medical resources, etc.

[0088] In one embodiment of the present invention, the S1 includes:

[0089] Collect clinical data from each medical system to obtain clinical collection data from each medical system;

[0090] Preprocess each clinical collection data to obtain corresponding clinical processing data;

[0091] Acquire preset feature requirement information, perform feature extraction on clinical processing data of each medical system according to the preset feature requirement information, and obtain system feature extraction data of each medical system;

[0092] The feature extraction data of each system is aggregated to obtain aggregated feature data.

[0093] The working principle of the above technical solution is: clinical data is collected from each medical system. These medical systems may include electronic medical record systems, laboratory information systems, imaging systems, pharmacy management systems, etc., each of which stores various types of information about patients during the medical process, such as diagnostic records, test results, medication status, etc. The collected clinical data is preprocessed to eliminate noise, redundancy and missing values ​​in the data. The preprocessing steps may include data cleaning (such as removing duplicate records and correcting erroneous data), data transformation (such as data standardization and normalization), data reduction (such as feature selection and dimensionality reduction), etc. According to the preset feature requirement information, key features are extracted from the preprocessed clinical data. These features are diagnostic codes, test result values, and types of medication related to specific diseases. The preset feature requirement information is set based on medical knowledge, clinical experience, and data analysis requirements. The system feature data extracted from each medical system is aggregated to form a unified, cross-system aggregated feature data set. The purpose of this step is to integrate data with different features from different sources.

[0094] The technical effects of the above technical scheme are as follows: through data aggregation, the integration of clinical data between different medical systems is realized, information silos are broken, and the availability and value of data are improved. The preprocessing step effectively eliminates noise and redundancy in the data, and improves the quality and accuracy of the data. Feature extraction is performed according to the preset feature requirement information to ensure that the extracted features are targeted and practical, and can meet specific clinical analysis and decision-making needs. The aggregated feature data set facilitates subsequent data analysis, reduces the complexity and time cost of data processing, and improves analysis efficiency. The integrated aggregated feature data set contains rich clinical information, which can provide doctors with a comprehensive patient view, support more accurate clinical decision-making and personalized treatment plan formulation. Through effective data collection, preprocessing, feature extraction and data aggregation steps, the integration and quality improvement of clinical data are achieved.

[0095] In one embodiment of the present invention, S2 includes:

[0096] In the aggregated feature data, data association is performed on the system feature extraction data of each medical system to obtain feature association information;

[0097] Obtaining a preset feature association category, classifying all feature association information according to the preset feature association category, and obtaining category feature association information of multiple categories;

[0098] Generate a medical data view corresponding to a preset feature association category according to each category feature association information;

[0099] Obtain a preset correlation threshold, compare the total number of correlations between each data feature in the medical data view and other features with the preset correlation threshold, and obtain a feature correlation comparison result;

[0100] According to the feature association comparison result, the data features of the category feature association information are screened for feature subsets to obtain multiple feature subsets of each category feature association information. The feature subsets include data feature information whose total number of associations with other features is greater than a preset association threshold.

[0101] The working principle of the above technical solution is as follows: in the aggregated feature data, the system feature extraction data from different medical systems are analyzed for associations between data features. This association can be based on statistical methods, machine learning algorithms or domain knowledge to determine the correlation or mutual dependence between features. According to the preset feature association categories, all feature association information is classified. The purpose of classification is to classify features with similar association characteristics into one category, which is convenient for subsequent data view generation and feature subset screening. For each category feature association information, a corresponding medical data view is generated. These views are presented in the form of graphs, tables or reports, aiming to intuitively show the association relationship between features and the significance of these relationships in a specific medical category. A preset association threshold is obtained, which is used to measure the threshold of the number of associations between features. Then, the total number of associations between each data feature in the medical data view and other features is compared with the preset association threshold. The purpose of this step is to identify feature information whose association strength exceeds the threshold and is therefore of great significance in a specific category. According to the feature association comparison results, the data features in the category feature association information are screened to form feature subsets. These feature subsets include data feature information whose total number of associations with other features is greater than the preset association threshold. These feature subsets represent sets of features with high relevance and importance in different medical categories.

[0102] The technical effect of the above technical solution is: through data feature association and classification, the understanding of the complex relationship between data is enhanced, and the potential medical laws and patterns can be revealed. The generated medical data view provides an intuitive data association display, providing doctors and other medical professionals with a powerful decision support tool. The feature subset screening process ensures that only those features with significant correlation are included, thereby optimizing the feature set and improving the accuracy and efficiency of data analysis. This method can discover new medical knowledge, such as the association between diseases, the effectiveness of treatment methods, etc. By identifying key feature subsets, more personalized medical decisions can be supported. Through the steps of data feature association, classification, view generation, association comparison and feature subset screening, the association information in clinical data is effectively mined and utilized. The above scheme can obtain data features that are more correlated than other features.

[0103] In one embodiment of the present invention, S3 includes:

[0104] The disease database is constructed by aggregating feature data, category feature association information and feature subsets to obtain a category feature disease database;

[0105] Obtain data dynamic information of the category characteristic disease database, based on the data dynamic information;

[0106] Calculate the data dynamic coefficient of the category characteristic disease database based on data dynamic information;

[0107] The calculation formula of the data dynamic coefficient is:

[0108]

[0109] Where DT is the data dynamic coefficient, N is the number of features considered, M is the number of association categories considered, P is the number of other parameters or indicators that may affect the data dynamics (such as time change rate, data source stability, etc.), and w i ,v j ,u k are the weights of the corresponding features, categories and other indicators, which reflect their importance to the dynamic coefficient of the data, ΔF i is the change in the ith feature (which can be the number of new features, the change in feature values, etc.), F ih is the basic quantity of the i-th feature. ΔC j is the change in the jth associated category (which can be the increase or decrease in the number of features within a category, the change in the strength of association between categories, etc.), C jh is the basic quantity of the jth association category, Rel k The relative change of the kth other parameter or indicator that may affect the dynamics of the data;

[0110] Comparing the data dynamic coefficient with a preset dynamic threshold to obtain a data dynamic comparison result;

[0111] Triggering a sorting instruction according to the dynamic comparison result of the data;

[0112] Obtaining inclusion conditions and exclusion conditions according to the inclusion and exclusion instructions;

[0113] The category feature database is included and excluded according to the inclusion conditions and exclusion conditions to obtain inclusion and exclusion information.

[0114] The working principle of the above technical solution is: by aggregating feature data (these feature data may come from multiple medical systems, databases or studies), combined with category feature association information (i.e., the association between different features on specific diseases), a category feature disease library is constructed. This library contains different diseases and their related feature sets, each of which reflects the characteristics of the disease in a specific aspect. The data dynamic information of the category feature disease library is obtained. This information includes newly added feature data, changes in associations between features, and the update frequency of the disease library. This data dynamic information reflects the changes in the disease library over time, that is, the changes in the data inside the disease library, and the dynamic change information of the data after exclusion and inclusion of the disease library; based on the acquired data dynamic information, the data dynamic coefficient of the category feature disease library is calculated. This coefficient is a quantitative indicator used to measure the degree of change of the disease library in a specific time period. The calculated data dynamic coefficient is compared with the preset dynamic threshold. If the data dynamic coefficient exceeds the threshold, it means that the disease library has changed significantly in the recent period, and it is necessary to trigger the exclusion instruction to update the disease library. The exclusion instruction is an operation instruction used to guide how to include and exclude the disease library based on the new data dynamic information. According to the triggered inclusion and exclusion instructions, specific inclusion and exclusion conditions are obtained. These conditions may be set based on new feature data, feature associations, or the update requirements of the disease database. Then, according to these conditions, the category feature database is included and excluded to obtain updated inclusion and exclusion information. This information reflects the latest status of the disease database after the inclusion of new features and the exclusion of old features.

[0115] In the above calculation formula, the feature change part reflects the degree of change in the quantity or value of each feature. These changes may be directly related to the update and dynamics of the disease library content. Medium i is the feature weight, reflecting the importance of the feature. i is the feature change, which can be the number of new features, feature value change, etc. ih It is the basic quantity or initial quantity of the feature, which is used to standardize the change quantity. The result is the sum of weighted standardized feature changes, which reflects the dynamics of the feature level. The category or associated category change part reflects the increase or decrease in the number of features in different categories or associated categories, the change in the strength of association between categories, etc. These changes are crucial to understanding the overall structure and dynamics of the disease library. Medium j is the weight of the category or associated category, reflecting its importance. j is the change in category or associated category. jhIt is the basic quantity or initial quantity of the category or associated category. The result is the sum of weighted standardized category changes, reflecting the dynamics at the category level. Other influencing parameters consider other factors that may affect the dynamics of data besides feature and category changes, such as time change rate, data source stability, etc. middleu k is the weight of other influencing parameters, Rel k It is a relative change, reflecting the degree of change of the parameter. The result is a weighted sum of changes in other influencing parameters, reflecting the impact of other factors on data dynamics. The comprehensive calculation part of the data dynamic coefficient normalizes the results of the above three parts to obtain a coefficient that comprehensively reflects the data dynamics. The result is a normalized data dynamic coefficient, which comprehensively considers the influence of features, categories and other factors, and is used to quantify data dynamics. By dividing it into multiple calculation parts, the physical meaning and calculation effect of the data dynamic coefficient (DDC) calculation formula can be more clearly understood. Each part reflects different aspects of data dynamics, and finally a comprehensive quantitative indicator is obtained.

[0116] The technical effect of the above technical solution is: by acquiring data dynamic information in real time and calculating data dynamic coefficients, the changing trend of the disease library can be discovered in time, thereby triggering the inclusion and exclusion instructions for updating. This can ensure the timeliness and accuracy of the disease library. Through inclusion and exclusion operations, the feature set in the disease library can be optimized to make it more in line with actual clinical needs. This can help medical institutions allocate medical resources more reasonably, improve diagnosis and treatment efficiency and patient satisfaction. The construction and updating of the category feature disease library provides valuable data resources for medical research. Through in-depth analysis of these data, new medical laws can be discovered, and new treatment methods and strategies can be proposed. With the continuous improvement and updating of the disease library, the disease characteristics and needs of different patients can be more accurately identified. This provides strong support for personalized medicine and can provide patients with more accurate and effective treatment plans. By aggregating feature data, category feature association information and feature subsets to construct a disease library, and updating and optimizing it according to data dynamic information, the timeliness and accuracy of the disease library can be significantly improved, and the allocation of medical resources can be optimized.

[0117] In one embodiment of the present invention, the S4 includes:

[0118] The inclusion and exclusion information are compared with the preset inclusion and exclusion thresholds, respectively;

[0119] When the inclusion and exclusion information is greater than the preset inclusion threshold or exclusion threshold, a database update instruction is triggered;

[0120] When the included information is greater than the preset inclusion threshold, the database incremental update instruction is triggered;

[0121] When the exclusion information is greater than the preset exclusion threshold, a database reduction update instruction is triggered.

[0122] The working principle of the above technical solution is as follows: the preset inclusion threshold and exclusion threshold are standard values ​​set according to business needs, data characteristics or expert experience. These thresholds are used to measure the significance or importance of the inclusion and exclusion information. The inclusion and exclusion information is compared with the preset inclusion threshold and exclusion threshold. The purpose of this step is to determine whether the inclusion and exclusion operations have reached the preset significance level, so as to decide whether to trigger the database update instruction. When the inclusion information is greater than the preset inclusion threshold or the exclusion information is greater than the preset exclusion threshold, it means that the inclusion or exclusion operation has a significant impact on the database and the database update instruction needs to be triggered. Different database update instructions are triggered according to the specific circumstances of the inclusion and exclusion information. When the inclusion information is greater than the preset inclusion threshold, the database incremental update instruction is triggered, that is, new data or features are added to the database; when the exclusion information is greater than the preset exclusion threshold, the database decrement update instruction is triggered, that is, data or features that no longer meet the conditions are deleted or removed from the database.

[0123] The steps to sort out the parts include:

[0124] The information of the disease databases created and participated by users is displayed in the form of cards, including: the name of the disease database, the number of patients, the number of scientific research projects created with the disease database as the data source, the creator, and the creation time. When the mouse moves into the corresponding disease database card, the operation buttons "Rename, Manage Team, Set, Delete" are displayed. Click the card to enter the disease database details page. You can enter the disease database name to search.

[0125] A. Disease database settings for data source "all data":

[0126] a. Enter the basic information of the disease database: enter the name of the disease database and select the data source "All Data".

[0127] b. Enter inclusion and exclusion criteria

[0128] Currently, the searchable sorting conditions are all enabled data table fields configured in [Backend Management System] - [Source Data Management]. The display format is "①Data Table Cascade Selection Box + ②Field Selection Box + ③Condition Selection Box + ④User Input Box".

[0129] Since different field selections correspond to different condition selection boxes, you need to select the search field first, and then select the search condition. When the data type of the selected field configured in [Backend Management System] - [Source Data Management] is "Text", the condition selection box has "Contains, Equals, Starts with..., Ends with..." options. If the selected field is configured with a value domain in [Backend Management System] - [Source Data Management], then ④ the user input box is an input box with input suggestions; when the data type is "Number", the condition selection box has "Between, Equals, Not Equal, Greater Than, Greater Than or Equal, Less Than and Less Than or Equal" options; when the data type is "Date", the condition selection box has "Equals, Earlier Than, Later Than, Between" options.

[0130] Inclusion conditions: Initially, users can "add" condition groups or combination medication conditions, where a condition group consists of single or multiple field conditions in the same data table. Each time a new field condition or condition group, or combination medication condition is added, a selection box will automatically appear above the newly added condition / group: "AND" or "OR;

[0131] When the logical relationship between two condition groups is "AND", adding event relationship conditions is optional.

[0132] B. Create: Click "Create" to enter the disease database details page. Once the disease database is successfully created and there are cases, the inclusion and exclusion conditions cannot be changed.

[0133] 3.1.3 Details of the disease database

[0134] It is divided into data navigation area and data content display area.

[0135] 3.1.4 Disease database case management

[0136] Cases can be added, removed, and searched.

[0137] 1) Add case: Enter the patient add interface, filter by patient name / first letter / medical card number / medical type / registration date to form a list of candidate patients, and check the patients in the list to add them to the disease database.

[0138] If a case exists in the removed cases of this disease database, it will be prompted when checking:

[0139] If the patient has been added to the disease database, the status is checked and cannot be cancelled:

[0140] If the selected case is in the "pending confirmation" list, you can directly select it. After importing, it will be added to the disease library, and the pending confirmation list needs to be updated at the same time.

[0141] When the patient cannot be found, you can select "Manually import patients" to add patients by manually filling in the patient number (required), name (required), gender, age, and institution. When performing the data scanning task on a daily basis, if the patient number filled in manually matches the patient's empiseqid in RDR, the patient data in RDR will be used to overwrite the manually filled in data.

[0142] 2) Batch removal: Check the cases to be removed in the table and click "Batch removal". After the cases are removed, the corresponding cases in the research projects that use this disease database as the data source will be removed simultaneously. The removed cases will no longer participate in data collection, but the collected data will still be retained in the disease database and research projects.

[0143] The technical effect of the above technical solution is: by setting a threshold and comparing the inclusion and exclusion information, it is possible to efficiently determine whether the database needs to be updated, avoiding unnecessary update operations and improving the efficiency of database updates. The triggering of the incremental and decremental update instructions is based on strict threshold comparison, ensuring that only data that meets the significance requirements will be included or excluded, thereby maintaining the data quality of the database. Triggering different update instructions (increment or decrement) can flexibly adjust the content of the database according to actual conditions, avoiding waste of resources and optimizing resource utilization. The method supports dynamic management of database content, enabling it to be continuously updated and optimized as business needs and data characteristics change, providing more accurate and timely data support for business decisions. By setting clear inclusion and exclusion criteria and corresponding update instructions, a standardized data governance process can be established to improve the transparency and traceability of data management. By comparing the inclusion and exclusion information with the preset threshold and triggering different database update instructions based on the comparison results, the database content can be effectively managed, data quality can be improved, resource utilization can be optimized, and the development of dynamic data management and data governance can be supported.

[0144] In one embodiment of the present invention, the system comprises:

[0145] A data feature aggregation module is used to collect and preprocess clinical data from different medical systems, obtain clinical processing data from multiple medical systems, perform feature extraction and data aggregation on the clinical processing data, and obtain aggregated feature data;

[0146] A category data association extraction module is used to perform data association on the aggregated feature data and classify it according to the preset feature association categories, obtain category feature association information, generate a medical data view of each preset feature association category, and then obtain a feature subset;

[0147] The database construction processing module is used to construct a category characteristic disease database, obtain data dynamic information, calculate data dynamic coefficients, analyze data dynamic coefficients, and then trigger exclusion instructions to include and exclude the category characteristic database and obtain inclusion and exclusion information;

[0148] The regional update module is used to compare the inclusion and exclusion information with the preset inclusion threshold and exclusion threshold respectively, trigger the database update instruction according to the comparison result, and then perform the database incremental update or database decrement update.

[0149] The working principle of the above technical solution is as follows: clinical data is collected from multiple medical systems, which may include multiple electronic medical record systems, imaging systems, laboratory information systems, etc. of multiple hospitals. The collected data is cleaned, formatted, and missing values ​​are processed to ensure data quality and consistency. Key features such as patient age, gender, diagnosis information, treatment process, etc. are extracted from the preprocessed data, and these features are aggregated to form aggregated feature data. The relationship between different features in the aggregated feature data is analyzed, such as the relationship between disease and symptoms, and between disease and treatment methods. The associated information is classified according to the preset feature association categories (such as disease type, treatment method classification, etc.), and medical data views are generated, which help understand and analyze data features. Specific feature subsets are extracted from the classified data for subsequent construction of the disease library. The initial disease library is constructed based on the feature subset, including feature descriptions of different diseases. New clinical data is continuously monitored and acquired, and the data dynamic coefficient is calculated, which reflects the update frequency and degree of change of the data. The data dynamic coefficient is analyzed, and when the data change reaches the preset conditions, the inclusion or exclusion instructions are triggered to update the disease library. The inclusion and exclusion information is compared with the preset inclusion and exclusion thresholds. Based on the comparison results, if the data meets the inclusion criteria, the database is incrementally updated (adding new data or disease types); if the exclusion criteria are met, the database is decrementally updated (deleting data or disease types that no longer meet the criteria).

[0150] The technical effects of the above technical scheme are as follows: through data preprocessing and feature extraction, the accuracy and consistency of the data used to construct the disease library are ensured. Data feature association and classification can reveal the potential relationship between different features and provide more comprehensive information for the construction of the disease library. By continuously monitoring data dynamics and triggering exclusion instructions, the disease library can reflect the latest clinical data and disease characteristics in real time, improving the timeliness and accuracy of the disease library. The construction of a disease library based on big data can provide disease distribution data and patient demand data, thereby optimizing resource allocation and improving medical service efficiency. The disease library provides doctors with rich disease information and feature descriptions, which can help doctors diagnose diseases, formulate treatment plans and evaluate prognosis more accurately. The disease library construction method and system based on big data realize the dynamic update and optimization of the disease library through efficient data processing and analysis technology.

[0151] In one embodiment of the present invention, the data feature aggregation module includes:

[0152] Collect clinical data from each medical system to obtain clinical collection data from each medical system;

[0153] Preprocess each clinical collection data to obtain corresponding clinical processing data;

[0154] Acquire preset feature requirement information, perform feature extraction on clinical processing data of each medical system according to the preset feature requirement information, and obtain system feature extraction data of each medical system;

[0155] The feature extraction data of each system is aggregated to obtain aggregated feature data.

[0156] The working principle of the above technical solution is: clinical data is collected from each medical system. These medical systems may include electronic medical record systems, laboratory information systems, imaging systems, pharmacy management systems, etc., each of which stores various types of information about patients during the medical process, such as diagnostic records, test results, medication status, etc. The collected clinical data is preprocessed to eliminate noise, redundancy and missing values ​​in the data. The preprocessing steps may include data cleaning (such as removing duplicate records and correcting erroneous data), data transformation (such as data standardization and normalization), data reduction (such as feature selection and dimensionality reduction), etc. According to the preset feature requirement information, key features are extracted from the preprocessed clinical data. These features are diagnostic codes, test result values, and types of medication related to specific diseases. The preset feature requirement information is set based on medical knowledge, clinical experience, and data analysis requirements. The system feature data extracted from each medical system is aggregated to form a unified, cross-system aggregated feature data set. The purpose of this step is to integrate data with different features from different sources.

[0157] The technical effects of the above technical scheme are as follows: through data aggregation, the integration of clinical data between different medical systems is realized, information silos are broken, and the availability and value of data are improved. The preprocessing step effectively eliminates noise and redundancy in the data, and improves the quality and accuracy of the data. Feature extraction is performed according to the preset feature requirement information to ensure that the extracted features are targeted and practical, and can meet specific clinical analysis and decision-making needs. The aggregated feature data set facilitates subsequent data analysis, reduces the complexity and time cost of data processing, and improves analysis efficiency. The integrated aggregated feature data set contains rich clinical information, which can provide doctors with a comprehensive patient view, support more accurate clinical decision-making and personalized treatment plan formulation. Through effective data collection, preprocessing, feature extraction and data aggregation steps, the integration and quality improvement of clinical data are achieved.

[0158] In one embodiment of the present invention, the category data association extraction module includes:

[0159] In the aggregated feature data, data association is performed on the system feature extraction data of each medical system to obtain feature association information;

[0160] Obtaining a preset feature association category, classifying all feature association information according to the preset feature association category, and obtaining category feature association information of multiple categories;

[0161] Generate a medical data view corresponding to a preset feature association category according to each category feature association information;

[0162] Obtain a preset correlation threshold, compare the total number of correlations between each data feature in the medical data view and other features with the preset correlation threshold, and obtain a feature correlation comparison result;

[0163] According to the feature association comparison result, the data features of the category feature association information are screened for feature subsets to obtain multiple feature subsets of each category feature association information. The feature subsets include data feature information whose total number of associations with other features is greater than a preset association threshold.

[0164] The working principle of the above technical solution is as follows: in the aggregated feature data, the system feature extraction data from different medical systems are analyzed for associations between data features. This association can be based on statistical methods, machine learning algorithms or domain knowledge to determine the correlation or mutual dependence between features. According to the preset feature association categories, all feature association information is classified. The purpose of classification is to classify features with similar association characteristics into one category, which is convenient for subsequent data view generation and feature subset screening. For each category feature association information, a corresponding medical data view is generated. These views are presented in the form of graphs, tables or reports, aiming to intuitively show the association relationship between features and the significance of these relationships in a specific medical category. A preset association threshold is obtained, which is used to measure the threshold of the number of associations between features. Then, the total number of associations between each data feature in the medical data view and other features is compared with the preset association threshold. The purpose of this step is to identify feature information whose association strength exceeds the threshold and is therefore of great significance in a specific category. According to the feature association comparison results, the data features in the category feature association information are screened to form feature subsets. These feature subsets include data feature information whose total number of associations with other features is greater than the preset association threshold. These feature subsets represent sets of features with high relevance and importance in different medical categories.

[0165] The technical effect of the above technical solution is: through data feature association and classification, the understanding of the complex relationship between data is enhanced, and the potential medical laws and patterns can be revealed. The generated medical data view provides an intuitive display of data association, providing a powerful decision support tool for doctors and other medical professionals. The feature subset screening process ensures that only those features with significant associations are included, thereby optimizing the feature set and improving the accuracy and efficiency of data analysis. This method can discover new medical knowledge, such as the association between diseases, the effectiveness of treatment methods, etc. By identifying key feature subsets, more personalized medical decisions can be supported. Through the steps of data feature association, classification, view generation, association comparison and feature subset screening, the association information in clinical data is effectively mined and utilized.

[0166] In one embodiment of the present invention, the database construction processing module includes:

[0167] The disease database is constructed by aggregating feature data, category feature association information and feature subsets to obtain a category feature disease database;

[0168] Obtain data dynamic information of the category characteristic disease database, based on the data dynamic information;

[0169] Calculate the data dynamic coefficient of the category characteristic disease database based on data dynamic information;

[0170] The calculation formula of the data dynamic coefficient is:

[0171]

[0172] Where DT is the data dynamic coefficient, N is the number of features considered, M is the number of association categories considered, P is the number of other parameters or indicators that may affect the data dynamics (such as time change rate, data source stability, etc.), and w i ,v j ,u k are the weights of the corresponding features, categories and other indicators, which reflect their importance to the dynamic coefficient of the data, ΔF i is the change in the ith feature (which can be the number of new features, the change in feature values, etc.), F ih is the basic quantity of the i-th feature. ΔC j is the change in the jth associated category (which can be the increase or decrease in the number of features within a category, the change in the strength of association between categories, etc.), C jh is the basic quantity of the jth associated category, Relk is the relative change of the kth other parameter or indicator that may affect the dynamics of the data;

[0173] Comparing the data dynamic coefficient with a preset dynamic threshold to obtain a data dynamic comparison result;

[0174] Triggering a sorting instruction according to the dynamic comparison result of the data;

[0175] Obtaining inclusion conditions and exclusion conditions according to the inclusion and exclusion instructions;

[0176] The category feature database is included and excluded according to the inclusion conditions and exclusion conditions to obtain inclusion and exclusion information.

[0177] The working principle of the above technical solution is: by aggregating feature data (these feature data may come from multiple medical systems, databases or studies), combined with category feature association information (i.e., the association between different features on specific diseases), a category feature disease library is constructed. This library contains different diseases and their related feature sets, each of which reflects the characteristics of the disease in a specific aspect. The data dynamic information of the category feature disease library is obtained. This information includes newly added feature data, changes in associations between features, and the update frequency of the disease library. This data dynamic information reflects the changes in the disease library over time, that is, the changes in the data inside the disease library, and the dynamic change information of the data after exclusion and inclusion of the disease library; based on the acquired data dynamic information, the data dynamic coefficient of the category feature disease library is calculated. This coefficient is a quantitative indicator used to measure the degree of change of the disease library in a specific time period. The calculated data dynamic coefficient is compared with the preset dynamic threshold. If the data dynamic coefficient exceeds the threshold, it means that the disease library has changed significantly in the recent period, and it is necessary to trigger the exclusion instruction to update the disease library. The exclusion instruction is an operation instruction used to guide how to include and exclude the disease library based on the new data dynamic information. According to the triggered inclusion and exclusion instructions, specific inclusion and exclusion conditions are obtained. These conditions may be set based on new feature data, feature associations, or the update requirements of the disease database. Then, according to these conditions, the category feature database is included and excluded to obtain updated inclusion and exclusion information. This information reflects the latest status of the disease database after the inclusion of new features and the exclusion of old features.

[0178] The technical effect of the above technical solution is: by acquiring data dynamic information in real time and calculating data dynamic coefficients, the changing trend of the disease library can be discovered in time, thereby triggering the inclusion and exclusion instructions for updating. This can ensure the timeliness and accuracy of the disease library. Through inclusion and exclusion operations, the feature set in the disease library can be optimized to make it more in line with actual clinical needs. This can help medical institutions allocate medical resources more reasonably, improve diagnosis and treatment efficiency and patient satisfaction. The construction and updating of the category feature disease library provides valuable data resources for medical research. Through in-depth analysis of these data, new medical laws can be discovered, and new treatment methods and strategies can be proposed. With the continuous improvement and updating of the disease library, the disease characteristics and needs of different patients can be more accurately identified. This provides strong support for personalized medicine and can provide patients with more accurate and effective treatment plans. By aggregating feature data, category feature association information and feature subsets to construct a disease library, and updating and optimizing it according to data dynamic information, the timeliness and accuracy of the disease library can be significantly improved, and the allocation of medical resources can be optimized.

[0179] In one embodiment of the present invention, the area update module includes:

[0180] The inclusion and exclusion information are compared with the preset inclusion and exclusion thresholds, respectively;

[0181] When the inclusion and exclusion information is greater than the preset inclusion threshold or exclusion threshold, a database update instruction is triggered;

[0182] When the included information is greater than the preset inclusion threshold, the database incremental update instruction is triggered;

[0183] When the exclusion information is greater than the preset exclusion threshold, a database reduction update instruction is triggered.

[0184] The working principle of the above technical solution is as follows: the preset inclusion threshold and exclusion threshold are standard values ​​set according to business needs, data characteristics or expert experience. These thresholds are used to measure the significance or importance of the inclusion and exclusion information. The inclusion and exclusion information is compared with the preset inclusion threshold and exclusion threshold. The purpose of this step is to determine whether the inclusion and exclusion operations have reached the preset significance level, so as to decide whether to trigger the database update instruction. When the inclusion information is greater than the preset inclusion threshold or the exclusion information is greater than the preset exclusion threshold, it means that the inclusion or exclusion operation has a significant impact on the database and the database update instruction needs to be triggered. Different database update instructions are triggered according to the specific circumstances of the inclusion and exclusion information. When the inclusion information is greater than the preset inclusion threshold, the database incremental update instruction is triggered, that is, new data or features are added to the database; when the exclusion information is greater than the preset exclusion threshold, the database decrement update instruction is triggered, that is, data or features that no longer meet the conditions are deleted or removed from the database.

[0185] The technical effect of the above technical solution is: by setting a threshold and comparing the inclusion and exclusion information, it is possible to efficiently determine whether the database needs to be updated, avoiding unnecessary update operations and improving the efficiency of database updates. The triggering of the incremental and decremental update instructions is based on strict threshold comparison, ensuring that only data that meets the significance requirements will be included or excluded, thereby maintaining the data quality of the database. Triggering different update instructions (increment or decrement) can flexibly adjust the content of the database according to actual conditions, avoiding waste of resources and optimizing resource utilization. The method supports dynamic management of database content, enabling it to be continuously updated and optimized as business needs and data characteristics change, providing more accurate and timely data support for business decisions. By setting clear inclusion and exclusion criteria and corresponding update instructions, a standardized data governance process can be established to improve the transparency and traceability of data management. By comparing the inclusion and exclusion information with the preset threshold and triggering different database update instructions based on the comparison results, the database content can be effectively managed, data quality can be improved, resource utilization can be optimized, and the development of dynamic data management and data governance can be supported.

[0186] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A method for constructing a disease database based on big data, characterized in that: The method comprises: S1. Collect and preprocess clinical data from different medical systems to obtain clinical processing data from multiple medical systems, perform feature extraction and data aggregation on the clinical processing data to obtain aggregated feature data; S2, performing data association of data features on the aggregated feature data and classifying them according to preset feature association categories, obtaining category feature association information, generating a medical data view of each preset feature association category, and then obtaining a feature subset; S3. Build a category characteristic disease database, obtain data dynamic information, calculate data dynamic coefficients, analyze the data dynamic coefficients, and then trigger exclusion instructions to include and exclude the category characteristic database to obtain inclusion and exclusion information; S4. Compare the inclusion and exclusion information with the preset inclusion threshold and exclusion threshold respectively, trigger a database update instruction according to the comparison result, and then perform a database incremental update or a database decrement update.

2. According to the method for constructing a disease database based on big data as described in claim 1, it is characterized in that: The S1 includes: Collect clinical data from each medical system to obtain clinical collection data from each medical system; Preprocess each clinical collection data to obtain corresponding clinical processing data; Acquire preset feature requirement information, perform feature extraction on clinical processing data of each medical system according to the preset feature requirement information, and obtain system feature extraction data of each medical system; The feature extraction data of each system is aggregated to obtain aggregated feature data.

3. According to the method for constructing a disease database based on big data as described in claim 1, it is characterized in that: The S2 includes: In the aggregated feature data, data association is performed on the system feature extraction data of each medical system to obtain feature association information; Obtaining a preset feature association category, classifying all feature association information according to the preset feature association category, and obtaining category feature association information of multiple categories; Generate a medical data view corresponding to a preset feature association category according to each category feature association information; Obtain a preset correlation threshold, compare the total number of correlations between each data feature in the medical data view and other features with the preset correlation threshold, and obtain a feature correlation comparison result; The feature subsets of the data features of the category feature association information are screened according to the feature association comparison result to obtain multiple feature subsets of each category feature association information.

4. According to the method for constructing a disease database based on big data as described in claim 1, it is characterized in that: The S3 includes: The disease database is constructed by aggregating feature data, category feature association information and feature subsets to obtain a category feature disease database; Obtain data dynamic information of the category characteristic disease database, based on the data dynamic information; Calculate the data dynamic coefficient of the category characteristic disease database based on data dynamic information; Comparing the data dynamic coefficient with a preset dynamic threshold to obtain a data dynamic comparison result; Triggering a sorting instruction according to the dynamic comparison result of the data; Obtaining inclusion conditions and exclusion conditions according to the inclusion and exclusion instructions; The category feature database is included and excluded according to the inclusion conditions and exclusion conditions to obtain inclusion and exclusion information.

5. According to the method for constructing a disease database based on big data as described in claim 1, it is characterized in that: The S4 includes: The inclusion and exclusion information are compared with the preset inclusion and exclusion thresholds, respectively; When the inclusion and exclusion information is greater than the preset inclusion threshold or exclusion threshold, a database update instruction is triggered; When the included information is greater than the preset inclusion threshold, the database incremental update instruction is triggered; When the exclusion information is greater than the preset exclusion threshold, a database reduction update instruction is triggered.

6. A disease database construction system based on big data, characterized in that: The system comprises: A data feature aggregation module is used to collect and preprocess clinical data from different medical systems, obtain clinical processing data from multiple medical systems, perform feature extraction and data aggregation on the clinical processing data, and obtain aggregated feature data; A category data association extraction module is used to perform data association on the aggregated feature data and classify it according to the preset feature association categories, obtain category feature association information, generate a medical data view of each preset feature association category, and then obtain a feature subset; The database construction processing module is used to construct a category characteristic disease database, obtain data dynamic information, calculate data dynamic coefficients, analyze data dynamic coefficients, and then trigger exclusion instructions to include and exclude the category characteristic database and obtain inclusion and exclusion information; The regional update module is used to compare the inclusion and exclusion information with the preset inclusion threshold and exclusion threshold respectively, trigger the database update instruction according to the comparison result, and then perform the database incremental update or database decrement update.

7. According to the big data-based disease database construction system of claim 6, it is characterized in that: The data feature aggregation module includes: Collect clinical data from each medical system to obtain clinical collection data from each medical system; Preprocess each clinical collection data to obtain corresponding clinical processing data; Acquire preset feature requirement information, perform feature extraction on clinical processing data of each medical system according to the preset feature requirement information, and obtain system feature extraction data of each medical system; The feature extraction data of each system is aggregated to obtain aggregated feature data.

8. According to the big data-based disease database construction system of claim 6, it is characterized in that: The category data association extraction module includes: In the aggregated feature data, data association is performed on the system feature extraction data of each medical system to obtain feature association information; Obtaining a preset feature association category, classifying all feature association information according to the preset feature association category, and obtaining category feature association information of multiple categories; Generate a medical data view corresponding to a preset feature association category according to each category feature association information; Obtain a preset correlation threshold, compare the total number of correlations between each data feature in the medical data view and other features with the preset correlation threshold, and obtain a feature correlation comparison result; The feature subsets of the data features of the category feature association information are screened according to the feature association comparison result to obtain multiple feature subsets of each category feature association information.

9. According to the big data-based disease database construction system of claim 6, it is characterized in that: The database construction processing module includes: The disease database is constructed by aggregating feature data, category feature association information and feature subsets to obtain a category feature disease database; Obtain data dynamic information of the category characteristic disease database, based on the data dynamic information; Calculate the data dynamic coefficient of the category characteristic disease database based on data dynamic information; Comparing the data dynamic coefficient with a preset dynamic threshold to obtain a data dynamic comparison result; Triggering a sorting instruction according to the dynamic comparison result of the data; Obtaining inclusion conditions and exclusion conditions according to the inclusion and exclusion instructions; The category feature database is included and excluded according to the inclusion conditions and exclusion conditions to obtain inclusion and exclusion information.

10. A disease database construction system based on big data according to claim 6, characterized in that: The area update module includes: The inclusion and exclusion information are compared with the preset inclusion and exclusion thresholds, respectively; When the inclusion and exclusion information is greater than the preset inclusion threshold or exclusion threshold, a database update instruction is triggered; When the included information is greater than the preset inclusion threshold, the database incremental update instruction is triggered; When the exclusion information is greater than the preset exclusion threshold, a database reduction update instruction is triggered.