Special disease medical data queue collection method, system and equipment based on human body index set and storage medium
By constructing a hierarchical set of human indicators and a four-level structured framework, the problems of inconsistent terminology and incomplete temporal characteristics in medical data integration were solved. This enabled automatic conversion and precise correlation of multi-source data, improved data management efficiency and accuracy, and supported high-quality data applications in clinical practice and scientific research.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGXI MEDICAL UNIVERSITY
- Filing Date
- 2025-12-23
- Publication Date
- 2026-04-10
AI Technical Summary
Existing medical data integration technologies lack a unified data standardization system, resulting in the need for manual work to align terms and convert units between different data sources, which is inefficient and difficult to guarantee accuracy. Insufficient attention is paid to the temporal characteristics of medical data, making it impossible to completely preserve historical data, which affects efficacy evaluation and treatment plan adjustment. Data collection lacks connection with clinical business processes, resulting in mixed data storage and cumbersome operations.
A hierarchical set of human body indicators is constructed and a two-way mapping table is established. Data standardization and time-series management are achieved through a four-level structured framework and time-series labels. Terminology alignment and unit conversion are automatically completed to ensure the accurate association between data and business processes. Historical data storage areas are set up in each link to completely save multi-time-series indicator data.
It achieves efficient and accurate conversion of multi-source heterogeneous data, ensures close integration of data with business processes, avoids mixed data storage, improves data integrity and utilization, and provides high-quality data support for clinical decision-making and scientific research analysis.
Smart Images

Figure CN121838985A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical and health data processing technology, and in particular to a method, system, device and storage medium for collecting disease-specific medical data queues based on human indicator sets. Background Technology
[0002] In clinical medical and scientific research practices, the integration of medical data directly affects the accuracy of diagnostic and treatment decisions and the reliability of research results. Medical institutions accumulate a large amount of patient data during daily diagnosis and treatment. This data is scattered and stored in different systems such as Hospital Information System (HIS), Laboratory Information System (LIS), and Picture Archiving and Communication System (PACS). However, different systems use their own corresponding data formats, terminology standards, and storage structures. When the same test indicator is described differently in different systems, for example, blood glucose is called "serum glucose" in some systems and "FBG" in others, and the units of measurement also differ between mg / dL and mmol / L, such inconsistent data formats can easily lead to a large amount of manual comparison and format conversion when integrating data across systems. This reduces integration efficiency and is also prone to data errors due to manual comparison and format conversion.
[0003] Current medical data integration technologies suffer from three main problems. First, there is a lack of a unified data standardization system. Terminology alignment and unit conversion between different data sources rely on manual work, resulting in low data exchange efficiency and difficulty in ensuring accuracy. Second, insufficient attention is paid to the temporal characteristics of medical data. Often, only the latest measurements of patient indicators are retained, while historical data is overwritten. This prevents doctors from tracking disease progression and affects efficacy assessment and treatment plan adjustments. For example, diabetic patients need multiple blood glucose monitoring data to determine treatment effectiveness, but existing systems cannot fully save and manage these historical measurement records. Third, data collection lacks integration with clinical workflows. Data from different diagnostic and treatment stages is stored in a mixed manner. Researchers need to manually filter data from specific stages when constructing disease data cohorts, a cumbersome process that easily leads to the omission of key information and incomplete data cohorts. Furthermore, when multiple data sources record inconsistent values for the same indicator for the same patient at the same time, current technologies lack conflict identification and handling mechanisms, relying solely on manual verification, which further increases the complexity of data management. Summary of the Invention
[0004] In view of the problems existing in the prior art, the present invention is proposed.
[0005] Therefore, the problems to be solved by this invention are how to construct a standardized indicator system to achieve automatic conversion of multi-source data, how to design a reasonable data structure to completely preserve time-series data and establish a precise correlation between data and business processes, and how to ensure data quality and consistency. These are technical problems that urgently need to be solved in the field of medical data integration.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, embodiments of the present invention provide a method for collecting disease-specific medical data queues based on human body indicator sets, which includes: constructing a hierarchical human body indicator set, receiving raw medical data from multiple data sources, establishing a mapping relationship between raw data fields and standardized indicators in the human body indicator set through a bidirectional mapping table, and converting the raw medical data into standardized indicator data. A four-level structured framework including lines, links, actions, and indicators is constructed. Time-series tags are added to the standardized indicator data. The time-series tags record the business affiliation information and time information of the data. Based on the time information in the time-series tags and the time period attributes, the target link and target action to which the standardized indicator data belongs are determined through a time association algorithm. The standardized indicator data is collected into the storage location corresponding to the target link and the target action, and a historical data action storage area is set in each link to store multi-time series indicator data of the action in that link. By integrating basic patient information, standardized indicator data of each step, and multi-time-series data from historical data actions, a disease-specific data queue is constructed.
[0007] As a preferred embodiment of the method for collecting disease-specific medical data queues based on human body indicator sets as described in this invention, the human body indicator set is organized according to a three-level hierarchical relationship of human body systems, organs and indicators, and each standardized indicator includes indicator code, standard name, synonym list, data type, unit of measurement and reference value range. The step of establishing a mapping relationship between the original data fields and the standardized indicators in the human body indicator set through a bidirectional mapping table includes: for the original indicator names that fail to be mapped through the bidirectional mapping table, extracting the original indicator names and counting the number of mapping failures; when the number of mapping failures exceeds a preset threshold, performing semantic association calculation and medical rule matching between the original indicator names and the names of each standardized indicator in the human body indicator set; when the comprehensive confidence meets the preset conditions, adding the original indicator name to the synonym list of the corresponding standardized indicator and updating the bidirectional mapping table.
[0008] The beneficial effects of this preferred technical solution are as follows: By constructing a three-level hierarchical human body indicator set, each standardized indicator is equipped with complete attributes such as indicator code, standard name, and synonym list; by statistically analyzing the number of mapping failures and setting trigger thresholds, frequently occurring unknown terms can be automatically identified; the original indicator name is semantically correlated with each standardized indicator in the human body indicator set and matched with medical rules, and synonyms are automatically added to the indicator list and the bidirectional mapping table is updated based on comprehensive confidence level; the processing method reduces the reliance on manual intervention during data conversion, reduces the data loss rate caused by inconsistent terminology, improves the conversion efficiency and accuracy of multi-source heterogeneous data, and can adapt to the differences in terminology habits among different medical institutions.
[0009] As a preferred embodiment of the method for collecting disease-specific medical data queues based on human body indicator sets according to the present invention, the step of converting raw medical data into standardized indicator data includes parsing the fields of the raw medical data and extracting the indicator name, measurement value, measurement time and unit of measurement. The standard indicator code corresponding to the indicator name is retrieved using the bidirectional mapping table; Obtain the standard unit of measurement according to the standard index code, determine whether the unit of measurement of the measured value is consistent with the standard unit of measurement, and if they are inconsistent, perform unit conversion according to the preset conversion formula. The converted measurement values are associated with the standard indicator codes to generate standardized indicator data.
[0010] As a preferred embodiment of the disease-specific medical data queue collection method based on human body indicator sets described in this invention, wherein: the lines correspond to disease diagnosis and treatment paths, the links correspond to diagnosis and treatment stages and have start and end times, the actions correspond to specific clinical operations and are associated with one or more standardized indicators; the time sequence labels record patient identifier, line identifier, link identifier, action identifier, data collection time, indicator code, data source and data version number; The time correlation algorithm used to determine the target stage and target action to which standardized indicator data belongs includes: obtaining the collection time of the standardized indicator data, obtaining the start time and end time of each stage in the line, determining whether the collection time falls within the time range of a certain stage, and when the collection time falls within the time range of multiple stages, calculating the belonging priority score of each stage based on the time overlap ratio, stage priority coefficient and indicator correlation strength, and aggregating the standardized indicator data to the stage with the highest belonging priority score.
[0011] The beneficial effects of this preferred technical solution are as follows: By constructing a four-level structured framework of lines, links, actions, and indicators, the disease diagnosis and treatment pathway, treatment stage, specific clinical operation, and standardized indicators are linked layer by layer, making data collection closely integrated with business processes; Time-series labels containing multi-dimensional information such as patient identifier, line identifier, link identifier, action identifier, and data collection time are added to the standardized indicator data to establish a precise correspondence between data and business processes; By obtaining the collection time of the standardized indicator data and matching it with the start and end times of each link, the time attribution of the data is determined. When the collection time falls into multiple links simultaneously, the attribution priority score is calculated based on the time overlap ratio, link priority coefficient, and indicator association strength, and the data is collected to the link with the highest score; The dynamic association algorithm ensures that each indicator data can be accurately mapped to a specific diagnosis and treatment link and clinical operation, avoiding business logic confusion caused by mixed data storage.
[0012] As a preferred embodiment of the method for collecting disease-specific medical data queues based on human body indicator sets as described in this invention, the priority coefficient of each link is set according to the link type, with the priority coefficient of treatment links being higher than that of diagnosis links, and the priority coefficient of diagnosis links being higher than that of screening links. The time overlap ratio is the proportion of the overlap between the data acquisition time and the time range of the process to the total time of the process. The correlation strength of the indicator represents the degree of correlation between the indicator and the actions within the process; the higher the degree of correlation, the greater the correlation strength of the indicator.
[0013] As a preferred embodiment of the method for collecting disease-specific medical data queues based on human body indicator sets described in this invention, the step of setting a historical data action storage area in each stage includes storing multiple indicator data generated by each action in the stage at different time points. Each indicator data generates a unique data identifier formed by the combination of patient identifier, indicator code, collection time and data source. Before storage, it is checked whether indicator data with the same unique data identifier already exists. If it already exists, the business association information is updated without storing the indicator value repeatedly.
[0014] As a preferred embodiment of the method for collecting disease-specific medical data queues based on human body indicator sets according to the present invention, the method further includes, after constructing the disease-specific data queue, calculating the data quality score of the disease-specific data queue, wherein the data quality score is obtained by weighted calculation of data integrity score, data accuracy score and data consistency score; The data integrity score is calculated based on the ratio of the number of collected indicators to the number of indicators that should be collected; the data accuracy score is calculated based on the error rate between the measured value and the standard reference value; and the data consistency score is calculated based on the consistency judgment results of multi-source data. When the data quality score is lower than the preset quality threshold, the corresponding data will be pushed to the data quality control platform for review and correction.
[0015] Secondly, embodiments of the present invention provide a disease-specific medical data queue collection system based on human body indicator sets, which includes an indicator set construction module, which constructs a hierarchical human body indicator set, receives raw medical data from multiple data sources, establishes a mapping relationship between raw data fields and standardized indicators in the human body indicator set through a bidirectional mapping table, and converts the raw medical data into standardized indicator data. The time-series tag generation module constructs a four-level structured framework including lines, links, actions, and indicators. It adds time-series tags to standardized indicator data. The time-series tags record the business affiliation information and time information of the data. Based on the time information in the time-series tags and the time period attributes, the target link and target action to which the standardized indicator data belongs are determined through a time association algorithm. The data collection module collects the standardized indicator data to the storage location corresponding to the target link and target action, and sets up a historical data action storage area in each link to store multi-time series indicator data of the action in that link. The queue construction module integrates basic patient information, standardized indicator data of each step of the process, and multi-time series data from historical data actions to construct a disease-specific data queue.
[0016] Thirdly, embodiments of the present invention provide a computer device, including a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program instructions are executed by the processor, they implement the steps of the method for collecting disease-specific medical data queues based on human indicator sets as described in the first aspect of the present invention.
[0017] Fourthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program instructions are executed by a processor, they implement the steps of the method for collecting disease-specific medical data queues based on human body indicator sets as described in the first aspect of the present invention.
[0018] The beneficial effects of this invention are as follows: By constructing a hierarchical set of human indicators and establishing a bidirectional mapping table, this invention provides a unified conversion standard for multi-source heterogeneous data, automatically completing terminology alignment and unit conversion, eliminating data format differences between different medical systems, and reducing reliance on manual intervention during data conversion; it constructs a four-level structured framework including lines, links, actions, and indicators, closely integrating data collection with disease diagnosis and treatment pathways; by adding time-series labels containing business affiliation information and time information to standardized indicator data, combined with time association algorithms, it achieves precise association between data and business processes, avoiding business logic confusion caused by mixed data storage; it sets up a historical data action storage area in each link, completely saving multiple indicator data generated by the same action at different time points, and avoiding duplicate storage through unique data identifiers, ensuring data integrity and improving storage efficiency, enabling doctors to trace the trend of patient indicator changes; it integrates basic patient information, standardized indicator data of actions in each link, and historical data to construct a disease-specific data queue, providing a clear structure for clinical decision optimization and scientific research data analysis; and it enhances the value and utilization rate of medical data. Attached Figure Description
[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 The flowchart shows a method for collecting disease-specific medical data queues based on human indicator sets. Figure 2 A diagram of computer equipment used for a method of collecting disease-specific medical data queues based on human indicator sets; Figure 3 This is a schematic diagram of a three-level hierarchical structure for a disease-specific medical data queue collection method based on human body indicator sets. Detailed Implementation
[0021] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0022] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0023] Secondly, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single embodiment or an embodiment selectively excluded from other embodiments.
[0024] Example 1 Reference Figure 1 - Figure 2 This is the first embodiment of the present invention, which provides a method for collecting disease-specific medical data queues based on human indicator sets, including: S100: Construct a hierarchical human body indicator set, receive raw medical data from multiple data sources, establish a mapping relationship between raw data fields and standardized indicators in the human body indicator set through a two-way mapping table, and convert the raw medical data into standardized indicator data.
[0025] S200: Construct a four-level structured framework that includes lines, links, actions, and indicators. Add time-series tags to standardized indicator data. The time-series tags record the business affiliation information and time information of the data. Based on the time information and time period attributes in the time-series tags, determine the target link and target action to which the standardized indicator data belongs through a time association algorithm.
[0026] S300: Standardized indicator data is collected into the storage location corresponding to the target link and target action, and a historical data action storage area is set in each link to store multi-time series indicator data of the action in that link.
[0027] S400: Integrates basic patient information, standardized indicator data of each step of the process, and multi-time series data from historical data actions to construct a disease-specific data queue.
[0028] It should be noted that medical data integration lacks a unified data standardization system. Terminology alignment and unit conversion between different data sources rely on manual work, resulting in low data exchange efficiency and difficulty in ensuring accuracy. Insufficient attention is paid to the temporal characteristics of medical data, often only retaining the latest measurements of patient indicators while covering historical data. This makes it impossible for doctors to track the trend of disease changes, affecting efficacy evaluation and treatment plan adjustments. Data collection lacks connection with clinical business processes, and data generated from different diagnosis and treatment stages are stored in a mixed manner. When researchers construct disease data cohorts, they need to manually filter data from specific stages, which is cumbersome and prone to missing key information, resulting in incomplete data cohorts.
[0029] To address the aforementioned issues, S100 constructs a hierarchical set of human indicators and establishes a bidirectional mapping table to convert multi-source heterogeneous data into a standardized format, resolving the problems of inconsistent terminology and unit differences. S200 constructs a four-level structured framework and adds time-series labels to the data, combining time association algorithms to determine the business affiliation of the data, achieving precise association between data and the diagnosis and treatment process. S300 aggregates standardized indicator data to corresponding target steps and actions, and sets up a historical data action storage area to completely save multiple indicator data generated by the same action at different time points, solving the problem of incomplete time-series data management. S400 integrates basic patient information, standardized indicator data of actions in each step, and historical data to construct a clearly structured and logically rigorous disease-specific data queue, providing high-quality data support for clinical decision-making and scientific research analysis.
[0030] Example 2 Reference Figure 1 - Figure 3 This is the second embodiment of the present invention.
[0031] In this embodiment, step S100 involves constructing a hierarchical human body indicator set, receiving raw medical data from multiple data sources, establishing a mapping relationship between the raw data fields and standardized indicators in the human body indicator set through a bidirectional mapping table, and converting the raw medical data into standardized indicator data. This includes the following steps A1-A2: A1: The human body indicator set is organized according to the three-level hierarchical relationship of human body systems, organs and indicators. Each standardized indicator includes indicator code, standard name, synonym list, data type, unit of measurement and reference value range. Establishing a mapping relationship between raw data fields and standardized indicators in the human body indicator set through a two-way mapping table includes: for raw indicator names that cannot be mapped through the two-way mapping table, extracting the raw indicator name and counting the number of mapping failures; when the number of mapping failures exceeds a preset threshold, calculating the semantic relevance of the raw indicator name with the names of each standardized indicator in the human body indicator set and matching it with medical rules; when the overall confidence level meets the preset conditions, adding the raw indicator name to the synonym list of the corresponding standardized indicator and updating the two-way mapping table.
[0032] Specifically, the human body index set is based on the "Health and Health Information Data Element Directory", "Human Anatomy", "Clinical Path Management Guidelines" and the SNOMED-CT medical terminology standard to construct a three-level hierarchical structure of "human body system-organ-index" to ensure the scientificity and authority of the index system. Each index contains basic attributes such as a unique index code, standard name, synonyms, the organ it belongs to, data type, measurement unit, reference value range, etc. For example, the index code 01080400300046 corresponds to the standard name "fasting blood glucose value", synonyms include "blood glucose value, FBG", the data type is numerical, the measurement unit is "mmol / L", and the reference value range is 3.9-6.1 mmol / L. The management of the index set is realized through the system background, supporting operations such as adding, modifying, deleting and classifying indexes. For example, when adding an index of "glycated hemoglobin", it is necessary to clarify its belonging to the "hematological system - red blood cells" level and automatically update the index binding relationship in all associated actions.
[0033] To achieve the dynamic update of the index set, a rule engine is introduced to build an intelligent closed-loop mechanism for the dynamic discovery and update of synonyms; continuously monitor the multi-source data conversion process. When the mapping of an unrecognized index name (such as "serum glucose") fails repeatedly, the rule engine automatically triggers the synonym discovery process.
[0034] First, the engine conducts multi-dimensional analysis on unknown terms, performs pattern matching through preset medical rules (such as "terms containing 'glucose' and with a measurement unit of mmol / L or mg / dL are determined to be blood glucose-related indexes"), and at the same time queries external medical terminology libraries (such as the UMLS terminology library) to obtain standardized correspondence relationships.
[0035] Subsequently, the Word2Vec model is used to train the word vectors of unknown terms and standard index names. The training corpus includes, the model window size is set to 5, the number of iterations is 100, the learning rate is 0.025, the text is converted into 300-dimensional word vectors, and 10,000 medical treatment index records in the clinical electronic medical records of tertiary hospitals are selected as the training corpus; In the corpus preprocessing stage, the medical-specific word segmentation tool jieba-medical is used for word segmentation. For medical abbreviations such as "FBG (fasting blood glucose)" and "OGTT (oral glucose tolerance test)", standardized replacement is completed by establishing a dedicated abbreviation mapping table. At the same time, general stop words without semantic contributions such as "of", "and", "conduct" are removed, as well as fields that are not core information of the index such as "department for submission" and "reporting doctor", and only the core text directly related to the semantics of the term is retained; The training set and the test set are randomly divided in an 8:2 ratio, with 80% of the corpus used for model training and 20% used for model performance validation. During the division process, it is ensured that the distribution ratio of various data sources in the training set and the test set is consistent to avoid data bias affecting the model performance. Model validation employs a dual-metric evaluation approach: "semantic similarity recall + term matching accuracy". Recall is calculated as the proportion of known synonym pairs in the test set that are identified by the model, while accuracy is calculated as the proportion of synonym pairs identified by the model that are actually valid synonyms. In addition, manual evaluation by medical experts is combined, requiring a recall rate of ≥90% and an accuracy rate of ≥88% before the model can be put into use, ensuring that word vectors can accurately reflect the semantic associations of medical terms.
[0036] The semantic association strength is then calculated using the cosine similarity formula, which is: cosθ=(A·B) / (|A|×|B|) Here, A and B are the word vectors of the unknown term and the standard indicator name, respectively. The comprehensive confidence score is calculated by weighting semantic similarity and rule matching degree, with a weight allocation of 0.7 for semantic similarity and 0.3 for rule matching degree. At the same time, the weight allocation ratio is combined with the experience judgment of three senior medical data experts to confirm that the ratio conforms to the matching logic of most medical terms.
[0037] This weight can be dynamically adjusted, and the triggering conditions for its adjustment are as follows: If the overall accuracy of a certain type of indicator (such as blood test indicators or symptom description indicators) is lower than 85% for 500 consecutive matching data points, an adaptive weight adjustment algorithm will be activated.
[0038] The adjustment algorithm is based on the historical matching accuracy to construct its optimization formula: in, For semantic similarity weights, For rule matching degree weight, This represents the matching accuracy based solely on the current semantic similarity. This represents the average accuracy of the last three historical matches. Domain fit coefficient (blood test indicators) Symptom description indicators Other categories of indicators ).
[0039] The above formula allows for dynamic optimization of weights based on matching performance and domain characteristics. For example, unit matching is more important in blood test indicators, but when the accuracy is not up to standard, the weight of rule matching degree will automatically increase. In contrast, semantic association is more critical in symptom description indicators, so the weight of semantic similarity will increase accordingly.
[0040] When the overall confidence level is ≥0.85, unknown terms are automatically added to the synonym list of the corresponding indicator; when the confidence level is between 0.6 and 0.85, the process is transferred to manual review, where medical experts confirm and decide whether to add the term; when the confidence level is <0.6, the term is marked as an unmatched term and sent to the management backend for further processing. All verified updates are synchronized to the knowledge base to continuously optimize the system's semantic recognition capabilities.
[0041] A2: Converting raw medical data into standardized indicator data includes parsing the fields of the raw medical data and extracting the indicator name, measurement value, measurement time and unit of measurement. Query the standard indicator code corresponding to the indicator name through the two-way mapping table; Obtain the standard unit of measurement based on the standard indicator code, determine whether the unit of measurement of the measured value is consistent with the standard unit of measurement, and if not, convert the unit according to the preset conversion formula. The converted measurement values are associated with the standard indicator codes to generate standardized indicator data.
[0042] Specifically, the data from multiple sources, including HIS, LIS, and PACS, is first parsed, with corresponding parsing methods used for different data source types. The medical order data in the HIS system is stored in XML format, and the DOM4J parser is used to extract fields such as patient ID, examination name, measurement value, measurement time, and unit of measurement.
[0043] The LIS system stores test data in an SQL Server database. It connects to the database via JDBC and executes the SQL query "SELECT Patient ID, Test Item Name, Measurement Result, Measurement Time, Unit FROM Test Record Table WHERE Data Status = 'Reviewed'" to extract valid data.
[0044] The PACS system uses optical character recognition (OCR) technology to extract text information from image report data, and then uses regular expressions to match key indicator data.
[0045] After data parsing is completed, the system enters the indicator mapping stage. The system associates the original indicators with the standard indicators in the human body indicator set through a two-way mapping table. The mapping table records information such as the original indicator name, data source type, standard indicator code, and mapping confidence. The mapping confidence is calculated by weighting semantic similarity and unit matching degree. The formula is: confidence = 0.7 × semantic similarity + 0.3 × unit matching degree (the unit matching degree is 1 when the units are consistent, 0.8 when they are inconsistent but can be converted, and 0 when they cannot be converted).
[0046] For example, mapping the "blood glucose value (mg / dL)" in the LIS system to the standard index 01080400300046 yields a confidence level of 0.92.
[0047] After mapping, data conversion is performed, including data type conversion, unit unification, and missing value handling. Data type conversion converts string values in the original data to floating-point values, for example, converting "65" to 65.0. Unit unification follows a preset conversion formula, such as the formula for converting blood glucose values from mg / dL to mmol / L: mmol / L = mg / dL × 0.0555, converting 65 mg / dL to 3.61 mmol / L.
[0048] Missing values were handled using a multiple imputation method. Based on other indicator data from the same patient at the same stage and the indicator distribution characteristics of similar patients, three sets of imputed values were generated. The average value was taken as the final imputed value, and a "imputation" label was added to ensure data integrity without affecting data accuracy. After data transformation, a standardized indicator list containing information such as indicator code, indicator value, time series label, data source, and data quality score was generated.
[0049] The initial data quality score in the data quality scoring transformation stage mainly addresses the core sources of error in the transformation process, and is calculated using a weighted summation formula: in, Field mapping accuracy (number of successfully mapped metrics / total number of metrics). These are the converted indicator values. This refers to a reference value calibrated using standard methods (such as laboratory standard measurements). The relative error of unit conversion (when (Time error meter is 0) This represents the number of missing indicators. This represents the total number of indicators.
[0050] The scoring here is inconsistent with the quality scoring in the data collection phase. The scoring in the data collection phase is as follows: exist This is based on an assessment of the compatibility between data and business processes. Will as middle" "The core part of the dimension (accounting for 60% of the accuracy score)."
[0051] Their core difference lies in the fact that the scoring in the transformation phase only focuses on the "errors in the standardized process," while the collection phase needs to add "data and business process compatibility assessment"—including the accuracy of matching data with lines, processes, and actions, as well as new dimensions such as consistency verification of multi-source data.
[0052] To reflect the error propagation during the conversion process, an error propagation coefficient k can be introduced, where k = 0.8 + 0.2 × That is, when When k ≥ 0.9, k = 1.0. When k < 0.6, the accuracy score for the aggregation stage = 0.6 × ×k+0.4× other accuracy check items ensure that the impact of conversion errors on the final data quality is quantifiable and traceable.
[0053] In this embodiment, step S200 constructs a four-level structured framework including lines, links, actions, and indicators. Time-series tags are added to the standardized indicator data. These tags record the business affiliation and time information of the data. Based on the time information and time period attributes in the time-series tags, a time association algorithm is used to determine the target link and target action to which the standardized indicator data belongs. This includes the following steps B1-B2: B1: The lines correspond to the disease diagnosis and treatment pathway, the links correspond to the diagnosis and treatment stages and have start and end times, and the actions correspond to specific clinical operations and are associated with one or more standardized indicators; the time sequence label records the patient identifier, line identifier, link identifier, action identifier, data collection time, indicator code, data source and data version number; The time correlation algorithm used to determine the target link and target action to which standardized indicator data belongs includes: obtaining the collection time of standardized indicator data and the start and end times of each link in the line; determining whether the collection time falls within the time range of a certain link; when the collection time falls within the time range of multiple links at the same time, calculating the belonging priority score of each link based on the time overlap ratio, link priority coefficient and indicator correlation strength, and aggregating the standardized indicator data to the link with the highest belonging priority score.
[0054] Specifically, the four-level structure of "line-step-action-indicator" is the core innovation of this invention in achieving structured data management. The lines are divided according to the "Guidelines for Clinical Pathway Management," such as the "Type 2 Diabetes Line" and the "Bladder Cancer Line." Each line contains multiple steps, and the division of steps is consistent with the clinical diagnosis and treatment process or research implementation steps. Each step has a start time and an effective time period. The time period can be dynamically adjusted according to the patient's actual treatment progress or the research project schedule. For example, the screening step of the "diabetes line" starts at the patient's first visit and ends when all screening items are completed. If the patient completes the screening within one month, the effective time period of the step is one month. If the screening is not completed within the time limit, the system automatically sends a reminder to medical staff.
[0055] The process involves multiple actions, each corresponding to a clinical procedure or research task. Each action comprises one or more core indicators from a set of human indicators. For example, the "blood glucose measurement" action in the screening process is associated with the "fasting blood glucose level" indicator, while the "OGTT test" action in the diagnostic process is associated with three indicators: "fasting blood glucose level," "1-hour postprandial blood glucose level," and "2-hour postprandial blood glucose level." The design of these actions strictly adheres to clinical operating procedures or research trial protocols, ensuring a high degree of alignment between data collection and business processes. Indicators, as the smallest unit describing a living organism, are closely linked to business processes through actions, ensuring that each indicator data point accurately corresponds to a specific business scenario.
[0056] B2: The priority coefficient of each step is set according to the type of step. The priority coefficient of treatment steps is higher than that of diagnosis steps, and the priority coefficient of diagnosis steps is higher than that of screening steps. The time overlap ratio is the proportion of the overlap between the data collection time and the time range of the process to the total time of the process. The correlation strength of an indicator indicates the degree of correlation between the indicator and the actions within a process; the higher the degree of correlation, the greater the correlation strength.
[0057] In this embodiment, step S300 involves aggregating standardized indicator data to the storage location corresponding to the target stage and target action, and setting up a historical data action storage area in each stage to store multi-time series indicator data of actions within that stage, including the following step C1: C1: In each stage, a historical data action storage area is set up. The historical data action storage includes multiple indicator data generated by each action in the stage at different time points. Each indicator data generates a unique data identifier formed by the combination of patient identification, indicator code, collection time and data source. Before storage, it is checked whether indicator data with the same unique data identifier already exists. If it already exists, the business association information is updated instead of storing the indicator value repeatedly.
[0058] Specifically, given the multi-time-series characteristics of medical data, time-series tags can add multi-dimensional identifiers to each indicator data, including fields such as patient ID, line ID, process ID, action ID, data collection timestamp, indicator code, data source, and data version number. For example, the patient ID is P2023001, the line ID is D2023002 (type 2 diabetes line), the process ID is S2023003 (screening process), the action ID is A2023004 (blood glucose measurement), the timestamp is 2023-05-10 08:30:25.123, the indicator code is 01080400300046, the data source is the LIS system, and the version number is 1.0. Time-series tags enable precise association between data and processes and time dimensions.
[0059] Each step includes a new "Historical Data" action, specifically designed to store historical indicator data for that step. For example, the "Historical Data_Screening" action in the screening step stores historical records of all blood glucose and blood pressure measurements performed during that step, enabling multi-version data management. Data storage employs a collaborative architecture combining a time-series database (InfluxDB) and a relational database (MySQL). The time-series database stores indicator data in the format "measurement=indicator code, tags={patient ID, line ID, step ID, action ID, data source}, fields={indicator value, data quality score}, time=timestamp," ensuring efficient storage and retrieval of time-series data. The relational database stores metadata such as lines, steps, and actions. Table structures include "Line Information Table" (Line ID, Line Name, Creation Time, Applicable Diseases, Status), "Step Information Table" (Step ID, Line ID, Step Name, Start Time, End Time, Description), and "Action Information Table" (Action ID, Step ID, Action Name, Associated Indicator Code, Operation Specification), ensuring consistency between metadata and time-series data.
[0060] Data storage follows a time-series storage principle, implementing storage operations based on the sequence of processes and the time of data collection. For example, blood glucose data from the screening process is stored first, followed by OGTT test data from the diagnostic process, with data within the same process arranged in ascending order of timestamps.
[0061] The dynamic association algorithm ensures that indicator data is automatically included in the "historical data" of the corresponding stage. The algorithm process is as follows: First, the patient ID is used to query the treatment path records in the electronic medical record to determine the path the patient is currently involved in; Secondly, the "Segment Information Table" in the relational database is called to obtain all segments of the line and their corresponding time period attributes; For each standardized indicator data, its collection timestamp is extracted, and the time interval matching algorithm is used to determine whether the collection timestamp is within the time range of a certain stage (i.e., stage start time ≤ collection timestamp ≤ stage end time). When a data collection timestamp falls within the time range of multiple stages, it is not allowed for a single data point to be associated with multiple stages simultaneously. Therefore, a stage attribution priority formula is used to determine its unique attribution. The formula is as follows: in, The priority score for the i-th stage is given. These are the weighting coefficients. The overlap ratio between the data collection time and the time range of the process is (O(i) = min(collection timestamp, process end time) - max(collection timestamp, process start time)) / (process end time - process start time)). Priority coefficient for treatment steps Diagnostic Screening Follow-up ), The correlation strength between indicators and actions in a process (obtained from the "Action-Indicator Correlation Table", with a value range of 0.6-1.0, and the higher the value for a more core correlation).
[0062] Calculate all overlapping elements The highest-scoring stage is assigned as the data stage; if stages overlap in time, the system will first prioritize them according to their priority coefficients. A preliminary sorting is performed, with higher priority steps participating first in the attribution determination. Same and then by overlapping ratio and correlation strength Further distinctions are needed.
[0063] For indicators such as "fasting blood glucose level" that are associated with multiple actions, in order to avoid duplicate storage, the system generates a unique data identifier for each indicator data (composed of patient ID + indicator code + collection timestamp + data source). Before storage, it checks whether the identifier already exists in the time series database. If it already exists, it only updates the stage ID, action ID and association relationship in the time series label, without duplicate storage of core data such as indicator values, ensuring storage efficiency while retaining complete business association traces.
[0064] If the collected timestamp falls within the time range of a certain stage, the "Action-Indicator Association Table" is queried to determine whether the indicator belongs to a certain action in that stage. If an association exists, the indicator data is automatically stored in the "Historical Data" action of the corresponding action, and the stage ID and action ID in the time sequence label are updated to ensure accurate matching between the data and the business process.
[0065] In this embodiment, step S400 integrates basic patient information, standardized indicator data of each step, and multi-time-series data from historical data actions to construct a disease-specific data queue, including the following step D1: D1: After constructing the disease-specific data queue, the following steps are also included: calculating the data quality score of the disease-specific data queue. The data quality score is calculated by weighting the data integrity score, data accuracy score, and data consistency score. The data integrity score is calculated based on the ratio of the number of collected indicators to the number of indicators that should be collected; the data accuracy score is calculated based on the error rate between the measured value and the standard reference value; and the data consistency score is calculated based on the consistency judgment results of multi-source data. When the data quality score is lower than the preset quality threshold, the corresponding data will be pushed to the data quality control platform for review and correction.
[0066] Specifically, the construction process of the disease-specific data queue revolves around four stages: data collection, queue construction, queue optimization, and queue application. Data collection utilizes a streaming processing framework (Flink) to receive standardized indicator data in real time. Based on the line ID, stage ID, and action ID in the time-series label, the data is routed to the corresponding collection channel. Each channel corresponds to a unique combination of "line-stage-action," such as the "Type 2 Diabetes Line-Screening Stage-Blood Glucose Measurement Action" channel. Data within a channel is sorted by timestamp, and duplicate records are removed using a duplicate data determination rule (patient ID, indicator code, and timestamp must be identical to be considered duplicate data), ensuring data uniqueness. The queue construction stage integrates basic patient information (name, gender, age, ID number), treatment information (indicator data for each stage of action), and historical data (historical records of each action) to form a structured disease-specific data queue. The queue data model fields include "Queue ID, Patient ID, Line ID, Basic Information, Summary of Data for Each Stage, Historical Data Index, Data Integrity Score, Creation Time, and Update Time." The summary of data for each stage stores the indicator data and time distribution of all actions in that stage in JSON format.
[0067] Queue optimization is based on data quality scoring, which uses a weighted scoring method with the following weights: 40% for data integrity, 30% for data accuracy, and 30% for data consistency. The data integrity score is calculated as follows: Data integrity score = number of collected indicators / [(number of indicators associated with mandatory actions - number of reasonable but not executed mandatory action indicators) + number of indicators associated with executed optional actions].
[0068] The actions are divided into mandatory actions (such as "fasting blood glucose measurement" in the diabetes screening process) and optional actions (such as "insulin antibody detection") according to the standards of clinical diagnosis and treatment. The number of related indicators for mandatory actions is a fixed value, while the related indicators for optional actions are only included in the "number of indicators to be collected" when they are actually performed. If a mandatory action is not performed due to the patient's special condition (such as allergies or contraindications), the relevant indicator number of that action will be deducted from the denominator after the attending physician marks it as "reasonably not performed" in the system and uploads the medical record evidence, thus avoiding score distortion due to reasonable differences in diagnosis and treatment. For dynamically adjusted treatment pathways (such as when a patient's condition worsens and additional treatment steps are needed), the system will monitor for adjustments in real time. When the pathway changes, the "number of mandatory action-related indicators" and "list of optional actions" will be automatically updated. The system will also record the adjustment time, reason for adjustment, and operator to ensure that the "number of indicators to be collected" is consistent with the actual treatment pathway. Considering the timeliness of indicator collection, a collection time window is set for each indicator (such as "within 24 hours after surgery" or "within 8-12 hours of fasting"). If an indicator is not collected within the time window or the collection time exceeds the window range, then the indicator is counted as missing and will be included in the denominator but not in the numerator. If the time window needs to be extended due to the patient's special circumstances, then the window parameters need to be adjusted and the record kept after approval by the medical quality administrator. The data accuracy score is calculated by comparing the data with the gold standard data (such as the laboratory test standard value). Error rate = |measured value - standard value| / standard value. Data accuracy score = 1 - error rate (1 when the error rate is ≤10%).
[0069] Data consistency scoring is determined by comparing multi-source data. Data from the same indicator at the same time point is considered consistent if the difference between multi-source data is ≤ a set threshold (blood glucose ≤ 0.5 mmol / L, blood pressure ≤ 10 mmHg), and the consistency score is 1; otherwise, it is 0. The final data quality score is calculated as follows: 0.4 × data integrity score + 0.3 × data accuracy score + 0.3 × data consistency score. A score ≥ 80 indicates high-quality data, 60-79 indicates acceptable data, and < 60 indicates data requiring correction. Data requiring correction is automatically pushed to the data quality control platform for review and correction by medical staff. After correction, the data quality score is recalculated.
[0070] The application scenarios of the cohort include clinical decision-making and scientific research analysis. For example, in clinical decision-making, based on the blood glucose change trends of patients in a diabetes-specific disease data cohort, an LSTM neural network is used to construct a complication risk prediction model. The input is the patient's blood glucose time-series data for the past 6 months (features include blood glucose value, measurement time interval, dietary records, and medication use), and the output is the probability of developing diabetic nephropathy in the next 3 months (a value between 0 and 1). The model training process is as follows: Data preprocessing normalizes blood glucose values to the [0,1] interval, and converts dietary records and medication use into binary features; the network structure is designed with 128 neurons in the input layer, 3 hidden layers (64 neurons per layer), and 1 neuron in the output layer; the loss function uses the cross-entropy loss function. For imbalanced datasets (where the complication rate may be low), it is necessary to consider that a low complication rate can lead to dataset imbalance. The cross-entropy loss function is then optimized using a weighted binary classification cross-entropy loss function, with the specific formula as follows: in, The total number of training samples, The true label of sample i ( This indicates the occurrence of complications. (Indicates that it did not happen). The model predicts the probability of complication occurring in sample i (within the range of 0-1). and Positive samples ( ) and negative samples ( The weight of ) is calculated as follows: in, To determine the number of positive samples in the training set, The number of negative samples in the training set ( ).
[0071] The weighting method can balance the loss contribution of positive and negative samples, and also prevent the model from being biased towards predicting non-complications due to an excessive proportion of negative samples. After optimization, the model's recall rate for identifying positive samples is improved by more than 25%.
[0072] Optimizing time-series data queries is primarily achieved through database selection, index design, and preprocessing and caching mechanisms. Database selection should combine the advantages of both time-series and relational databases. Using InfluxDB to store time-series metric data enables efficient time-range queries and batch data read / write operations; using MySQL to store metadata ensures stable business relationships.
[0073] Index optimization involved creating composite indexes in InfluxDB with index fields combining (Patient ID + Timestamp) and (Line ID + Stage ID + Timestamp), using a B+ tree index structure. This reduces disk I / O operations during queries and improves the efficiency of range queries. In the MySQL database, single indexes were created for fields such as "Line ID," "Stage ID," and "Action ID" to accelerate metadata queries.
[0074] The preprocessing and caching mechanism uses Redis to cache hot query results. Hot queries include "all indicator data of a certain patient in a certain line" and "all patient data of a certain line in a certain stage". The cache validity period is set to 1 hour. If the data is updated during this period (such as modification of indicator values or addition of historical data), the corresponding cache will be automatically invalidated and recalculated and cached.
[0075] For example, the frequently asked query "querying blood glucose measurement records for the most recent 3 months in the diabetes screening process" will be pre-calculated according to patient ID and time range, and the results will be stored in Redis. When querying, it will be read directly from the cache, and the response time will be controlled within 100ms to meet the real-time query requirements.
[0076] Data conflict handling resolves inconsistencies in multi-source data. The timestamp priority mechanism compares multi-source data of the same action at the same time point (error ≤ 10ms), and takes the data with the largest timestamp as valid data. If the timestamps are completely consistent, the priority of the data source is further compared (clinical laboratory data > imaging report data > medical order data), and the data from the higher priority data source is taken.
[0077] Version control assigns a version number to each indicator data point, using the format "major version number.minor version number". The initial entry is 1.0, and the major version number increments by 1 for each modification. Minor corrections (such as unit conversion errors) increment the minor version number by 1. The system establishes a "Data Version Information Table" to record the version number, patient ID, indicator code, modification time, modifier ID, reason for modification, original data, and modified data, ensuring that data modifications are traceable.
[0078] Conflict marking is used to identify cases where the difference between multiple sources of data for the same indicator for the same patient at the same time point exceeds a set threshold (blood glucose ≥1.0 mmol / L, blood pressure ≥20 mmHg). These data are automatically marked as conflicting data, and a conflict report is generated. The report includes information such as data source, value, timestamp, difference, and confidence level. It is pushed to the corresponding attending physician through the hospital's messaging system. After review, the physician can choose to retain the correct data or merge the data (take the average value and mark it with "merged") to ensure that data conflicts are resolved reasonably.
[0079] In summary, by introducing a rule engine to construct an intelligent closed loop, using the Word2Vec model for word vector training and calculating semantic relevance, and combining this with medical rule matching to form a comprehensive confidence judgment, the system achieves dynamic discovery and automatic updating of indicator synonyms, enabling it to continuously adapt to newly emerging terminology. Different parsing methods are used for different types of data sources such as HIS, LIS, and PACS. A bidirectional mapping table maps raw indicators to standard indicators, units are unified according to a preset conversion formula, and missing values are handled using multiple interpolation methods, generating a standardized indicator list containing complete information such as indicator codes, indicator values, and time series labels. A four-level structured framework is constructed, and the start and end times of each stage are set. A time interval matching algorithm determines the time allocation of data. When data collection time falls into multiple stages simultaneously, the allocation priority score is calculated based on the time overlap ratio, stage priority coefficient, and indicator correlation strength to ensure accurate data allocation to the corresponding stage. A collaborative storage architecture combining time series databases and relational databases is adopted, generating a unique data identifier for each indicator data. Before storage, it checks whether the same identifier already exists to avoid duplicate storage, preserving complete business relationship traces while ensuring storage efficiency. Cohort optimization is performed based on data quality scores. High-quality data is selected through weighted calculations of data integrity, accuracy, and consistency to construct a high-quality disease-specific data cohort, providing reliable data support for clinical applications and scientific research analysis.
[0080] Example 3 The above is an illustrative scheme for a method of collecting disease-specific medical data queues based on human body indicator sets. It should be noted that the technical solution of this system for collecting disease-specific medical data queues based on human body indicator sets belongs to the same concept as the technical solution of the aforementioned method for collecting disease-specific medical data queues based on human body indicator sets. Details not described in detail in the technical solution of the system for collecting disease-specific medical data queues based on human body indicator sets in this embodiment can be found in the description of the aforementioned method for collecting disease-specific medical data queues based on human body indicator sets.
[0081] This embodiment also provides a disease-specific medical data queue collection system based on human indicator sets, including: The indicator set construction module constructs a hierarchical human indicator set, receives raw medical data from multiple data sources, establishes a mapping relationship between raw data fields and standardized indicators in the human indicator set through a bidirectional mapping table, and converts the raw medical data into standardized indicator data. The time-series tag generation module constructs a four-level structured framework including lines, links, actions, and indicators. It adds time-series tags to standardized indicator data. The time-series tags record the business affiliation information and time information of the data. Based on the time information in the time-series tags and the time period attributes, the target link and target action to which the standardized indicator data belongs are determined through a time association algorithm. The data collection module collects the standardized indicator data to the storage location corresponding to the target link and target action, and sets up a historical data action storage area in each link to store multi-time series indicator data of the action in that link. The queue construction module integrates basic patient information, standardized indicator data of each step of the process, and multi-time series data from historical data actions to construct a disease-specific data queue.
[0082] This embodiment also provides an electronic device suitable for the collection of disease-specific medical data queues based on human body indicator sets, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the method for collecting disease-specific medical data queues based on human body indicator sets as proposed in the above embodiment.
[0083] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, it implements the method for collecting disease-specific medical data queues based on human indicator sets as proposed in the above embodiments.
[0084] The storage medium proposed in this embodiment and the method for collecting disease-specific medical data queues based on human indicator sets proposed in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0085] Based on the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.
[0086] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for collecting disease-specific medical data queues based on human body indicator sets, characterized in that: This includes constructing a hierarchical human body indicator set, receiving raw medical data from multiple data sources, establishing a mapping relationship between raw data fields and standardized indicators in the human body indicator set through a two-way mapping table, and converting raw medical data into standardized indicator data. A four-level structured framework including lines, links, actions, and indicators is constructed. Time-series tags are added to the standardized indicator data. The time-series tags record the business affiliation information and time information of the data. Based on the time information in the time-series tags and the time period attributes, the target link and target action to which the standardized indicator data belongs are determined through a time association algorithm. The standardized indicator data is collected into the storage location corresponding to the target link and the target action, and a historical data action storage area is set in each link to store multi-time series indicator data of the action in that link. By integrating basic patient information, standardized indicator data of each step, and multi-time-series data from historical data actions, a disease-specific data queue is constructed.
2. The method for collecting disease-specific medical data queues based on human indicator sets as described in claim 1, characterized in that: The human body indicator set is organized according to a three-level hierarchical relationship of human body systems, organs and indicators. Each standardized indicator includes indicator code, standard name, synonym list, data type, unit of measurement and reference value range. The step of establishing a mapping relationship between the original data fields and the standardized indicators in the human body indicator set through a bidirectional mapping table includes: for the original indicator names that fail to be mapped through the bidirectional mapping table, extracting the original indicator names and counting the number of mapping failures; when the number of mapping failures exceeds a preset threshold, performing semantic association calculation and medical rule matching between the original indicator names and the names of each standardized indicator in the human body indicator set; when the comprehensive confidence meets the preset conditions, adding the original indicator name to the synonym list of the corresponding standardized indicator and updating the bidirectional mapping table.
3. The method for collecting disease-specific medical data queues based on human indicator sets as described in claim 2, characterized in that: The process of converting raw medical data into standardized indicator data includes parsing the fields of the raw medical data and extracting the indicator name, measurement value, measurement time and unit of measurement. The standard indicator code corresponding to the indicator name is retrieved using the bidirectional mapping table; Obtain the standard unit of measurement according to the standard index code, determine whether the unit of measurement of the measured value is consistent with the standard unit of measurement, and if they are inconsistent, perform unit conversion according to the preset conversion formula. The converted measurement values are associated with the standard indicator codes to generate standardized indicator data.
4. The method for collecting disease-specific medical data queues based on human indicator sets as described in claim 3, characterized in that: The lines correspond to disease diagnosis and treatment pathways, the steps correspond to diagnosis and treatment stages and have start and end times, and the actions correspond to specific clinical operations and are associated with one or more standardized indicators; the time sequence labels record patient identifier, line identifier, step identifier, action identifier, data collection time, indicator code, data source, and data version number; The time correlation algorithm used to determine the target stage and target action to which standardized indicator data belongs includes: obtaining the collection time of the standardized indicator data, obtaining the start time and end time of each stage in the line, determining whether the collection time falls within the time range of a certain stage, and when the collection time falls within the time range of multiple stages, calculating the belonging priority score of each stage based on the time overlap ratio, stage priority coefficient and indicator correlation strength, and aggregating the standardized indicator data to the stage with the highest belonging priority score.
5. The method for collecting disease-specific medical data queues based on human indicator sets as described in claim 4, characterized in that: The priority coefficient of each step is set according to the step type. The priority coefficient of treatment steps is higher than that of diagnosis steps, and the priority coefficient of diagnosis steps is higher than that of screening steps. The time overlap ratio is the proportion of the overlap between the data acquisition time and the time range of the process to the total time of the process. The correlation strength of the indicator represents the degree of correlation between the indicator and the actions within the process; the higher the degree of correlation, the greater the correlation strength of the indicator.
6. The method for collecting disease-specific medical data queues based on human indicator sets as described in claim 5, characterized in that: The provision of a historical data action storage area in each stage includes storing multiple indicator data generated at different time points for each action within that stage. Each indicator data generates a unique data identifier formed by a combination of patient identification, indicator code, collection time, and data source. Before storage, it is checked whether indicator data with the same unique data identifier already exists. If it already exists, the business association information is updated without storing the indicator value repeatedly.
7. The method for collecting disease-specific medical data queues based on human indicator sets as described in claim 6, characterized in that: After constructing the disease-specific data queue, the method further includes calculating the data quality score of the disease-specific data queue. The data quality score is obtained by weighting the data integrity score, data accuracy score, and data consistency score. The data integrity score is calculated based on the ratio of the number of collected indicators to the number of indicators that should be collected; the data accuracy score is calculated based on the error rate between the measured value and the standard reference value; and the data consistency score is calculated based on the consistency judgment results of multi-source data. When the data quality score is lower than the preset quality threshold, the corresponding data will be pushed to the data quality control platform for review and correction.
8. A disease-specific medical data queue collection system based on human body indicator sets, based on the disease-specific medical data queue collection method based on human body indicator sets as described in any one of claims 1 to 7, characterized in that: It also includes an indicator set construction module, which constructs a hierarchical human indicator set, receives raw medical data from multiple data sources, establishes a mapping relationship between raw data fields and standardized indicators in the human indicator set through a bidirectional mapping table, and converts the raw medical data into standardized indicator data. The time-series tag generation module constructs a four-level structured framework including lines, links, actions, and indicators. It adds time-series tags to standardized indicator data. The time-series tags record the business affiliation information and time information of the data. Based on the time information in the time-series tags and the time period attributes, the target link and target action to which the standardized indicator data belongs are determined through a time association algorithm. The data collection module collects the standardized indicator data to the storage location corresponding to the target link and target action, and sets up a historical data action storage area in each link to store multi-time series indicator data of the action in that link. The queue construction module integrates basic patient information, standardized indicator data of each step of the process, and multi-time series data from historical data actions to construct a disease-specific data queue.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the method for collecting disease-specific medical data queues based on human body indicator sets as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the method for collecting disease-specific medical data queues based on human body indicator sets as described in any one of claims 1 to 7.