Disease data processing method and device and computer readable storage medium

CN122619236BActive Publication Date: 2026-09-29FUWAI HOSPITAL CHINESE ACAD OF MEDICAL SCI & PEKING UNION MEDICAL COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611095915.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-22
Publication Date
2026-09-29
Estimated Expiration
2046-07-22

AI Technical Summary

Technical Problem

[0004]然而,现有技术多聚焦于数据采集、数据治理和数据质控中的某一单点,难以完全自动化构建临床数据库,而且对各专病的结构化抽取的过程中抽取标准不一致,难以复现,需要对各专病单独处理,以构建对应的临床数据库

Benefits of technology

[0011]本申请实施例的疾病数据处理方法、装置及计算机可读存储介质,应用于疾病数据处理系统,疾病数据处理系统中预定义了数据元模式,数据元模式中包括至少两个结构化疾病属性,结构化疾病属性包括临床后果严重等级属性和决策属性,通过疾病数据处理系统的采集通道调度引擎,采集多源病例数据,通过疾病指标提取模型和疾病指标提取模型对应的指标优选规则,对多源病例数据进行指标抽取,得到包括至少一个疾病指标的待治理数据,从而能够使多源病例数据均在同一个维度上,以结构化多源病例数据,提高采集效率,根据待治理数据确定数据元模式的至少一个结构化疾病属性的属性值,以将属性值直接存储入专病数据库,得到直观的结论,通过数据元模式确定待治理数据的字段风险等级,根据字段风险等级确定待治理数据的各字段对应的质控判定规则,从而可以通过字段风险等级分级进行判定,提升自动化程度,且指标优选规则透明,提高可解释性,通过质控判定规则,通过至少一个验证推理模型和验证推理模型对应的指标优选规则对待治理数据进行验证,多维度验证待治理数据能够提高准确性,在待治理数据通过验证的情况下,根据待治理数据确定治理结果,实现高质量的自动结构化入库方式,通过数据元模式能够解耦数据处理阶段和入库阶段,提高疾病数据处理系统的泛化能力,同时提升疾病数据处理的准确度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122619236B_ABST
    Figure CN122619236B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a disease data processing method and device and a computer readable storage medium. The method is applied to a disease data processing system, the disease data processing system predefines a data element mode, the data element mode includes at least two structured disease attributes, the structured disease attributes include a clinical consequence severity level attribute and a decision attribute, and the method includes: collecting multi-source case data; performing index extraction on the multi-source case data to obtain to-be-managed data containing at least one disease index; determining attribute values of at least one structured disease attribute based on the to-be-managed data; determining a field risk level of the to-be-managed data and a quality control determination rule corresponding to each field; verifying the to-be-managed data based on the quality control determination rule; and in a case where the to-be-managed data passes the verification, determining a management result based on the to-be-managed data. According to the method provided in the embodiments of the present application, the generalization capability of the disease data processing system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of data processing technology, and in particular relates to a disease data processing method, apparatus and computer-readable storage medium. Background Technology

[0002] With technological advancements, various case data can now be stored in corresponding clinical databases in the medical field.

[0003] In the existing technology, there are clinical databases built for various diseases. In the process of building a clinical database, it is possible to connect with electronic medical record systems and extract clinical indicators from unstructured medical record texts based on rule templates or natural language processing technology of general large language models.

[0004] However, existing technologies mostly focus on a single point in data collection, data governance, and data quality control, making it difficult to fully automate the construction of clinical databases. Moreover, the extraction standards are inconsistent during the structured extraction process for each disease, making it difficult to reproduce. Each disease needs to be processed separately to construct the corresponding clinical database. Summary of the Invention

[0005] This application provides a disease data processing method, apparatus, and computer-readable storage medium, which can process various disease data through data element model and store them in the corresponding disease-specific database, thereby improving the versatility and generalization of disease data processing.

[0006] In a first aspect, embodiments of this application provide a disease data processing method. The method is applied to a disease data processing system, which predefines a data element schema. The data element schema includes at least two structured disease attributes, namely, a clinical consequence severity level attribute and a decision attribute. The method includes: Multi-source case data is collected through the data acquisition channel scheduling engine of the disease data processing system; By using the disease indicator extraction model and the corresponding indicator optimization rules, indicators are extracted from multi-source case data to obtain data to be treated that contains at least one disease indicator. The attribute value of at least one structured disease attribute of the data element pattern is determined based on the data to be governed; The risk level of each field in the data to be governed is determined by the data element model, and the quality control judgment rules corresponding to each field of the data to be governed are determined based on the field risk level. Based on the quality control judgment rules, the data to be governed is verified through at least one verification inference model and the corresponding indicator selection rules of the verification inference model; If the data to be governed passes verification, the governance outcome is determined based on the data to be governed.

[0007] Secondly, embodiments of this application provide a disease data processing apparatus, which is applied to a disease data processing system. The apparatus includes: The acquisition module is used to collect multi-source case data through the acquisition channel scheduling engine of the disease data processing system; The extraction module is used to extract indicators from multi-source case data through the disease indicator extraction model and the indicator optimization rules corresponding to the disease indicator extraction model, so as to obtain data to be treated containing at least one disease indicator. A structured module for determining the attribute value of at least one structured disease attribute of a data element pattern based on the data to be governed; The quality control module is used to determine the field risk level of the data to be governed through the data element model, and to determine the quality control judgment rules corresponding to each field of the data to be governed based on the field risk level. The verification module is used to verify the data to be governed based on the quality control judgment rules, through at least one verification inference model and the indicator selection rules corresponding to the verification inference model; The output module is used to determine the governance result based on the data to be governed, provided that the data to be governed has passed verification.

[0008] Thirdly, embodiments of this application provide a terminal device, the device including: a processor and a memory storing computer program instructions; the processor reads and executes the computer program instructions to implement the disease data processing method as described in the first aspect.

[0009] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the disease data processing method as described in the first aspect.

[0010] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the disease data processing method as described in the first aspect.

[0011] The disease data processing method, apparatus, and computer-readable storage medium of this application are applied to a disease data processing system. The system predefines a data element schema, which includes at least two structured disease attributes, including a clinical consequence severity level attribute and a decision attribute. The system collects multi-source case data through its acquisition channel scheduling engine. Using a disease indicator extraction model and corresponding indicator optimization rules, indicators are extracted from the multi-source case data to obtain data to be processed, including at least one disease indicator. This ensures that the multi-source case data is on the same dimension, thus structuring the multi-source case data and improving acquisition efficiency. The attribute value of at least one structured disease attribute of the data element schema is determined based on the data to be processed, and the attribute value is directly stored. By inputting data into a disease-specific database, intuitive conclusions are obtained. The risk level of each field in the data to be treated is determined through the data element schema. Based on the risk level, quality control judgment rules are determined for each field in the data to be treated. This allows for judgment based on the risk level of each field, improving automation. The indicator selection rules are transparent, enhancing interpretability. Through the quality control judgment rules, the data to be treated is verified using at least one verification inference model and the corresponding indicator selection rules. Multi-dimensional verification of the data to be treated improves accuracy. If the data to be treated passes verification, the treatment result is determined based on the data, achieving a high-quality, automated, structured data entry method. The data element schema decouples the data processing stage from the data entry stage, improving the generalization ability of the disease data processing system and enhancing the accuracy of disease data processing. Attached Figure Description

[0012] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a flowchart illustrating a disease data processing method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the architecture of a disease data processing system provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a disease data processing device provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation

[0014] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.

[0015] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.

[0016] It should be noted that the acquisition, storage, use, and processing of data in this application embodiment all comply with the relevant provisions of laws and regulations.

[0017] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.

[0018] This application relates to the field of medical information processing and clinical data governance technology, specifically to an automatic construction system and method for a specialized clinical database. This system comprises a data acquisition subsystem, a data governance subsystem, a data quality control subsystem, a disease-specific database, and a dynamic maintenance subsystem working collaboratively, forming a closed loop through data feedback between subsystems. This invention uses the construction of a coronary heart disease-specific database as an example for illustration, but its architecture can be generalized to other diseases by replacing the disease-specific field dictionary and scenario rules.

[0019] In existing technologies, data collection in multicenter clinical studies has long relied on manual entry or single-point docking. The collection process requires manual review of paper or electronic medical records and entry of each item into the electronic data collection system, resulting in high labor costs, high transcription error rates, and errors in single fields that can easily contaminate downstream analysis. The collection efficiency is low during the follow-up phase, and patients' subsequent treatments are scattered across different medical institutions, regions, and time windows, making it difficult to monitor patients' subsequent conditions.

[0020] In addition, the acquisition channels are independent of each other and lack a data acquisition mechanism for individual data elements or individual patients. At the same time, it is difficult to automatically extract and uniformly map unstructured modal data such as surgical videos, heart sound waveforms, and image-based medical records to the same data dimension.

[0021] Natural language processing can also be used to extract structured clinical texts. However, this method relies on large-scale manually labeled training sets and task-specific model fine-tuning, resulting in long development cycles, weak cross-institutional generalization ability, and limited processing capabilities for image-based medical records. Zero-shot extraction schemes based on general large language models are usable for simple indicators, but their accuracy is low when faced with complex indicators that require inference based on time windows, symptom evolution, and risk factors. Furthermore, the heterogeneity of information systems in different institutions can lead to the same medical record being assigned different labels in different centers, impairing the reproducibility of multi-center studies.

[0022] More importantly, in existing solutions, the extraction rules (prompt words) are generally designed and fixed by humans according to their own standards, which can easily lead to inconsistent extraction standards, lack of transparency, and difficulty in reproducing the results.

[0023] In addition, to address the unreliability of the output of large language models, the results can be manually reviewed. However, manual review does not reduce efficiency in essence. Alternatively, the model self-reported confidence threshold can be used for filtering, but the model self-reported confidence has a weak correlation with the true accuracy rate, and there is a risk of systematic misjudgment in high-risk scenarios.

[0024] In addition, existing quality control schemes are mostly independent processes, and their judgment thresholds are not automatically generated in conjunction with the disease-specific data element model, making it difficult to integrate them with other aspects of database construction.

[0025] Furthermore, existing technologies focus on a single point in data collection, data governance, or data quality control. Even if the performance of one link is improved, the overall efficiency and quality of the process are still limited by other links and cannot be continuously optimized as the technology is used.

[0026] To address the problems in the prior art, embodiments of this application provide a disease data processing method, apparatus, and computer-readable storage medium.

[0027] The disease data processing method provided in the embodiments of this application will be introduced first below.

[0028] Figure 1 A flowchart illustrating a disease data processing method according to an embodiment of this application is shown. Figure 1 As shown, the method is applied to a disease data processing system. The disease data processing system predefines a data element schema, which includes at least two structured disease attributes, namely a clinical consequence severity level attribute and a decision attribute.

[0029] The disease data processing method provided in this application embodiment can be applied to a disease data processing system. The core design of the disease data processing system lies in its pre-built configurable data element schema. The data element schema is an abstract and structured definition of the concepts of disease data involved in the diagnosis and treatment of specific diseases. The data element schema contains at least two dimensions of structured disease attributes, including clinical consequence severity level attributes and decision attributes. The clinical consequence severity level attributes are used to identify the degree of impact of misjudgment of extracted fields on the patient's treatment path. For example, the coronary heart disease classification field, which is used to directly guide the antithrombotic treatment plan, has a severity level of high risk, while fields such as height and weight, which are basic physical signs, have a severity level of low risk. The decision attributes are used to determine whether the field directly enters the clinical treatment path or is only used as background reference information.

[0030] The method may include the following steps: S101: Collect multi-source case data through the acquisition channel scheduling engine of the disease data processing system.

[0031] Specifically, the acquisition channel scheduling engine in the disease data processing system performs the acquisition of multi-source data to obtain multi-source case data. During the acquisition process, the availability signal and the patient's compliance signal can be determined based on the data element to be acquired and the corresponding individual patient. The availability signal is used to distinguish whether the data element to be acquired is compatible, and the compliance signal is used to distinguish the patient's willingness to upload the data to be acquired.

[0032] For example, the acquisition channel scheduling engine can select acquisition channels and automatically degrade them according to a predetermined priority when the selected channels are unavailable. Its judgment logic includes, but is not limited to: for baseline data elements to be acquired, if the information system of this hospital can be connected, the structured fields are automatically synchronized through the information system adapter; if the information system cannot be connected (e.g., image-based medical records from other hospitals), the image-based medical record structuring module automatically structures the records taken by photos, scans, or screenshots. For multimodal data elements to be collected (such as surgical videos, heart sound waveforms, etc.), the data is collected and structured by the corresponding dedicated terminal through the multimodal intake and structuring module. For follow-up data, if the patient's compliance signal is high, the data can be automatically returned through the patient's self-management process in the chronic disease management application; if the patient's compliance signal is low, such as failure to return the data on time, the system will automatically downgrade and trigger an intelligent voice outbound call, which will then be transcribed and stored in the database.

[0033] Furthermore, the disease data processing system also includes an information system adapter, an image-based medical record structuring module, a multimodal intake and structuring module, and chronic disease management applications.

[0034] The information system adapter establishes a continuous connection with the hospital's information system through a predefined interface, automatically synchronizes structured fields in an event-driven manner (such as discharge, test report generation, and medical order changes), and automatically ignores fields unrelated to the specific disease by constraining the mapping table through a disease-specific field dictionary.

[0035] The image-based medical record structuring module uses a visual language model as its underlying engine. It performs single-step reasoning on de-identified medical record images, directly completing image recognition, layout understanding, and field location, and outputting structured text with layout semantics preserved, along with a timestamp.

[0036] The multimodal intake and structuring module receives non-textual data from dedicated terminals and extracts it into data to be processed within the data metadata model. Specifically, surgical videos can be extracted using a surgical stage and event annotation extractor to obtain structured annotations such as surgical stages and key events; heart sound waveforms can be extracted using an acoustic feature extractor to obtain corresponding acoustic feature indicators; and wearable and monitoring data can be extracted into corresponding indicators using a feature extractor.

[0037] Each modality uses its own extractor, but all are mapped to the same data element pattern and associated with the corresponding patient record. For cross-hospital surgical videos, a two-stage transmission architecture of local encoding and encryption at edge computing nodes and dedicated line backhaul can be used to allow video data to be accessed and annotated without leaving the original institution's encryption boundary.

[0038] The chronic disease management application is used for patient self-management and automatic feedback of follow-up indicators. The scheduling engine dynamically adjusts the trigger frequency of questionnaires and outbound calls based on the patient's risk stratification and the achievement of previously reported indicators. That is, the frequency is reduced for low-risk patients who have achieved the indicators, and the frequency is increased for high-risk patients who have not achieved multiple indicators.

[0039] This application embodiment can collect multi-source case data through the acquisition channel scheduling engine, thereby reducing acquisition costs and improving acquisition efficiency.

[0040] S102: Using the disease indicator extraction model and the corresponding indicator optimization rules, indicators are extracted from multi-source case data to obtain data to be treated that contains at least one disease indicator.

[0041] Specifically, through the disease indicator extraction model and the corresponding indicator optimization rules, structured indicator extraction is performed on multi-source case data to obtain data to be managed that includes at least one disease indicator from a data element pattern.

[0042] Among them, the disease indicator extraction model can be based on a pre-trained large language model to convert unstructured medical record text into structured fields that conform to the data element pattern. The indicator selection rule is a structured extraction instruction that has been iteratively corrected by the test set. Different indicators correspond to independent indicator selection rules.

[0043] Taking the classification indicators of coronary heart disease as an example, the indicator selection rules clearly define the judgment boundaries of the four types of classification, the processing logic of boundary scenarios such as troponin deficiency, and the constraint condition that prohibits the inference of classification based on medication history.

[0044] After the text enters the model, it is first normalized to complete the mapping of medical terminology synonyms and the unification of numerical units. Then, it is distributed to the corresponding extraction logic according to the field type. After extraction, the structured field entries are output to obtain the data to be processed.

[0045] Furthermore, the extraction optimization rules can be set to a semantic constraint format that is independent of the base, thereby adapting to different large model bases. This eliminates the need to rewrite the extraction optimization rules when switching bases, thereby eliminating the differences between annotators in the manual extraction process and making the results reproducible.

[0046] This application embodiment extracts multi-source case data through a disease indicator extraction model and corresponding indicator optimization rules, thereby obtaining structured data to be treated, and can extract multi-source case data into a unified data dimension.

[0047] S103: Determine the attribute value of at least one structured disease attribute of the data meta-pattern based on the data to be governed.

[0048] Specifically, by using the data element pattern in the disease data processing system, the corresponding structured disease attributes are matched field by field according to the data to be treated, thereby obtaining the attribute value of at least one structured disease attribute corresponding to the data to be treated.

[0049] For example, using the coronary heart disease classification field, the matching result shows a high level of clinical consequence severity and a yes decision attribute; using the height and weight field, the matching result shows a low level of clinical consequence severity and a no decision attribute.

[0050] This application embodiment enhances the disease data processing system's contextual understanding of the data to be managed by using attribute values, enabling the acquisition of structured disease attributes corresponding to specific diseases, thereby improving interpretability and intuitiveness.

[0051] S104: Determine the field risk level of the data to be governed through the data element model, and determine the quality control judgment rules corresponding to each field of the data to be governed based on the field risk level.

[0052] Specifically, determining the risk level of a field in data to be governed through the data element schema can be achieved by automatically calculating the risk level of the field corresponding to the data to be governed based on the structured attributes already registered within the data element schema.

[0053] In addition, the quality control judgment rules corresponding to each field of the data to be governed are determined according to the field risk level. The classification logic of the field risk level is related to the field attribute value. Therefore, it is not necessary to manually configure the corresponding field risk level judgment method for each new disease or each new field. It can be directly generated by the attribute definition in the data element pattern, realizing the automatic adaptation of the quality control strategy.

[0054] The mapping relationship between field risk levels and quality control judgment rules in this application embodiment is maintained uniformly through the data element model, which reduces the workload of manually configuring the threshold of quality control judgment rules for each field. When replacing specific diseases, the field dictionary and corresponding attributes are updated, which can directly complete the synchronization and adaptation, and improve the versatility of the disease data processing system.

[0055] S105: Based on the quality control judgment rules, the data to be governed is verified through at least one verification inference model and the corresponding indicator selection rules of the verification inference model.

[0056] Specifically, according to the quality control judgment rules, the data to be governed can be verified through at least one verification inference model and the corresponding indicator selection rules.

[0057] The verification inference model can be the current disease indicator extraction model, the disease indicator extraction model, or a large model base with a significantly different architecture from the disease indicator extraction model. It can also be a lightweight verification inference model. The verification results corresponding to each verification inference model can be obtained. The verification results can be combined according to the Boolean logic corresponding to the quality control judgment rules to determine whether the field passes the verification.

[0058] The minimum configuration is to use one validation inference model and the corresponding indicator optimization rules to validate the data to be governed. Further optimization configurations can be added on this basis, such as using two validation models and the corresponding indicator optimization rules to validate the data to be governed.

[0059] Furthermore, the selection of the base for the verification inference model can be adjusted according to the deployment scenario. For local deployment scenarios, different series of verification inference models with different parameter scales can be used, while for cloud deployment scenarios, different model services can be used. Ensuring architectural differences can achieve the effect of consistency verification.

[0060] This application embodiment uses at least one verification inference model to perform verification inference on multi-source case data, and uses quality control judgment rules to determine whether the data to be treated has passed verification, thereby improving the accuracy of the treatment results corresponding to the data to be treated and reducing labor costs.

[0061] S106: If the data to be governed passes verification, determine the governance result based on the data to be governed.

[0062] Specifically, if the data to be treated passes the verification corresponding to the quality control judgment rules, the data to be treated will be determined as the treatment result of the patient's multi-source case data.

[0063] The results of the treatment can be stored in a disease-specific database or used for downstream reasoning.

[0064] The embodiments of this application determine the treatment results by using the data to be treated, thereby improving the accuracy of the treatment results.

[0065] This application embodiment separates the business semantics of specific diseases from the disease data processing logic at the system architecture level through the data element pattern, making the disease data processing system universal across specific diseases. When converting specific diseases for disease data processing, the conversion can be completed by replacing the data element pattern definition, reducing the workload of reconstruction, lowering maintenance costs, and improving portability.

[0066] In some embodiments, the disease data processing system further includes an error attribution module, which also includes: An optimization test set is constructed based on the governance results, according to the preset optimization cycle and preset test set conditions. The test set is processed and optimized based on the index selection rules to obtain the test extraction results; If the performance value of any test result is lower than the preset performance threshold, the index selection rule corresponding to the attribute value is optimized through the restricted action space in the error attribution module to obtain the optimized index selection rule.

[0067] Specifically, the disease data processing system also includes error attribution. The error attribution module is used to determine the error type statistics and attribution of the indicator selection rules during the initialization phase of the disease data processing system. After going online, it can also maintain the performance of the disease data processing system.

[0068] Based on the pre-set optimization period and pre-set test set conditions, an optimized test set is constructed according to the governance results. The pre-set optimization period can be set according to the update frequency of disease-specific data, and the pre-set test set conditions are used to constrain the scope of the extracted test set. These conditions can be constrained according to the time range and the review status to ensure the correctness and validity of the data in the optimized test set.

[0069] In addition, the currently effective online indicator optimization rule is invoked to perform a complete inference on the optimized test set to obtain the test extraction results of the current indicator optimization rule under the new data distribution.

[0070] In addition, the error attribution module performs performance evaluation on the test extraction results. This can be achieved by comparing the test extraction results with the labeled test results of the gold standard annotation included in the optimized test set, and statistically analyzing quantitative performance indicators such as accuracy, recall, or F1 score of each indicator on the current optimized test set.

[0071] If the performance value of the test sampling result of any indicator is lower than the preset performance threshold, for example, if the accuracy rate drops from 98% to 92% in the baseline period, the indicator selection rule is determined to have experienced performance drift.

[0072] For indicators that fail to meet performance standards, the indicator selection rules are re-optimized in a targeted manner. The error attribution module performs error type attribution analysis on the error output of the indicator on the current optimization test set, identifies the main distribution of the current error, and the restricted action space selects the repair action that matches the error type from the preset repair action set based on the error signature obtained from the attribution. The repair action is applied to the corresponding part of the current indicator selection rule to obtain the repaired indicator selection rule, and the performance value of the repaired indicator selection rule is re-evaluated on the optimization test set.

[0073] If the performance value recovers to above the preset performance threshold, the indicator optimization rule is updated according to the repaired rule. If the performance value does not recover to the preset performance threshold, the indicator optimization rule is iterated until the performance value exceeds the preset performance threshold or the preset maximum number of iterations is reached, and the optimized indicator optimization rule is obtained.

[0074] The embodiments of this application realize the iteration of the indicator optimization rules, eliminating the need for manual periodic checks for performance degradation and manual adjustment of extraction rules. The periodic construction of optimization test sets can promptly capture performance degradation caused by data heterogeneity and time drift. The combination of error attribution and limited action space improves the targeting and controllability of the indicator optimization rules.

[0075] In some embodiments, the quality control judgment rules corresponding to each field of the data to be governed are determined based on the field risk level, including: When the field risk level is low, the quality control judgment rule is determined to be that the data to be governed and the verification output of at least one verification inference model are consistent. When the field risk level is medium risk, the quality control judgment rule is determined to be that the data to be treated and the verification output of at least two verification inference models are consistent. When the risk level of a field is high, the quality control judgment rule is set as the data to be governed and the validation output of all validation inference models are consistent.

[0076] Specifically, based on the clinical consequence severity level attribute and decision attribute of each field in the data element model, the fields are divided into high-risk, medium-risk, and low-risk levels, and different verification pass conditions are assigned to each risk level.

[0077] Among them, the quality control judgment rule represents the verification conditions for the data to be governed. If the verification conditions are met, the data to be governed will be determined as the governance result. That is, the data to be governed and the output of the verification inference model must satisfy a consistency relationship in order to be judged as passing the quality control rule.

[0078] When the risk level of a field is low, the corresponding quality control judgment rule is that the data to be governed must be consistent with the verification output of at least one verification inference model. That is, if the output of any one of the three verification inference models is consistent with the data to be governed, it means that the field has passed quality control.

[0079] Low-risk fields carry less weight in clinical decision-making, and the consequences of misjudgment are manageable. Therefore, more lenient quality control criteria can be applied to achieve higher processing efficiency. Examples of low-risk fields include patient height, weight, and admission time.

[0080] When a field's risk level is medium risk, the corresponding quality control judgment rule is that the field can be judged to have passed quality control only if the data to be treated is the same as the verification output of at least two verification inference models.

[0081] Among them, the medium-risk level fields usually involve the selection of conventional treatment plans or the stratification of disease severity, such as the determination of the type of antithrombotic regimen or the assessment of blood lipid target achievement. Misjudgment of medium-risk level fields will not directly lead to a fundamental error in the treatment direction, but may affect the follow-up strategy or the fine-tuning of drug dosage. Therefore, it is necessary to reach a consensus on most validation signals before release to reduce the risk of deviation of a single validation model on the boundary case.

[0082] When a field has a high risk level, the corresponding quality control judgment rule is that the field can only pass quality control if the data to be treated is the same as the verification output of all verification inference models.

[0083] High-risk fields directly influence the diagnosis or key treatment path selection in clinical decision-making, such as the classification of coronary heart disease and the judgment of post-stent complications. Misjudgment may lead to patients receiving inappropriate treatment plans and irreversible clinical consequences. Therefore, if there is an inconsistency between the output of any validation inference model and the data to be treated, it means that the output of that field is uncertain and will not be directly allowed into the database. Instead, it will be routed to manual review through a selective rejection router.

[0084] The embodiments of this application, through hierarchical quality control judgment rules, eliminate the need for separate configuration of quality control logic when adding new fields. They can be synchronously adapted with the expansion of the disease-specific field dictionary, thereby reducing the workload of manual review while maintaining accuracy.

[0085] In some embodiments, the verification inference model includes a disease indicator extraction model, a disease indicator verification model, and a link inference model, and the verification output includes first verification data, second verification data, and logical review verification data. The first validation data is obtained by extracting secondary indicators from multi-source case data using the disease indicator extraction model and the corresponding indicator optimization rules. The second validation data is obtained by extracting indicators from multi-source case data through the disease indicator validation model and the corresponding indicator optimization rules. The disease indicator extraction model obtains the inference chain information of the data to be treated by reasoning according to the indicator selection rules. The inference chain information is then logically reviewed by the link inference model to obtain logical review verification data. The logical review includes at least one of logical break, skip step and contradiction.

[0086] Specifically, the verification inference model includes a disease indicator extraction model, a disease indicator verification model, and a link inference model. The disease indicator extraction model is a secondary call performed during the verification phase, inferring again under the same input and parameter configuration to obtain the first verification data.

[0087] The second call indicates that the disease indicator extraction model is a separate inference session independent of the governance phase. There is no cache reuse or result reuse between the two inferences, thus ensuring that the consistency difference between the two outputs can reflect the stability of the inference process itself.

[0088] The disease indicator validation model differs substantially from the disease indicator extraction model in terms of architecture or inference method. For example, the disease indicator extraction model adopts a dense Transformer architecture with a large parameter scale, while the disease indicator validation model can choose a sparse expert hybrid architecture with a smaller parameter scale, or choose a similar architecture model based on different pre-training strategies.

[0089] By using a disease indicator validation model and its corresponding indicator optimization rules, independent indicator extraction is performed on the same multi-source case data to obtain second validation data.

[0090] The link inference model is used to verify the intermediate step sequence generated by the disease indicator extraction model during the inference process of the governance phase. When calling the disease indicator extraction model to infer the data to be governed, the output parameters of the link inference model can be configured to include backtracking inference chain information while returning structured fields.

[0091] Among them, the inference chain information records the deductive relationship between the evidence fragments identified by the disease indicator extraction model from the input text, the extraction and selection rules cited, and the intermediate judgment conclusions.

[0092] The link reasoning model receives reasoning chain information, performs logical review on it, and obtains logical review verification data. The logical review includes at least whether there is a logical break between the premise and the conclusion in the reasoning chain information, that is, jumping directly to the conclusion without the necessary intermediate judgments, whether there are skipped steps, that is, whether the evidence on which a certain intermediate conclusion is based cannot be fully supported by the input text, and whether there are self-contradictions among the judgments in the reasoning chain information. For example, the same medical record is classified into one subtype and another mutually exclusive subtype in the same reasoning chain.

[0093] This application embodiment verifies the data to be governed from three dimensions: output stability, cross-base consistency, and the rationality of reasoning logic, through quality control judgment rules. The detected anomalies can be directly correlated to specific error types. It can not only determine whether a field is qualified, but also provide error signals for the optimization of extraction and selection rules. It can apply different strengths of verification consensus requirements to fields with different risk levels, thereby balancing the accuracy of data entry and the cost of manual review.

[0094] In some embodiments, the disease data processing system further includes a selective rejection router and an error attribution module; After validating the data to be governed based on the quality control judgment rules and through at least one validation inference model and the corresponding indicator optimization rules of the validation inference model, the process also includes: By selectively discarding the data, the router routes the data to be governed to the verification unit if the data fails verification. The verification results corresponding to the data to be treated are obtained based on the verification unit; The verification result is fed back to the error attribution module as an error signal, so that the error attribution module can optimize the index selection rule corresponding to the attribute value based on the error signal and through the limited action space in the error attribution module, and obtain the optimized index selection rule.

[0095] Specifically, the disease data processing system also includes a selective rejection router and an error attribution module.

[0096] Among them, the selective rejection router is a decision distribution node. Its input end receives the verification status of each field output by the quality control verification process, that is, the consistency judgment result between the verification data output by each verification inference model and the data to be governed, and the risk level corresponding to the field. Its output end directs the data flow to different paths according to the preset Boolean combination rules.

[0097] The error attribution module performs error type attribution during the rule iteration optimization in the governance phase, and secondly, in the feedback loop of the actual application phase, it receives the review results and triggers the re-optimization of the rules.

[0098] After completing all verification reasoning for the data to be governed during the quality control verification phase, the pass or fail status of each verification signal is sent to the selective rejection router, so that the selective rejection router can make a decision on whether the data to be governed meets the release conditions based on the quality control judgment rules corresponding to the risk level of the field.

[0099] Low-risk fields pass verification when at least one verification signal is consistent, medium-risk fields pass verification when at least two verification signals are consistent, and high-risk fields pass verification when all verification signals are consistent. If verification fails, meaning the data to be governed fails to meet the verification consensus strength required for its risk level, the selective rejection router marks the data to be governed and the failed verification status and routes it to the review unit.

[0100] The review unit can be a manual review station, which is equipped with a graphical interface. The graphical interface simultaneously displays the multi-source case data of the data to be treated, the output results of each validation inference model, the abnormal markers of each validation signal, and the definition information of the data element pattern of the field. The reviewer reviews the data based on the above data, obtains the review results, and enters the review results into the disease database.

[0101] Simultaneously, the review results are encapsulated as error signals and fed back to the error attribution module. The error attribution module then sends the data to be addressed, the verification data of each verification model, the review results given by the reviewers, and the specific verification anomaly types that trigger selective rejection routing to the cumulative buffer of the error attribution module.

[0102] The error signal is a structured record that includes all the contextual information for subsequent attribution analysis, enabling the error attribution module to accurately determine the error type based on the error signal and then update the indicator selection rules.

[0103] When the number of error signals of the same indicator and the same error type in the cumulative buffer reaches a preset error threshold, the error attribution module initiates a re-optimization process for the indicator's selection rule. Specifically, the error attribution module clusters the accumulated error signals, performs error type attribution analysis on the error output of the indicator in the clustering results, identifies the main distribution of the current error, and selects a repair action matching the error type from the preset repair action set based on the error signature obtained from the attribution in the restricted action space. The repair action is then applied to the corresponding part of the current indicator selection rule to obtain the repaired indicator selection rule, and the performance value of the repaired indicator selection rule is re-evaluated on the optimization test set.

[0104] If the performance value recovers to above the preset performance threshold, the indicator optimization rule is updated according to the repaired rule. If the performance value does not recover to the preset performance threshold, the indicator optimization rule is iterated until the performance value exceeds the preset performance threshold or the preset maximum number of iterations is reached, and the optimized indicator optimization rule is obtained.

[0105] This application embodiment achieves hierarchical dispatch of review tasks through a selective rejection router, reducing the manpower consumption caused by indiscriminate full review, ensuring the processing priority of high-risk diagnostic fields, and the review results are fed back to the error attribution module in the form of error signals, so that the results of manual review are transformed into the input of the extraction and optimization rule iteration, forming a positive loop, thereby improving the adaptability and sustainability of the extraction and optimization rules.

[0106] In some embodiments, the review result is fed back to the error attribution module as an error signal, so that the error attribution module optimizes the index selection rule corresponding to the attribute value based on the error signal through the limited action space in the error attribution module, and obtains the optimized index selection rule, including: The error attribution module encapsulates the review results into error signal records, which include multi-source case data, disease indicators, data to be treated, treatment results, and failed verification methods corresponding to the review results. Accumulate the corresponding error signal records for each disease indicator separately, and perform cluster statistics on the error signal records of each disease indicator that failed the validation method according to the preset clustering period to obtain the cluster statistics value. When the clustering statistics of the same disease indicator reach the preset iterative optimization threshold, the indicator selection rule is iteratively optimized within the limited action space to obtain the optimized indicator selection rule.

[0107] Specifically, after the review unit completes the review of the rejected medical records, the review result and its associated context information are fed back to the error attribution module for encapsulation, resulting in a structured error signal record. The error signal record includes the original multi-source case data corresponding to the review result, the disease indicator used to identify the error signal, the data to be treated, the treatment result given by the review unit, and the verification method that the error signal record failed, such as being marked as failing heterogeneous model consistency verification or failing logical review.

[0108] Error signals are recorded and categorized according to their respective disease indicators. Samples for different disease indicators are independent of each other and do not interfere with each other. Full-indicator clustering statistics are performed according to the preset clustering period. During the clustering process, statistics are collected based on the type and method of failure to pass validation, and the clustering statistics of each type of error under each disease indicator are obtained.

[0109] The cluster statistics include the cumulative number of single-class errors and the proportion of such errors to the total error of the disease indicator.

[0110] For example, in the error sample pool of the coronary heart disease classification index, 11 valid records were stored in a single week. After clustering, 7 of them belonged to the missing value boundary misjudgment class. The corresponding failed verification methods all contained inference chain review anomalies, and the clustering statistics of the failed verification methods were 7 cases.

[0111] For each type of error, a corresponding preset iterative optimization threshold is set. The preset iterative optimization threshold is related to the risk level of the field. The trigger threshold is lower for high-risk fields and higher for low-risk fields. When the clustering statistics of a certain type of error under the same disease indicator reach the corresponding preset iterative optimization threshold, the targeted iterative optimization process of that disease indicator is triggered.

[0112] The iterative optimization process is confined to a limited action space. The corresponding repair action is matched according to the error type. For misjudgment of missing value boundaries, a missing value judgment branch and negative constraints are added. For missed detection of negative expression, the identification rules and counterexamples of negative sentence are supplemented. For unit confusion, a unit unification verification step is added.

[0113] The performance value of the corrected indicator selection rule is re-evaluated on the optimized test set. If the performance value recovers to above the preset performance threshold, the corrected indicator selection rule is updated. If the performance value does not recover to the preset performance threshold, the indicator selection rule is iterated until the performance value exceeds the preset performance threshold or the preset maximum number of iterations is reached, thus obtaining the optimized indicator selection rule.

[0114] This application embodiment realizes directional optimization based on error signal records. The directional optimization process is constrained by preset clustering period and preset iteration optimization threshold, which reduces the performance fluctuation caused by occasional errors in individual cases and frequent updates of the extraction optimization rules, and improves the accuracy and efficiency of the extraction optimization rule iteration.

[0115] In some embodiments, the process of determining the index preference rule includes: The indicator test set is structured and extracted using extraction rules to obtain the test extraction results. The extraction rules include operational decision rules, negative constraints, output format, and test samples. By comparing the sampling results of the test with the labeled results of the indicator test set, the error type of each indicator and the error signature corresponding to the error type are determined. The error type includes at least one of the following: unit confusion, missed negative statement, value illusion, boundary condition misjudgment and field crosstalk. Based on the error signature, the extraction rules are repaired within a restricted action space, wherein the restricted action space includes at least one of adding counterexamples, tightening constraints, splitting fields, and adjusting the decision threshold. Based on the repaired extraction rules, the indicator test set is re-extracted in a structured manner to update the test extraction results until the test error between the extraction results and the labeled results of the extraction rules on the indicator test set reaches the preset termination condition, and the extraction rules are determined as the indicator optimization rules.

[0116] Specifically, initial extraction rules are developed for each disease-specific indicator. These rules include operational decision rules, negative constraints, output format, and test samples.

[0117] Among them, the operational judgment rules clearly define the extraction logic and judgment criteria of the fields. Taking the serum low-density lipoprotein cholesterol target status field as an example, the extraction rules specify the target value boundaries for patients with different risk levels, the conversion relationship between different units, and the priority of values ​​from multiple test reports.

[0118] Negative constraints define the decision logic that should not be triggered, such as not inferring the current test value based on the description of the patient's medical history, and not using estimated values ​​to replace clearly recorded test results.

[0119] The output format is constrained to a structured data format, including field values, original evidence location, corresponding units, and other information, to reduce parsing bias in free text output.

[0120] The test sample represents a small number of typical boundary cases selected manually, with the extraction rules embedded to quickly determine the judgment boundary.

[0121] The indicator test set was structured and extracted using extraction rules to obtain the test extraction results. The indicator test set is a gold standard dataset with manual annotation covering multiple cooperating medical institutions and various medical record writing styles. The sample size was adjusted according to the complexity of the fields. All annotation results were cross-checked and confirmed by at least two specialist physicians.

[0122] The test sampling results are compared with the gold standard annotation results to determine the error types of each indicator and the corresponding structured error signatures. Among them, the error types include at least the following: unit confusion, that is, failing to correctly identify the unit identifier in the medical record and directly outputting values ​​with different units; missed negative statements, that is, failing to identify negative signs or negative test results clearly recorded in the medical record and misjudging them as positive or missing fields; value illusion, that is, generating non-existent values ​​or conclusions when the corresponding fields are not clearly recorded in the medical record; boundary condition misjudgment, that is, failing to perform judgment according to the standard boundary when the values ​​or symptoms are at the critical position of the judgment threshold; and field crosstalk, that is, incorrectly assigning information from adjacent fields to the current target field.

[0123] Taking the low-density lipoprotein cholesterol target status field as an example, after full comparison, 8 out of 12 erroneous samples were classified as boundary condition misjudgments. These were concentrated in the critical samples within 0.1 mol / L above and below the corresponding target threshold. The corresponding error signature marked the error type as boundary condition misjudgment, with a statistical frequency of 8. The typical characteristic was that the test value was close to the target threshold. The high-incidence scenario was the test result summary paragraph in the discharge summary.

[0124] Furthermore, based on the error signature matching corresponding repair actions, all repair operations are limited to execution within a predefined restricted action space. This restricted action space includes at least the following: adding counterexamples, i.e., supplementing typical samples of the corresponding error type into the example library of extraction rules to enhance the extraction rules' ability to identify boundary scenarios; tightening constraints, i.e., adding more explicit limiting clauses to operational rules to narrow the fuzzy judgment range; splitting fields, i.e., to address the field crosstalk problem caused by complex logic, splitting a single complex field into multiple subfields for separate extraction and then aggregation; and adjusting the judgment threshold, i.e., modifying the critical judgment criteria for numerical fields to match clinical guidelines.

[0125] The repaired extraction rules are reloaded into the indicator extraction module, and structured extraction is performed again on the same indicator test set. The test extraction results are updated, the error level of this extraction is recalculated, and compared with the results of the previous iteration. The iteration continues until the preset termination condition is met.

[0126] Among them, the preset termination condition can be that the extraction accuracy reaches the preset performance threshold, or the accuracy improvement of consecutive preset rounds of iteration is less than the preset iteration threshold, which is judged as performance convergence. If either termination condition is met, the current version of the extraction rule is determined as the indicator optimization rule of the corresponding indicator and is stored in the rule base simultaneously.

[0127] This application's embodiments are based on manual initial rules, completing error attribution, rule repair, and effect verification, shortening the development cycle of disease-specific field extraction rules, enabling structured error location through error signatures, and ensuring the logical controllability and auditability of the extraction and optimization rules through the limitation of repair actions, thereby reducing unpredictable deviations caused by open rewriting.

[0128] In some embodiments, after determining the governance outcome based on the data to be governed, the method further includes: Identify the source identifier for each piece of data to be processed. The source identifier includes information about the source institution and the location information of multi-source case data. Determine the governance path information for multi-source case data. The governance path information includes determining the indicator selection rules for the data to be governed and the model version number of the disease indicator extraction model. Determine the quality control status information of the governance results. The quality control status information includes the verification results and the review status. The review status includes the direct output status and the reviewed status. The source identifier, treatment path information, quality control status information, and treatment results are stored in the corresponding disease-specific database.

[0129] Specifically, each treatment result is identified by its corresponding source identifier. The source identifier is the original location marker of the treatment result before it enters the disease data processing system. The location information of multi-source case data is the original record locator of multi-source case data in the internal system of its source institution, such as the combination of hospital number and visit number, the unique report number of the test report, or the access path of the image sequence.

[0130] The governance path information, which refers to the processing parameters of the multi-source case data when it passes through the governance subsystem, includes at least two aspects: the version identifier of the indicator selection rule called when extracting the field and the model version number of the disease indicator extraction model used when performing inference.

[0131] The quality control status information is determined by the quality control link of the disease data processing system, including the verification results and the review status. The verification results record the verification pass status of the governance data on each verification inference model. The review status is used to distinguish how the governance data obtained the governance results. The direct output status indicates that the governance data meets the quality control verification rules and is directly entered into the database without the review unit. The reviewed status indicates that the governance data is routed to the review unit by the selective rejection router, and the reviewer determines its review result.

[0132] Source identifiers, governance path information, quality control status information, and governance results are written into the disease-specific database in a unified storage association format, so that when querying any governance result, its corresponding full-chain metadata can be returned together.

[0133] In this embodiment, the source identifier, governance path information, and quality control status information of the governance results are recorded simultaneously upon entry into the database, forming a complete audit trail of the lifecycle of each governance result. This ensures that all data in the disease-specific database corresponds to a traceable metadata chain. Simultaneously, the recording of source identifiers and governance path information provides the operational foundation for the dynamic maintenance subsystem, reducing the deviation in drift detection results caused by mixing data from different versions of indicator optimization rules into the same test set.

[0134] Figure 2 This is a schematic diagram of the architecture of a disease data processing system provided in an embodiment of this application, such as... Figure 2 As shown, the system includes a disease data processing system 100, a data acquisition subsystem 110, an acquisition channel scheduling engine 115, an information system adapter 116, an image-based medical record structuring module 117, a multimodal intake and structuring module 118, a chronic disease management application 119, a data governance subsystem 120, an indicator extraction module (normalization, self-correction, and diversion extraction) 121, a scenario rule knowledge base 123, a multi-base adaptation module 124, an error attribution module 125, a restricted action space iteration module 126, a convergence locking module 127, a data quality control subsystem 130, a heterogeneous model consistency verification module 131, a homogeneous sampling stability verification module 132, an inference chain review module 133, a selective rejection router 134, a review unit 135, a disease-specific database 140, a dynamic maintenance subsystem 150, a test set construction module 151, a drift detection module 152, and a directional re-optimization trigger module 153.

[0135] In some embodiments, the disease data processing system is deployed in the local data center of a tertiary-level cardiovascular specialty hospital, equipped with several inference servers with graphics processors for localized operation of the large language model serving as the main foundation of the governance subsystem; an auxiliary model is connected via an interface as a heterogeneous validation model for the quality control subsystem; and a lightweight model is deployed locally as an inference chain review model. All patient identification information is de-identified at the collection end, and only the de-identified data enters the governance and quality control stages, meeting the compliance requirements of the Personal Information Protection Law and the Good Clinical Practice for Drug Clinical Trials.

[0136] The scheduling engine 115 maintains a data element-channel priority table and a patient compliance status table. For each data element, available channels are tried sequentially according to priority: the information system adapter 116 connects to the hospital's information system using the medical information exchange standard and subscribes to events such as admission, laboratory report generation, imaging report generation, medical order changes, and discharge, triggering a synchronization for each event; when the target field cannot be obtained through the information system, it is downgraded to the image-based medical record structure module 117. For follow-up data elements, the scheduling engine selects automatic feedback from the chronic disease management application 119 or downgrades to trigger intelligent voice outbound calls based on the patient's compliance status. Table 1 is a schematic diagram of the dynamic adjustment of follow-up trigger frequency according to risk level.

[0137] Table 1. Schematic diagram of dynamic adjustment of follow-up trigger frequency according to risk level The indicator extraction module 121, targeting the logical field of coronary artery disease classification, invokes a structured instruction template that includes role settings, operational rules, negative constraints, output format, and few-sample examples. The error attribution module 125 compares the extraction results with the gold standard on the test set. For example, it finds two common errors in this field: first, missing troponin data incorrectly classifies it as non-ST-segment elevation myocardial infarction; second, recurrent chest pain after previous stent placement is directly classified as unstable angina. The restricted action space iteration module 126 selects repair actions accordingly: tightening constraints for the former (adding a missing value decision branch), and adding targeted counterexamples and restating the decision order for the latter. After each repair, the field is retested on the test set; when performance reaches a threshold, the convergence locking module 127 locks the optimal rule for this field. This process is executed separately for each indicator.

[0138] An example of a metric optimization rule structure is as follows: [Role] Coronary heart disease specialist data extraction assistant.

[0139]

Task

[0140]

Operational Rules

[0141] [Negative Constraint] N1: Subtype cannot be inferred based on medication history; N2: When troponin data is missing, it should not be automatically presumed as NSTEMI; N3: Chest pain recurring after previous stent placement should not be directly diagnosed as UA and must be reassessed according to R1-R4.

[0142] [Output Format] Strict JSON: {label, evidence, rationale}; [Examples of a small sample size] (Several typical borderline medical records).

[0143] The selective rejection router 134 applies different Boolean combination rules according to the risk level of the field. Table 2 is a schematic table of Boolean combination rules. The risk level is automatically assigned by the data element pattern.

[0144] Table 2. Schematic diagram of Boolean combination rules Taking the coronary heart disease classification field as an example, specific operational examples of upward and downward coupling of the data quality control subsystem 130 are given to more clearly illustrate the feasibility of this bidirectional coupling relationship.

[0145] In the disease data processing system, the quality control subsystem does not require engineers to manually configure the rejection threshold for each field. Instead, it is automatically derived from the field attributes pre-registered in the data element schema (i.e., the field dictionary constrained by the data acquisition subsystem).

[0146] The data element schema registers several attributes for each field, such as: severity of clinical consequences, whether it enters the treatment decision path, whether it is a composite logical field, and the downstream impact range of missing or misjudged data. The quality control subsystem reads these attributes, automatically calculates the risk level of the field according to predefined mapping rules, and automatically selects the corresponding signal Boolean combination rule and rejection threshold accordingly. Taking the coronary heart disease classification field as an example, Table 3 is a schematic table of the mapping derivation process.

[0147] Table 3. Schematic diagram of the mapping derivation process In contrast, for fields like "admission time," which are registered as having low clinical severity, not entering the treatment decision path, and are non-logical, the same mapping rule automatically derives "low risk" and automatically selects the strictest rejection condition S1 ∩ S2 ∩ S3 (rejection only when all three signals report abnormalities), minimizing unnecessary manual review. Therefore, the quality control threshold setting is entirely driven by the disease-specific data element model, eliminating the need for manual threshold configuration for new fields or diseases. In other words, as the field dictionary is replaced or expanded with the disease, the corresponding quality control threshold is automatically generated. This is upward coupling.

[0148] Continuing the previous example, for the coronary heart disease classification medical records that were selectively rejected by router 134 and sent to review unit 135, the correct label given by manual review, along with the signal status of each verification module at the time of rejection, is structured and encapsulated into an error signal record, which is automatically fed back to the error attribution module 125 of data governance subsystem 120. This error signal record includes at least: field name, main model output label, manually confirmed correct label, the signal that triggered rejection (which of S1 / S2 / S3), and original evidence text fragments. The error attribution module accumulates these records by field and performs periodic clustering. Once the cumulative count of a certain error type exceeds a preset threshold, the restricted action space iteration module 126 is automatically triggered to perform targeted re-optimization of that field. Table 4 is a complete backflow triggering link diagram.

[0149] Table 4. Complete Backflow Trigger Path Diagram During the continuous operation of the disease data processing system after its launch, the same error signal also flows back to the dynamic maintenance subsystem 150: when the drift detection module 152 finds that the performance of this field has dropped below the threshold on the newly constructed multi-institution test set, the targeted re-optimization trigger module 153 also restarts the above-mentioned error attribution and limited action space iteration process only for this field, without affecting other fields. Therefore, each anomaly detected in the quality control process is no longer a one-time manual correction, but rather serves as an error signal driving the self-improvement of the indicator selection rules, continuously improving the system's accuracy with use and resulting in optimized indicator selection rules, i.e., downward coupling.

[0150] Figure 3 This is a schematic diagram of the structure of a disease data processing device provided in an embodiment of this application. Figure 3 As shown, the disease data processing device is applied to a disease data processing system and may include an acquisition module 310, an extraction module 320, a structuring module 330, a quality control module 340, a verification module 350, and an output module 360.

[0151] The acquisition module 310 is used to acquire multi-source case data through the acquisition channel scheduling engine of the disease data processing system; The extraction module 320 is used to extract indicators from multi-source case data through the disease indicator extraction model and the indicator optimization rules corresponding to the disease indicator extraction model, so as to obtain data to be treated containing at least one disease indicator. Structured module 330 is used to determine the attribute value of at least one structured disease attribute of the data element pattern based on the data to be governed; The quality control module 340 is used to determine the field risk level of the data to be governed through the data element pattern, and to determine the quality control judgment rules corresponding to each field of the data to be governed based on the field risk level. The verification module 350 is used to verify the data to be governed based on the quality control judgment rules, through at least one verification inference model and the index selection rules corresponding to the verification inference model; The output module 360 ​​is used to determine the governance result based on the data to be governed if the data to be governed passes the verification.

[0152] In some embodiments, the disease data processing system further includes an error attribution module, which is also used for: An optimization test set is constructed based on the governance results, according to the preset optimization cycle and preset test set conditions. The test set is processed and optimized based on the index selection rules to obtain the test extraction results; If the performance value of any test result is lower than the preset performance threshold, the index selection rule corresponding to the attribute value is optimized through the restricted action space in the error attribution module to obtain the optimized index selection rule.

[0153] In some embodiments, the quality control module 340 determines the quality control judgment rules corresponding to each field of the data to be addressed based on the field risk level, for the purpose of: When the field risk level is low, the quality control judgment rule is determined to be that the data to be governed and the verification output of at least one verification inference model are consistent. When the field risk level is medium risk, the quality control judgment rule is determined to be that the data to be treated and the verification output of at least two verification inference models are consistent. When the risk level of a field is high, the quality control judgment rule is set as the data to be governed and the validation output of all validation inference models are consistent.

[0154] In some embodiments, in the quality control module 340, the verification reasoning model includes a disease indicator extraction model, a disease indicator verification model, and a link reasoning model, and the verification output includes first verification data, second verification data, and logical review verification data. The first validation data is obtained by extracting secondary indicators from multi-source case data using the disease indicator extraction model and the corresponding indicator optimization rules. The second validation data is obtained by extracting indicators from multi-source case data through the disease indicator validation model and the corresponding indicator optimization rules. The disease indicator extraction model obtains the inference chain information of the data to be treated by reasoning according to the indicator selection rules. The inference chain information is then logically reviewed by the link inference model to obtain logical review verification data. The logical review includes at least one of logical break, skip step and contradiction.

[0155] In some embodiments, the disease data processing system further includes a selective rejection router and an error attribution module; After the verification module 350 verifies the data to be governed based on the quality control judgment rules, using at least one verification inference model and the corresponding indicator optimization rules, it is also used for: By selectively discarding the data, the router routes the data to be governed to the verification unit if the data fails verification. The verification results corresponding to the data to be treated are obtained based on the verification unit; The verification result is fed back to the error attribution module as an error signal, so that the error attribution module can optimize the index selection rule corresponding to the attribute value based on the error signal and through the limited action space in the error attribution module, and obtain the optimized index selection rule.

[0156] In some embodiments, the verification module 350 feeds back the verification result as an error signal to the error attribution module, so that the error attribution module, based on the error signal, optimizes the index selection rule corresponding to the attribute value through the limited action space in the error attribution module, and obtains the optimized index selection rule, which is used for: The error attribution module encapsulates the review results into error signal records, which include multi-source case data, disease indicators, data to be treated, treatment results, and failed verification methods corresponding to the review results. Accumulate the corresponding error signal records for each disease indicator separately, and perform cluster statistics on the error signal records of each disease indicator that failed the validation method according to the preset clustering period to obtain the cluster statistics value. When the clustering statistics of the same disease indicator reach the preset iterative optimization threshold, the indicator selection rule is iteratively optimized within the limited action space to obtain the optimized indicator selection rule.

[0157] In some embodiments, the process of determining the index preference rule includes: The indicator test set is structured and extracted using extraction rules to obtain the test extraction results. The extraction rules include operational decision rules, negative constraints, output format, and test samples. By comparing the sampling results of the test with the labeled results of the indicator test set, the error type of each indicator and the error signature corresponding to the error type are determined. The error type includes at least one of the following: unit confusion, missed negative statement, value illusion, boundary condition misjudgment and field crosstalk. Based on the error signature, the extraction rules are repaired within a restricted action space, wherein the restricted action space includes at least one of adding counterexamples, tightening constraints, splitting fields, and adjusting the decision threshold. Based on the repaired extraction rules, the indicator test set is re-extracted in a structured manner to update the test extraction results until the test error between the extraction results and the labeled results of the extraction rules on the indicator test set reaches the preset termination condition, and the extraction rules are determined as the indicator optimization rules.

[0158] In some embodiments, after determining the governance result based on the data to be governed, the output module 360 ​​is further configured to: Identify the source identifier for each piece of data to be processed. The source identifier includes information about the source institution and the location information of multi-source case data. Determine the governance path information for multi-source case data. The governance path information includes determining the indicator selection rules for the data to be governed and the model version number of the disease indicator extraction model. Determine the quality control status information of the governance results. The quality control status information includes the verification results and the review status. The review status includes the direct output status and the reviewed status. The source identifier, treatment path information, quality control status information, and treatment results are stored in the corresponding disease-specific database.

[0159] Figure 4 A schematic diagram of the hardware structure of the terminal device provided in an embodiment of this application is shown.

[0160] The terminal device may include a processor 401 and a memory 402 storing computer program instructions.

[0161] Specifically, the processor 401 may include a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0162] Memory 402 may include mass storage for data or instructions. For example, and not limitingly, memory 402 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. In one instance, memory 402 may include removable or non-removable (or fixed) media, or memory 402 may be non-volatile solid-state memory. Memory 402 may be internal or external to the integrated gateway disaster recovery device.

[0163] In one instance, memory 402 may be read-only memory (ROM). In one instance, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or flash memory, or a combination of two or more of these.

[0164] Memory 402 may include read-only memory (ROM), random access memory (RAM), disk storage media device, optical storage media device, flash memory device, electrical, optical, or other physical / tangible memory storage device. Therefore, generally, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to the first aspect of this disclosure.

[0165] The processor 401 reads and executes computer program instructions stored in the memory 402 to achieve... Figure 1 The disease data processing method in the illustrated embodiment.

[0166] In one example, the terminal device may also include a communication interface 403 and a bus 404. Wherein, as... Figure 4 As shown, the processor 401, memory 402, and communication interface 403 are connected through bus 404 and complete communication with each other.

[0167] The communication interface 403 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.

[0168] Bus 404 includes hardware, software, or both, that couples components of an online data traffic metering device together. For example, and not as a limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 404 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, this application contemplates any suitable bus or interconnect.

[0169] Furthermore, in conjunction with the disease data processing methods in the above embodiments, this application embodiment can provide a computer-readable storage medium for implementation. This computer-readable storage medium stores computer program instructions; when executed by a processor, these computer program instructions implement any of the disease data processing methods in the above embodiments.

[0170] This application also provides a computer program product, including a computer program that, when executed by a processor, implements any of the disease data processing methods described in the above embodiments.

[0171] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0172] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, read-only memory (ROM), flash memory, erasable read-only memory (EROM), floppy disks, compact disc read-only memory (CD-ROM), optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.

[0173] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0174] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.

[0175] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.

Claims

1. A method for processing disease data, characterized in that, The method is applied to a disease data processing system, which predefines a data element schema. The data element schema includes at least two structured disease attributes, namely, a clinical consequence severity attribute and a decision attribute. The method includes: Multi-source case data are collected through the acquisition channel scheduling engine of the disease data processing system. By using the disease indicator extraction model and the indicator optimization rules corresponding to the disease indicator extraction model, indicators are extracted from the multi-source case data to obtain data to be treated that contains at least one disease indicator. Based on the data to be treated, determine the attribute value of at least one of the structured disease attributes of the data element pattern; The risk level of the fields in the data to be governed is determined by the data element pattern, and the quality control judgment rule corresponding to each field in the data to be governed is determined based on the field risk level. Based on the quality control judgment rules, the data to be treated is verified through at least one verification inference model and the index optimization rules corresponding to the verification inference model; If the data to be treated passes verification, the treatment result is determined based on the data to be treated.

2. The method according to claim 1, characterized in that, The disease data processing system also includes an error attribution module, and further includes: An optimization test set is constructed based on the governance results, according to a preset optimization cycle and preset test set conditions; The optimized test set is processed based on the aforementioned index selection rules to obtain the test extraction results; If the performance value of any of the test extraction results is lower than the preset performance threshold, the index selection rule corresponding to the attribute value is optimized through the restricted action space in the error attribution module to obtain the optimized index selection rule.

3. The method according to claim 1, characterized in that, The quality control judgment rules for determining each field of the data to be addressed based on the risk level of the field include: If the risk level of the field is low, the quality control judgment rule is determined to be that the data to be treated and the verification output of at least one of the verification inference models are consistent. When the risk level of the field is medium risk, the quality control judgment rule is determined to be that the data to be treated and the verification outputs of at least two of the verification inference models are consistent; If the risk level of the field is high, the quality control judgment rule is determined to be that the data to be treated and the verification output of all the verification inference models are consistent.

4. The method according to claim 3, characterized in that, The verification reasoning model includes a disease indicator extraction model, a disease indicator verification model, and a link reasoning model. The verification output includes first verification data, second verification data, and logical review verification data. The first verification data is obtained by performing secondary indicator extraction on the multi-source case data using the disease indicator extraction model and the indicator optimization rules corresponding to the disease indicator extraction model. The second verification data is obtained by extracting indicators from the multi-source case data using the disease indicator verification model and the indicator optimization rules corresponding to the disease indicator verification model. The disease indicator extraction model is used to obtain the inference chain information of the data to be treated according to the indicator selection rules. The inference chain information is then logically reviewed by the link inference model to obtain the logical review verification data. The logical review includes at least one of logical break, skip step, and contradiction.

5. The method according to claim 1, characterized in that, The disease data processing system also includes a selective rejection router and an error attribution module; After verifying the data to be treated based on the quality control judgment rule using at least one verification inference model and the index optimization rule corresponding to the verification inference model, the method further includes: If the data to be governed fails verification, the selective rejection router will route the data to be governed to the verification unit. Based on the verification unit, the verification result corresponding to the data to be treated is obtained; The verification result is fed back to the error attribution module as an error signal, so that the error attribution module can optimize the index selection rule corresponding to the attribute value based on the error signal and through the limited action space in the error attribution module, so as to obtain the optimized index selection rule.

6. The method according to claim 5, characterized in that, The verification result is fed back to the error attribution module as an error signal, so that the error attribution module can optimize the index selection rule corresponding to the attribute value based on the error signal and through the limited action space in the error attribution module, to obtain the optimized index selection rule, including: The error attribution module encapsulates the review results into an error signal record, wherein the error signal record includes multi-source case data, disease indicators, data to be treated, treatment results, and failed verification methods corresponding to the review results; The error signal records corresponding to each of the disease indicators are accumulated respectively, and the error signal records of each of the disease indicators that failed the verification method are clustered and statistically analyzed according to a preset clustering period to obtain clustering statistical values. When the clustering statistics of the same disease indicator reach a preset iterative optimization threshold, the indicator selection rule is iteratively optimized within a limited action space to obtain the optimized indicator selection rule.

7. The method according to claim 1, characterized in that, The process of determining the index optimization rules includes: The indicator test set is structured and extracted using extraction rules to obtain test extraction results. The extraction rules include operational decision rules, negative constraints, output format, and test samples. By comparing the test extraction results with the annotation results in the indicator test set, the error type of each indicator and the error signature corresponding to the error type are determined. The error type includes at least one of the following: unit confusion, missed negative expression, value illusion, boundary condition misjudgment, and field crosstalk. Based on the error signature, the extraction rule is repaired within a restricted action space, wherein the restricted action space includes at least one of adding counterexamples, tightening constraints, splitting fields, and adjusting the judgment threshold; Based on the repaired extraction rule, the indicator test set is re-extracted in a structured manner to update the test extraction results until the test error between the test extraction results and the annotation results on the indicator test set reaches a preset termination condition, and the extraction rule is determined as the indicator optimization rule.

8. The method according to claim 1, characterized in that, After determining the governance result based on the data to be governed, the method further includes: Determine the source identifier corresponding to each of the data to be treated, wherein the source identifier includes the source institution information and the location information of the multi-source case data; Determine the governance path information for the multi-source case data, wherein the governance path information includes determining the indicator optimization rules for the data to be governed and the model version number of the disease indicator extraction model; The quality control status information of the treatment results is determined, including the verification results and the review status, and the review status includes the direct output status and the reviewed status. The source identifier, the governance path information, the quality control status information, and the governance results are stored in the corresponding disease database.

9. A disease data processing device, characterized in that, The device is used in a disease data processing system, and the device includes: The acquisition module is used to acquire multi-source case data through the acquisition channel scheduling engine of the disease data processing system; The extraction module is used to extract indicators from the multi-source case data through the disease indicator extraction model and the indicator optimization rules corresponding to the disease indicator extraction model, so as to obtain data to be treated containing at least one disease indicator. A structured module is used to determine the attribute value of at least one structured disease attribute of the data element pattern based on the data to be treated. The quality control module is used to determine the field risk level of the data to be governed through the data element pattern, and to determine the quality control judgment rule corresponding to each field of the data to be governed based on the field risk level. The verification module is used to verify the data to be treated based on the quality control judgment rules, through at least one verification inference model and the index optimization rules corresponding to the verification inference model; The output module is used to determine the governance result based on the data to be governed if the data to be governed passes verification.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, implement the disease data processing method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Special disease database data element extraction method and system based on domain description language reasoning

    CN117912623A

  • Data management method, system and equipment for assisting brain tumor special disease and medium

    CN120257991A