Method and system for intelligently extracting tumor registration data based on artificial intelligence

By using an AI-based multi-agent collaboration framework and an international rule base, tumor registry data is automatically extracted from unstructured text, solving the problems of manual dependence and data consistency in traditional tumor registry systems. This achieves efficient and accurate data extraction and encoding, making it suitable for global cancer registry systems.

CN121306384APending Publication Date: 2026-01-09CANCER INST & HOSPITAL CHINESE ACADEMY OF MEDICAL SCI
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202511499242.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Traditional tumor registry systems rely on manual extraction of data from unstructured pathology/medical records, which is prone to omissions and inconsistencies, has low quality control efficiency, and makes it difficult to coordinate multi-source heterogeneous data, resulting in low efficiency and poor data quality.

Method used

An AI-based multi-agent collaboration framework is adopted, which integrates data processing agent preprocessing, quality control agent error correction and completion, inference agent parsing, and quality control audit agent verification. Combined with the CanReg5 tool and tumor knowledge base, it realizes automated data extraction, encoding and quality control, forming a closed-loop optimization process.

Benefits of technology

It significantly improves the intelligence and automation of tumor registry data processing, reduces manual intervention, improves data accuracy and timeliness, and ensures that the results meet international standards and are portable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121306384A_ABST
    Figure CN121306384A_ABST
Patent Text Reader

Abstract

The invention provides a tumor registration data intelligent extraction method and system based on artificial intelligence, and relates to the technical field of medical data processing. The method comprises the following steps: triggering a corresponding data processing Agent to preprocess different modal data in EMR data; triggering a quality control Agent to carry out data error correction and completion on the first set, and calling a calculation tool, a CanReg5 tool and a tumor knowledge base as required to generate a second set; a quality control auditing Agent is triggered to verify the current candidate vector set, and if verification is passed, the current candidate vector set is stored in a tumor registration database; otherwise, generating an error list and triggering the state coordination Agent to correct the second set so as to enter a new round of reasoning to update the candidate vector set, and further minimizing loop correction and efficient closed-loop control. According to the method, autonomous decomposition and cross-step scheduling of tasks are realized, and links such as extraction, coding and quality control are automatically completed in a multi-source and multi-modal data processing process, so that the intelligence and automation level of tumor registration data processing is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical data processing technology, specifically to an intelligent extraction method and system for tumor registry data based on artificial intelligence. Background Technology

[0002] Traditional registration relies primarily on manual processes, extracting structured data from unstructured sources within hospitals (such as pathology reports, medical records, and imaging results). Registration staff manually review patient records, collecting key information including demographic data (name, age at diagnosis, address, etc.) and tumor characteristics (anatomical location, histological type, diagnostic stage, etc.). Historically, this data was recorded on paper forms or report cards, compiled internally within the hospital, and periodically submitted to the regional cancer registry, where it was then manually entered into a digital database by registration staff.

[0003] This workflow has several limitations: it heavily relies on skilled data entry personnel, increasing the risk of transcription errors and inconsistencies; manual verification is labor-intensive and time-consuming, leading to inefficiency and potential delays; it lacks real-time processing capabilities, hindering the monitoring and timely analysis of emerging trends; and resource constraints limit the scope of collected variables, often causing valuable details such as treatment outcomes and genetic information to be overlooked. These challenges limit the efficiency, accuracy, and data breadth of traditional tumor registry systems.

[0004] With advancements in digitalization and information technology, registration processes in many countries are gradually shifting towards semi-automation. Various regions are beginning to adopt registration software that interfaces with Electronic Medical Records (EMR), extracting elements from structured and unstructured fields in a rule- or template-driven manner, and then standardizing and coding them (e.g., ICD-10). Registration personnel then manually review the data to create comprehensive cancer records. This semi-automated model is superior to purely manual processes in terms of sensitivity and effectiveness, but it still requires significant human intervention, and the continuous monitoring and consistency maintenance of data quality remain challenging. Summary of the Invention

[0005] (a) Technical problems to be solved To address the problems of existing tumor registries relying on manual extraction and encoding from unstructured pathology / medical records, which are prone to omissions and inconsistencies, have low quality control efficiency, and are difficult to coordinate with multi-source heterogeneous data, this invention proposes an AI-based intelligent extraction method and system. This system automatically extracts and standardizes core variables from unstructured text, and incorporates built-in quality control and closed-loop verification to reduce manual intervention and improve accuracy and timeliness.

[0006] (II) Technical Solution To achieve the above objectives, the present invention provides the following technical solution: An AI-based intelligent extraction method for tumor registry data includes: Acquire EMR data, trigger the corresponding data processing agent to preprocess different modal data in the EMR data, and generate the first set; The quality control agent is triggered to perform data correction and completion on the first set, generating the second set. The inference agent is triggered to parse the second set and obtain the current candidate vector set, including: Read the patient's registration number, gender, date of birth, ICD-10 code, date of admission, and duration of illness from the second set; If the duration of illness is available, the calculation tool is invoked to calculate the date of the first visit as the onset date based on the admission date and duration of illness; otherwise, the admission date is used as the onset date. The CanReg5 tool was invoked and combined with a tumor knowledge base to convert the ICD-10 code into ICD-O-3 code; wherein the ICD-O-3 code includes anatomical, morphological, behavioral, and grading codes. The quality control audit agent is triggered to verify the current candidate vector set. If the verification passes, the current candidate vector set is saved to the tumor registry database; otherwise, an error list is generated and the status coordination agent is triggered to correct the second set in order to enter a new round of inference to update the candidate vector set.

[0007] Preferably, the data processing agent is one or any combination of natural language agent, table processing agent, image processing agent and handwriting recognition agent.

[0008] Preferably, the triggering quality control agent performs data correction and completion on the first set to generate a second set, including: Based on a predefined family of deterministic verification functions, determine whether the data in the first set satisfies the corresponding format constraints, logical constraints, or code set constraints; For data that fails the validation or is missing, based on the data and the context of the modal data in which it is located, and in combination with the first set, corresponding candidate values ​​are generated to correct and complete the data. After iterating through all the data in the first set, obtain the second set.

[0009] Preferably, the step of calling the CanReg5 tool and combining it with a tumor knowledge base to convert the ICD-10 encoding to ICD-O-3 encoding includes: The CanReg5 tool is invoked to perform a standard mapping from the ICD-10 encoding to the ICD-O-3 encoding based on built-in rules. If the mapping result is ambiguous or lacks confidence, a query request is constructed based on the mapping result and the EMR data, and the RAG tool is invoked to perform dynamic retrieval in the tumor knowledge base through the query request to obtain domain knowledge for resolving ambiguity or improving confidence. The mapping result is compared with the domain knowledge to determine the ICD-O-3 encoding.

[0010] Preferably, the step of triggering the quality control audit agent to verify the current candidate vector set includes: Based on a predefined family of deterministic verification functions, determine whether the data in the current candidate vector set satisfies the corresponding format constraints, logical constraints, or code set constraints; The CanReg5 tool is invoked to verify the consistency between different candidate vectors. If all candidate vectors in the current candidate vector set pass the verification, then the verification is successful; otherwise, the verification fails.

[0011] Preferably, the consistency verification between different candidate vectors is one or any combination of the following: (1) Verify the logical consistency between demographic characteristics and anatomical codes, including: Verify the physiological compatibility between sex and primary site, and verify the consistency between age of onset and clinical pathogenesis patterns of specific sites and morphology; wherein the anatomical coding is divided into primary part and specific part; (2) Verifying the clinicopathological consistency between anatomical and morphological codes refers to: Based on the preset site-morphology association rules, determine whether the morphological code is a common, rare, or impossible pathological type of the primary site reported by the site. (3) Verifying the logical consistency between morphological encoding and behavioral encoding means: Determine whether the behavioral codes characterizing the benign or malignant nature of a tumor match the classification of its morphological codes; (4) Verifying the logical consistency between morphological encoding and hierarchical encoding refers to: The grading criteria for determining whether the grading code characterizing the degree of tumor differentiation conforms to the classification standard of its specific morphological code.

[0012] Preferably, the triggering state coordination agent corrects the second set, including: The error list is then categorized. If the problem is determined to be a data issue, the corresponding data processing agent or quality control agent is triggered to regenerate a new second set. If the problem is determined to be a reasoning problem, the error list is used as a directional constraint and query vector to supplement the original second set.

[0013] An AI-based intelligent extraction system for tumor registry data includes: The acquisition and preprocessing module is used to acquire EMR data, trigger the corresponding data processing agent to preprocess different modal data in the EMR data, and generate a first set; The error correction and completion module is used to trigger the quality control agent to perform data error correction and completion on the first set and generate the second set. The inference module is used to trigger the inference agent to parse the second set and obtain the current candidate vector set; it includes: The reading unit is used to read the patient's registration number, gender, date of birth, ICD-10 code, admission date, and duration of illness from the second set; The calculation unit is used to, if the duration of illness is available, invoke a calculation tool to back-calculate the date of the first visit as the date of onset based on the admission date and the duration of illness; otherwise, the admission date is used as the date of onset. The conversion unit is used to call the CanReg5 tool and combine it with the tumor knowledge base to convert the ICD-10 code into ICD-O-3 code; wherein the ICD-O-3 code includes anatomical, morphological, behavioral and grading codes; The verification and correction module is used to trigger the quality control audit agent to verify the current candidate vector set. If the verification passes, the current candidate vector set is saved to the tumor registry database; otherwise, an error list is generated and the status coordination agent is triggered to correct the second set so as to enter a new round of inference to update the candidate vector set.

[0014] A storage medium storing a computer program for intelligent extraction of tumor registry data based on artificial intelligence, wherein the computer program causes a computer to execute the intelligent extraction method of tumor registry data as described above.

[0015] An electronic device, comprising: One or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the programs including methods for performing intelligent extraction of tumor registry data as described above.

[0016] (III) Beneficial Effects This invention provides an intelligent extraction method and system for tumor registry data based on artificial intelligence. Compared with existing technologies, it has the following advantages: This invention employs a multi-agent collaborative framework combined with an inference mechanism, including: first, triggering the corresponding data processing agent to preprocess different modalities of EMR data; second, triggering the quality control agent to perform data error correction and completion on the first set, calling computing tools, CanReg5 tools, and the tumor knowledge base as needed to generate a second set; then, triggering the quality control audit agent to verify the current candidate vector set. If the verification passes, the current candidate vector set is saved to the tumor registry database; otherwise, an error list is generated and the state coordination agent is triggered to correct the second set, so as to enter a new round of inference to update the candidate vector set, thereby minimizing loop correction and achieving efficient closed-loop control. This invention achieves autonomous task decomposition and cross-step scheduling, and can automatically complete extraction, encoding, and quality control in the process of processing multi-source, multi-modal data, thereby significantly improving the intelligence and automation level of tumor registry data processing. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A block diagram illustrating an intelligent extraction method for tumor registry data based on artificial intelligence, provided in an embodiment of the present invention; Figure 2 The flowchart illustrates an intelligent extraction method for tumor registration data based on artificial intelligence, as provided in this embodiment of the invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are described clearly and completely. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] This application provides an AI-based intelligent extraction method and system for tumor registry data, which solves the technical problem of how to improve the intelligence and automation level of EMR data processing. By combining the self-built RAG knowledge base with the official CanReg5 rule base, it realizes the parsing, structuring transformation and standardized coding of unstructured data, ensuring that the results meet local needs and have global consistency and portability.

[0021] The technical solution in this application is to solve the above-mentioned technical problems, and the general idea is as follows: First, this invention aims to comprehensively introduce artificial intelligence technology into the population-based tumor registry data processing workflow. Addressing the inefficiencies, error-proneness, and poor scalability caused by traditional reliance on manual data entry and rule-based judgment, it constructs a fully automated processing system based on Retrieval Augmented Generation (RAG) and an agent workflow. During inference, the agent can dynamically decide whether to invoke specific tools (such as CanReg5) based on the context, and then backfill the results into the inference chain, forming a closed-loop optimization. The entire process achieves multi-round interactive parsing, structured mapping, and verification from EMR text to computable data, significantly reducing the proportion of human intervention and improving the accuracy, interpretability, and execution efficiency of data extraction.

[0022] Secondly, to ensure that the encoding results of structured data not only conform to international standards but also possess optimization capabilities, this embodiment of the invention integrates two types of authoritative knowledge sources in the inference process, both of which are called-upon tools and autonomously scheduled by the inference agent: First, a tool-based official rule base (centered on CanReg5 rules, which can be regarded as an "expert knowledge base"). The inference agent calls the CanReg5 tool (CanReg5 is a multi-user, multi-platform open-source tool used for inputting, storing, checking, and analyzing cancer registry data) to perform standard mapping from ICD-10 encoding to ICD-O-3 encoding based on built-in rules. Second, a self-constructed tumor knowledge base. If the mapping results are ambiguous or lack confidence, the inference agent can choose to call this tool during the inference process to retrieve domain knowledge related to the case. Based on the retrieval results, a query vector is generated, and the tumor knowledge base is dynamically retrieved using the RAG tool, thereby enhancing domain knowledge coverage and improving the ICD-O-3 encoding mapping results. This combination reduces the bias caused by model illusions, ensures that the final encoding strictly conforms to globally unified standards, and significantly reduces human intervention.

[0023] Finally, the data and coding results generated by the embodiments of this invention strictly adhere to international tumor classification and coding standards, possessing universality and portability globally, and can be seamlessly applied to population-based tumor registry systems in different countries and regions. This not only facilitates cross-border data sharing and scientific research collaboration, but also provides a reliable standardized data foundation for global cancer epidemiological research, policy making, and resource allocation.

[0024] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0025] Example 1: like Figure 1 As shown, this embodiment of the invention provides an intelligent extraction method for tumor registry data based on artificial intelligence, including: S1. Obtain EMR data and trigger the corresponding data processing agent to preprocess the different modal data in the EMR data to generate a first set; S2. Trigger the quality control agent to perform data correction and completion on the first set, and generate the second set; S3. Trigger the inference agent to parse the second set and obtain the current candidate vector set; including: S31. Read the patient's registration number, gender, date of birth, ICD-10 code, date of admission, and duration of illness from the second set; S32. If the duration of illness is available, then call the calculation tool to calculate the date of the first visit as the onset date based on the admission date and duration of illness; otherwise, use the admission date as the onset date. S33. Using the CanReg5 tool and in conjunction with the tumor knowledge base, the ICD-10 code is converted into the ICD-O-3 code; wherein the ICD-O-3 code includes anatomical, morphological, behavioral, and grading codes; S4. Trigger the quality control audit agent to verify the current candidate vector set. If the verification passes, save the current candidate vector set to the tumor registry database; otherwise, generate an error list and trigger the status coordination agent to correct the second set in order to enter a new round of inference to update the candidate vector set.

[0026] The embodiments of the present invention realize autonomous task decomposition and cross-step scheduling, which can automatically complete extraction, encoding, quality control and other links in the process of processing multi-source and multi-modal data, thereby significantly improving the intelligence and automation level of tumor registry data processing.

[0027] like Figure 2 As shown, Figure 2 A flowchart of an intelligent extraction method for tumor registry data based on artificial intelligence is disclosed.

[0028] Next, we will combine Figure 2 Detailed explanation of each of the above steps: First, it should be clarified that this embodiment of the invention constructs a multi-agent integrated system for multimodal data. This system adopts a multi-agent architecture, combining a manually constructed tumor knowledge base, computational tools, RAG tools, and CanReg5 tools to achieve unified processing and deep reasoning of free text, tables, images, and handwritten content from different data centers. Because clinical data is diverse in form and complex in path, a single model is insufficient. Therefore, this embodiment of the invention follows the concept of Agent AI, dividing the system into several dedicated agents, each collaborating to complete tasks according to a predetermined role. This multi-module architecture allows for flexible task allocation, improving interpretability and efficiency. Agents coordinate and schedule each other to complete processes such as data reception, preprocessing, knowledge reasoning, and result output. In particular, this embodiment of the invention can deploy a large language model within the hospital and localize all external tools such as the tumor registry knowledge base, RAG index, and CanReg5, meeting the hospital's data security and compliance requirements.

[0029] Regarding the computing framework, this embodiment of the invention deploys a large language model (e.g., DeepSeek-R1) locally on an internal server using Ollam (an open-source framework for running large local language models), with all inference and tool calls completed within the intranet. End-to-end multi-agent orchestration, knowledge base construction, tool writing, and invocation are implemented based on LangChain (a framework for developing applications driven by large language models). Model weights, vector indices, and intermediate results are all stored in locally controlled storage, with outbound dependencies disabled by default.

[0030] For data input, EMR data exported from the Hospital Information System (HIS) was used, including the patient record summary, admission record, initial progress note, discharge record, and various case reports, imaging reports, and biochemical data, specifically structured, unstructured, and image data. For data output, several basic coding information from existing tumor registries were included: registry number, gender, date of birth, ICD-10 code, date of onset, and ICD-O-3 codes (anatomical Topo, morphological Morp, behavioral Behavioral ...

[0031] Reference Figure 2 The method provided in this embodiment of the invention specifically includes: In step S1, EMR data is acquired, and the corresponding data processing agent is triggered to preprocess the different modal data in the EMR data to generate the first set.

[0032] In this step, the acquired EMR data is defined as the source set S = {txt, tbl, img, hw}, and the s-th data is denoted as... The set of candidate records corresponding to the S sources is K. s Then the total size of the source set is: Specifically, the above data is processed and stored entirely within a controlled environment within the hospital, and outbound access from the external network is disabled.

[0033] Next, this step invokes the corresponding data processing agent based on the data modality to perform preprocessing operations: The Natural Language Agent is responsible for cleaning and extracting text from medical records.

[0034] The table processing agent reads and parses structured tables.

[0035] Image processing agents use medical image models to process images from sources such as computer-to-myograph (CT) and magnetic resonance imaging (MRI).

[0036] The handwriting recognition agent uses an OCR tool to recognize handwritten medical records. The OCR process can be described as follows: I represents the handwritten input image, and T represents the output text.

[0037] Finally, the original preprocessed results are generated as the first set. Each element This indicates a standardized data item from one of the above sources.

[0038] In step S2, the quality control agent is triggered to perform data correction and completion on the first set, generating the second set.

[0039] After preprocessing generates standardized input, this step triggers the quality control agent to perform data correction and completion, including: S21. Based on a predefined family of deterministic verification functions, determine whether the data in the first set satisfies the corresponding format constraints, logical constraints, or code set constraints.

[0040] S22. For data that fails the verification or is missing, based on the context of the data and the modal data in which it is located, and in combination with the first set, generate corresponding candidate values ​​to correct and complete the data.

[0041] S23. After traversing all the data in the first set, obtain the second set.

[0042] Specifically: The system first defines a family of deterministic verification functions. Each of them Apply to standardized data items This is used to determine whether a data item meets predefined format constraints, logical constraints, or code set constraints. For example, format constraints include: resident ID numbers must be 18 digits and pass the check digit rules; date fields must use YYYY-MM-DD and be parsable as valid calendar dates; and ICD-10 encoding must begin with one uppercase letter followed by two digits and may include 1–4 decimal places. If these constraints exist... If the data item violates the rules, it means that the data item needs to be corrected or removed.

[0043] For fields that fail validation or are missing, the system, based on the data and the context of its modality (related text, tables, image summaries, etc. for the same patient), along with the current first set D, submits it to the quality control agent as input. The agent returns a candidate value. If the value passes format, logic, and code set validation, it is written back as an error correction or completion result, updating the corresponding data item in D. If the conditions are not met, the original state is maintained and marked as awaiting manual review.

[0044] Ultimately, it can be represented as a mapping. ,in This second set, after standardization, correction, and completion, ensures the consistency and integrity of the data input into the subsequent entity extraction and registration coding stages.

[0045] In step S3, the inference agent is triggered to parse the second set and obtain the current candidate vector set.

[0046] It should be noted that this step triggers the inference agent to directly read simple fields using a "start with the easy and then move to the difficult" strategy. For complex fields, it calls RAG, CanReg5, and the onset date calculation tool as needed, gradually completing all registered fields and forming verifiable candidate results in a loop. The specific steps are as follows: S31. Read the patient's registration number, gender, date of birth, ICD-10 code, date of admission, and duration of illness from the second set; it is understood that the "date of onset" here is defined as the date on which the patient first consults a doctor or is admitted to the hospital due to suspected cancer.

[0047] S32. If the duration of illness is available, the calculation tool is invoked to calculate the date of the first visit as the onset date based on the admission date and duration of illness; otherwise, the admission date is used as the onset date.

[0048] To improve the accuracy and consistency of date calculation, this invention provides a tool for calculating the date of onset. This tool calculates the date based on the admission date and the duration of illness recorded in the medical record. Specifically, if the duration of illness is recorded, the date of the first visit is considered the date of onset, and the calculation is performed as follows: "Date of onset = Date of admission". The duration of illness is used for calculation; if the duration of illness is not found, the onset date is taken as the admission date (onset date = admission date). Let the patient's admission date be... The duration of illness is The date of onset is then calculated as follows: in, The time unit can be a range of days, months, or years. The system will automatically convert the time unit and format the date to avoid date deviations in large models during free text inference and ensure the consistency of the calculated date of onset in logic and format.

[0049] S33. Using the CanReg5 tool and in conjunction with the tumor knowledge base, the ICD-10 code is converted into the ICD-O-3 code; wherein the ICD-O-3 code includes anatomical, morphological, behavioral, and grading codes.

[0050] To improve the accuracy of reasoning, this embodiment of the invention constructs a local tumor knowledge base, whose sources include cancer research papers downloaded from PubMed, clinical expert experience, and ICDs. O 3. Coding rules and teaching materials, etc. The construction process is as follows: First, professional knowledge in the field of tumor registry is manually integrated to form a knowledge base including PubMed literature, clinical expert experience, teaching materials, ICD-O-3 standard books, etc. Each document in the knowledge base is represented using vectorization: ,in, It is the first in the knowledge base l One document, This is a text embedding function, implemented using gte-Qwen2-7B-instruct (a state-of-the-art text embedding model based on the Qwen2-7B Large Language Model (LLM)) and the trained model. The generated tumor knowledge base is denoted as set. , where m is the number of documents.

[0051] For example, the step of calling the CanReg5 tool and combining it with a tumor knowledge base to convert the ICD-10 encoding to ICD-O-3 encoding includes: S331. Call the CanReg5 tool to perform standard mapping from ICD-10 encoding to ICD-O-3 encoding based on built-in rules.

[0052] To improve the accuracy and consistency of encoding conversion, the present invention provides the CanReg5 tool. The CanReg5 tool can not only be used to automatically convert topological and morphological codes between different versions of the International Classification of Diseases (ICD) and the International Classification of Cancer (ICD-O), but also check the legality of variable values, verify the consistency between age, gender, and disease location in subsequent verification stages, and identify possible duplicate tumor records.

[0053] Here, when invoking this tool, it will be from the second set. Using the extracted registry number, gender, and ICD-10 coding variables as input, this tool can automatically generate ICD-O-3 anatomical topo, morphological morp, behavioral beha, and grading grade based on built-in rules. The calculation process can be represented as follows: ,in These are mapping functions for the CanReg5 toolkit. For the extracted registration number, gender, and ICD-10 code, This enables the conversion of ICD-O-3 codes and related grading results, thereby achieving a standardized and efficient tumor registry data processing workflow.

[0054] S332. If the mapping result is ambiguous or lacks confidence, a query request is constructed based on the mapping result and the EMR data (i.e., pathological diagnosis reports, imaging examination reports, etc.), and the RAG tool is invoked to dynamically search the tumor knowledge base through the query request to obtain domain knowledge for resolving ambiguity or improving confidence. This invention designs and integrates the RAG tool to address the problems of large models lacking domain knowledge and being prone to generating illusions. By dynamically searching the local tumor registry knowledge base through RAG, external knowledge is introduced during the generation process, enabling the model to utilize updated domain knowledge to generate more accurate and reliable answers without retraining.

[0055] Here, during the current inference step, the inference agent automatically detects ambiguities or insufficient confidence in the mapping results; once additional support is needed, it invokes the RAG tool and constructs a query vector based on the mapping result c. Then, semantic similarity retrieval was performed: Where argmax represents the index of the maximum value found in a given function or array. Represents the cosine similarity function. The most relevant knowledge documents retrieved.

[0056] S333. Perform a consistency determination between the mapping result and the domain knowledge to determine the ICD-O-3 encoding.

[0057] Here, after returning the retrieved domain knowledge in the form of source annotations and evidence chains, it is concatenated with the mapping results to perform consistency adjudication to determine the variable values, ensuring that the output results are consistent with authoritative tumor registry standards, thereby ensuring that the system has interpretability and traceability.

[0058] In step S4, the quality control audit agent is triggered to verify the current candidate vector set. If the verification passes, the current candidate vector set is saved to the tumor registry database; otherwise, an error list is generated and the status coordination agent is triggered to correct the second set in order to enter a new round of inference to update the candidate vector set.

[0059] After obtaining the current candidate vector set, it is not directly stored in the database, but instead enters the source tracing and quality control stage. This step triggers the quality control audit agent to verify the current candidate vector set based on the expert rule set and the official rule base. The relevant steps are as follows: S10. Based on a predefined family of deterministic verification functions, determine whether the data in the current candidate vector set satisfies the corresponding format constraints, logical constraints, or code set constraints.

[0060] S20. Invoke the CanReg5 tool to verify the consistency between different candidate vectors. The verification of consistency between different candidate vectors includes: (1) Verify the logical consistency between demographic characteristics and anatomical codes, including: Verify the physiological compatibility between sex and primary site, and verify the consistency between age of onset and clinical pathogenesis patterns of specific sites and morphology; wherein the anatomical coding is divided into primary part and specific part; (2) Verifying the clinicopathological consistency between anatomical and morphological codes refers to: Based on the preset site-morphology association rules, determine whether the morphological code is a common, rare, or impossible pathological type of the primary site reported by the site. (3) Verifying the logical consistency between morphological encoding and behavioral encoding means: Determine whether the behavioral codes characterizing the benign or malignant nature of a tumor match the classification of its morphological codes; (4) Verifying the logical consistency between morphological encoding and hierarchical encoding refers to: The grading criteria for determining whether the grading code characterizing the degree of tumor differentiation conforms to the classification standard of its specific morphological code.

[0061] S30. If all candidate vectors in the current candidate vector set pass the verification, then the verification is successful; otherwise, the verification fails.

[0062] Specifically, the verification process can be represented as follows: in, Covers field integrity, formatting standards, date logic, and cross-field consistency; Let t be the current set of candidate vectors, and let t be the time index; Check is the test function.

[0063] If Check=1, output the final result. And save it to the tumor registry database.

[0064] If Check=0, the quality control audit agent generates an error list and triggers the status coordination agent to correct the second set, including: S100. Classify the error list.

[0065] S200. If the problem is determined to be a data issue, the corresponding data processing agent or quality control agent is triggered to regenerate a new second set.

[0066] Specifically, if the data issue is determined to be such as OCR misreading, table parsing error, non-compliant format, or cross-document conflict, the quality control agent or related preprocessing agent will be triggered to regenerate a new second set. and only for The variables involved are re-reasoned in the next round in a predetermined order.

[0067] S300. If the problem is determined to be a reasoning problem, the error list is used as a directional constraint and query vector to supplement the original second set.

[0068] If the problem is determined to be a reasoning issue such as a conflict in rule application, improper knowledge mapping, or incorrect date calculation, then the original second set should be retained. Unchanged, will As a directional constraint and query interaction inference agent, it only re-infers the relevant variables.

[0069] Finally, based on the modified second set, a new round of inference is initiated to update the candidate vector set until the corresponding candidate vector set passes the verification and is saved to the tumor registry database.

[0070] Thus, this embodiment of the invention completes the entire process of the intelligent extraction method for tumor registry data based on artificial intelligence. It can be understood that, through the aforementioned knowledge-driven and multi-agent collaboration scheme, this embodiment of the invention achieves automated and locally deployable high-quality integration of multi-source heterogeneous medical data; it utilizes the RAG tool to introduce authoritative knowledge, and employs computational tools and the CanReg5 tool to ensure accurate encoding; simultaneously, it uses an end-to-end closed loop of "reasoning—quality control—feedback—correction" throughout the entire process of acquisition, parsing, encoding, and database storage; ultimately improving the efficiency and reliability of tumor registry data capture, and saving the structured and verified results to the tumor registry database, providing a reliable data foundation for subsequent analysis and applications.

[0071] To verify the reliability and feasibility of the proposed method in real-world scenarios, the following specific experiment is provided. This experiment completes the entire closed-loop process of model inference, knowledge retrieval, encoding conversion, and quality control review in a local environment. The inference engine uses the open-source Ollam to deploy a large local language model, with DeepSeek-r1 as the base model; the workflow and tool orchestration are implemented based on LangChain. The system operates with a "multi-agent and retrieval enhancement (RAG) + rule calibration" structure, completing tasks such as information extraction, cross-document linking, ICD-10→ICD-O-3 encoding, and calculation of the onset date.

[0072] The data source is 130 real electronic health records from a hospital, including text fields such as admission and discharge records, initial medical history, pathology reports, imaging conclusions, and surgical / treatment records. All samples were de-identified, retaining only the medical content and necessary timestamps required for the study. To assess objectivity and accuracy, the data was independently labeled by coders with experience in tumor registry. All 130 cases served as the test set, used solely for objective evaluation and not for any parameter tuning.

[0073] The evaluation used only accuracy as a uniform metric, and statistics were calculated for each field individually. The results are shown in Table 1: Table 1 Experimental results demonstrate that a workflow based on a large language model can achieve better ICD-O-3 encoding performance than expert registrars, while significantly reducing manual extraction workload. Wider adoption of this approach has the potential to accelerate data acquisition and improve the integrity and quality of cancer surveillance systems.

[0074] Example 2: This invention provides an intelligent extraction system for tumor registry data based on artificial intelligence, comprising: The acquisition and preprocessing module is used to acquire EMR data, trigger the corresponding data processing agent to preprocess different modal data in the EMR data, and generate a first set; The error correction and completion module is used to trigger the quality control agent to perform data error correction and completion on the first set and generate the second set. The inference module is used to trigger the inference agent to parse the second set and obtain the current candidate vector set; it includes: The reading unit is used to read the patient's registration number, gender, date of birth, ICD-10 code, admission date, and duration of illness from the second set; The calculation unit is used to, if the duration of illness is available, invoke a calculation tool to back-calculate the date of the first visit as the date of onset based on the admission date and the duration of illness; otherwise, the admission date is used as the date of onset. The conversion unit is used to call the CanReg5 tool and combine it with the tumor knowledge base to convert the ICD-10 code into ICD-O-3 code; wherein the ICD-O-3 code includes anatomical, morphological, behavioral and grading codes; The verification and correction module is used to trigger the quality control audit agent to verify the current candidate vector set. If the verification passes, the current candidate vector set is saved to the tumor registry database; otherwise, an error list is generated and the status coordination agent is triggered to correct the second set so as to enter a new round of inference to update the candidate vector set.

[0075] Example 3: This invention provides a storage medium storing a computer program for intelligent extraction of tumor registry data based on artificial intelligence, wherein the computer program causes a computer to execute the intelligent extraction method of tumor registry data as described above.

[0076] Example 4: This invention provides an electronic device, comprising: One or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the programs including methods for performing intelligent extraction of tumor registry data as described above.

[0077] It is understood that the AI-based intelligent extraction system, storage medium, and electronic device for tumor registration data provided in the embodiments of the present invention correspond to the AI-based intelligent extraction method for tumor registration data provided in the embodiments of the present invention. The explanations, examples, and beneficial effects of the relevant content can be referred to the corresponding parts of the method, and will not be repeated here.

[0078] In summary, compared with existing technologies, it has the following beneficial effects: 1. The embodiments of the present invention realize autonomous task decomposition and cross-step scheduling, which can automatically complete extraction, encoding, quality control and other links in the process of processing multi-source and multi-modal data, thereby significantly improving the intelligence and automation level of tumor registration data processing.

[0079] 2. In the structured data generation stage, this invention innovatively combines internationally recognized coding rules (CanReg5) with self-constructed localized optimization rules (tumor knowledge base) to achieve dual calibration. This method maintains global comparability while optimizing for local data characteristics, effectively reducing coding errors and information loss. 3. This invention supports fully localized deployment and can run in hospitals or cancer registry centers while ensuring data compliance. Furthermore, the rule system adopted in this invention is internationally applicable and has the potential for cross-regional and cross-national promotion, adapting to the technical requirements and regulatory environments of cancer registry systems in different countries.

[0080] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0081] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for intelligent extraction of tumor registry data based on artificial intelligence, characterized in that, include: Acquire EMR data, trigger the corresponding data processing agent to preprocess different modal data in the EMR data, and generate the first set; The quality control agent is triggered to perform data correction and completion on the first set, generating the second set. The inference agent is triggered to parse the second set and obtain the current candidate vector set, including: Read the patient's registration number, gender, date of birth, ICD-10 code, date of admission, and duration of illness from the second set; If the duration of illness is available, the calculation tool is invoked to calculate the date of the first visit as the onset date based on the admission date and duration of illness; otherwise, the admission date is used as the onset date. The CanReg5 tool was invoked and combined with a tumor knowledge base to convert the ICD-10 code into ICD-O-3 code; wherein the ICD-O-3 code includes anatomical, morphological, behavioral, and grading codes. The quality control audit agent is triggered to verify the current candidate vector set. If the verification passes, the current candidate vector set is saved to the tumor registry database; otherwise, an error list is generated and the status coordination agent is triggered to correct the second set in order to enter a new round of inference to update the candidate vector set.

2. The intelligent extraction method for tumor registration data as described in claim 1, characterized in that, The data processing agent is one or any combination of natural language agent, table processing agent, image processing agent and handwriting recognition agent.

3. The intelligent extraction method for tumor registration data as described in claim 1, characterized in that, The triggering quality control agent performs data correction and completion on the first set to generate a second set, including: Based on a predefined family of deterministic verification functions, determine whether the data in the first set satisfies the corresponding format constraints, logical constraints, or code set constraints; For data that fails the validation or is missing, based on the data and the context of the modal data in which it is located, and in combination with the first set, corresponding candidate values ​​are generated to correct and complete the data. After iterating through all the data in the first set, obtain the second set.

4. The intelligent extraction method for tumor registration data as described in claim 1, characterized in that, The process of calling the CanReg5 tool and combining it with a tumor knowledge base to convert the ICD-10 encoding to ICD-O-3 encoding includes: The CanReg5 tool is invoked to perform a standard mapping from the ICD-10 encoding to the ICD-O-3 encoding based on built-in rules. If the mapping result is ambiguous or lacks confidence, a query request is constructed based on the mapping result and the EMR data, and the RAG tool is invoked to perform dynamic retrieval in the tumor knowledge base through the query request to obtain domain knowledge for resolving ambiguity or improving confidence. The mapping result is compared with the domain knowledge to determine the ICD-O-3 encoding.

5. The intelligent extraction method for tumor registration data as described in claim 4, characterized in that, The triggering quality control audit agent verifies the current candidate vector set, including: Based on a predefined family of deterministic verification functions, determine whether the data in the current candidate vector set satisfies the corresponding format constraints, logical constraints, or code set constraints; The CanReg5 tool is invoked to verify the consistency between different candidate vectors. If all candidate vectors in the current candidate vector set pass the verification, then the verification is successful; otherwise, the verification fails.

6. The intelligent extraction method for tumor registration data as described in claim 5, characterized in that, The consistency verification between different candidate vectors is one or any combination of the following: (1) Verify the logical consistency between demographic characteristics and anatomical codes, including: Verify the physiological compatibility between sex and primary site, and verify the consistency between age of onset and clinical pathogenesis patterns of specific sites and morphology; wherein the anatomical coding is divided into primary part and specific part; (2) Verifying the clinicopathological consistency between anatomical and morphological codes refers to: Based on the preset site-morphology association rules, determine whether the morphological code is a common, rare, or impossible pathological type of the primary site reported by the site. (3) Verifying the logical consistency between morphological encoding and behavioral encoding means: Determine whether the behavioral codes characterizing the benign or malignant nature of a tumor match the classification of its morphological codes; (4) Verifying the logical consistency between morphological encoding and hierarchical encoding refers to: The grading criteria for determining whether the grading code characterizing the degree of tumor differentiation conforms to the classification standard of its specific morphological code.

7. The intelligent extraction method for tumor registration data as described in claim 1, characterized in that, The triggering state coordination agent corrects the second set, including: The error list is categorized; If the problem is determined to be a data issue, the corresponding data processing agent or quality control agent is triggered to regenerate a new second set. If the problem is determined to be a reasoning problem, the error list is used as a directional constraint and query vector to supplement the original second set.

8. An intelligent extraction system for tumor registry data based on artificial intelligence, characterized in that, include: The acquisition and preprocessing module is used to acquire EMR data, trigger the corresponding data processing agent to preprocess different modal data in the EMR data, and generate a first set; The error correction and completion module is used to trigger the quality control agent to perform data error correction and completion on the first set and generate the second set. The inference module is used to trigger the inference agent to parse the second set and obtain the current candidate vector set; it includes: The reading unit is used to read the patient's registration number, gender, date of birth, ICD-10 code, admission date, and duration of illness from the second set; The calculation unit is used to, if the duration of illness is available, invoke a calculation tool to back-calculate the date of the first visit as the date of onset based on the admission date and the duration of illness; otherwise, the admission date is used as the date of onset. The conversion unit is used to call the CanReg5 tool and combine it with the tumor knowledge base to convert the ICD-10 code into ICD-O-3 code; wherein the ICD-O-3 code includes anatomical, morphological, behavioral and grading codes; The verification and correction module is used to trigger the quality control audit agent to verify the current candidate vector set. If the verification passes, the current candidate vector set is saved to the tumor registry database; otherwise, an error list is generated and the status coordination agent is triggered to correct the second set so as to enter a new round of inference to update the candidate vector set.

9. A storage medium, characterized in that, It stores a computer program for intelligent extraction of tumor registry data based on artificial intelligence, wherein the computer program causes a computer to execute the intelligent extraction method of tumor registry data as described in any one of claims 1 to 7.

10. An electronic device, characterized in that, include: One or more processors; Memory; And one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the programs including methods for performing the intelligent extraction method of tumor registry data as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Knowledge-based electronic medical record quality control method

    CN106682397A

  • Mapping method of diseases and human body parts

    CN107545143A

  • Disease code automatic conversion method and device

    CN111309703A

  • Artificial intelligence auditing quality control mode and system based on medical insurance disease category payment system ICD coding

    CN112992366A

  • Data extraction method and device, computer and readable storage medium

    CN114550190A