A medical identification and analysis method, equipment and medium
By vectorizing and analyzing medical identification information, structured data is generated and its credibility is verified, solving the problems of low efficiency and poor authority in traditional medical identification and achieving efficient and authoritative identification results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INSPUR ZHUOSHU BIG DATA IND DEV CO LTD
- Filing Date
- 2026-01-05
- Publication Date
- 2026-06-02
AI Technical Summary
Traditional medical assessments rely on manual processing, which is inefficient, time-consuming, and prone to errors. This makes it difficult to meet the dual requirements of efficiency and authority in medical dispute resolution. Furthermore, existing systems lack the ability to access legal knowledge in real time and perform structured verification, resulting in inconsistent analysis quality.
By vectorizing the medical assessment information uploaded by users, a reading extraction agent is used to generate structured data, a diagnostic reasoning agent is used to generate preliminary analysis opinions, and a review agent is used to review the credibility of the opinions. When the credibility of the preliminary analysis opinions is inconsistent, expert opinions are obtained for adjustment, and finally, the assessment results are generated.
It improves the efficiency of medical identification, reduces human intervention, and ensures the authority and compliance of identification results, meeting medical standards and legal requirements.
Smart Images

Figure CN122135872A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and medical information processing technology, and in particular to a medical identification and analysis method, device and medium. Background Technology
[0002] Traditional medical assessments rely heavily on manual processing, which suffers from drawbacks such as complex medical record formats, difficulties in information extraction, time-consuming legal reviews, and highly subjective analytical conclusions. These shortcomings make it difficult to meet the current dual demands for efficiency and authority in handling medical disputes.
[0003] With the rapid evolution of artificial intelligence and medical information technology, medical assessment, as a crucial link connecting medical behavior, legal responsibility, and patient rights, is facing unprecedented challenges and demands for transformation. On the one hand, the number of medical disputes is increasing year by year, and medical records are complex and diverse in form, encompassing multimodal data such as PDF scans, imaging films, handwritten records, and electronic medical records. Traditional manual processing methods are inefficient, time-consuming, and prone to errors. On the other hand, assessment conclusions must strictly adhere to regulations such as the "Regulations on the Handling of Medical Accidents" and the "Civil Code," as well as eighteen core medical systems, requiring extremely high levels of professionalism, compliance, and traceability. However, existing systems generally lack the ability to access and verify legal knowledge in real time, resulting in inconsistent analysis quality and failing to meet the high standards demanded by the judiciary, administration, and patients. Summary of the Invention
[0004] This application provides a medical identification analysis method, equipment, and medium to solve the following technical problem: how to solve the problems of low efficiency and poor authority in medical identification.
[0005] In a first aspect, embodiments of this application provide a medical identification analysis method, the method comprising: vectorizing raw data in medical identification information uploaded by a user to obtain target data, wherein the medical identification information further includes a medical identification task corresponding to the raw data; inputting the target data and the medical identification task into a reading extraction agent to obtain structured data output by the reading extraction agent, wherein the reading extraction agent is used to extract and standardize information from the target data based on the medical identification task to generate structured data; inputting the structured data into a diagnostic reasoning agent to obtain preliminary analysis opinions output by the diagnostic reasoning agent, wherein the diagnostic reasoning agent is used to analyze the structured data to generate preliminary analysis opinions corresponding to the medical identification task; inputting the preliminary analysis opinions into a review agent to obtain credibility tags output by the review agent, wherein the review agent is used to review the preliminary analysis opinions to generate credibility tags corresponding to the preliminary analysis opinions; if the credibility tags of the preliminary analysis opinions are inconsistent with preset tags, obtaining expert opinions corresponding to the preliminary analysis opinions; adjusting the preliminary analysis opinions based on the expert opinions to obtain the identification result corresponding to the medical identification task.
[0006] Secondly, embodiments of this application also provide an apparatus, the apparatus comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a medical identification analysis method as described in the first aspect above.
[0007] Thirdly, embodiments of this application also provide a computer storage medium storing computer-executable instructions, which, when executed, implement a medical identification and analysis method as described in the first aspect above.
[0008] The medical identification and analysis method, equipment, and medium provided in this application have the following beneficial effects: In this embodiment, the original data in the medical assessment information uploaded by the user can be vectorized to obtain target data. Then, the medical assessment task and the target data are input into a reading and extraction agent to obtain structured data output by the agent. This structured data is then input into a diagnostic reasoning agent to obtain preliminary analysis opinions. These preliminary analysis opinions can then be input into a review agent to obtain a credibility label corresponding to them. If the credibility label of the preliminary analysis opinion is inconsistent with a preset label, expert opinions are obtained. Finally, the preliminary analysis opinion is adjusted based on the expert opinions to obtain the assessment result. This method, by generating preliminary analysis opinions through reading and extraction agents and diagnostic reasoning agents, reduces manual intervention. Furthermore, low-credibility cases trigger expert review, preventing the waste of expert resources on simple cases and improving the efficiency of medical assessments. Simultaneously, after review by the review agent, experts can also manually verify low-credibility preliminary analysis opinions, ensuring the final assessment result is authoritative and complies with medical standards and legal requirements. Attached Figure Description
[0009] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart of a medical identification and analysis method provided in this application embodiment; Figure 2 A structural diagram of a medical identification and analysis system provided in this application embodiment; Figure 3 This is a schematic diagram of the internal structure of a medical identification and analysis device provided in an embodiment of this application. Detailed Implementation
[0010] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0011] In traditional medical forensic practice, experts typically spend a significant amount of time meticulously reviewing medical records, manually extracting key information, and then writing a summary of diagnosis and treatment and an analysis of responsibility based on their personal experience. This process is highly dependent on the individual expert's skill level, and suffers from drawbacks such as strong subjectivity, inconsistent standards, and outdated knowledge. Furthermore, due to inconsistent terminology and varying writing styles in medical records, manual compilation often leads to omissions or misjudgments, further affecting the accuracy and credibility of the forensic conclusions. In addition, legal provisions, treatment guidelines, and expert consensus are scattered across a vast amount of documents, lacking systematic integration and intelligent retrieval methods. Experts must manually consult paper or electronic materials, resulting in low efficiency and a high risk of overlooking crucial evidence.
[0012] This application provides a medical identification and analysis scheme. The technical solution proposed in this application will be described in detail below with reference to the accompanying drawings.
[0013] Figure 1 A flowchart illustrating a medical identification and analysis method provided in this application embodiment. Figure 1 As shown in the figure, a method provided in this application embodiment specifically includes the following steps: Step 101: Vectorize the original data in the medical appraisal information uploaded by the user to obtain the target data.
[0014] The medical assessment information also includes the medical assessment task corresponding to the original data.
[0015] In this embodiment, user-uploaded medical assessment information can be processed to obtain target data. For the original data corresponding to the user-uploaded medical assessment task, such as medical assessment reports, medical records, and examination forms, Optical Character Recognition (OCR) processing can be performed to convert the text in the image into editable and searchable electronic text, followed by vectorization processing. If the original data is already electronic text, it can be directly vectorized, thus reducing manual intervention and improving the efficiency of medical assessment analysis. The specific vectorization processing method is not specifically limited in this application; the main purpose is to convert the user-uploaded original data into target data that a computer can understand, thereby accelerating the assessment process and improving efficiency.
[0016] Step 102: Input the target data and the medical identification task into the reading extraction agent to obtain the structured data output by the reading extraction agent.
[0017] The reading extraction agent is used to extract and standardize information from the target data based on the medical identification task, generating structured data.
[0018] In this embodiment, a reading extraction agent can be used to extract and standardize target data based on a medical identification task, generating structured data. For example, key entities (such as patient name, age, symptoms, diagnosis, and examination indicators) can be identified and extracted from the target data, and then converted into standardized fields (such as "symptoms: fever; duration: 3 days"). This eliminates ambiguous expressions in natural language and provides accurate input for subsequent diagnostic reasoning.
[0019] Step 103: Input the structured data into the diagnostic reasoning agent to obtain the preliminary analysis opinion output by the diagnostic reasoning agent.
[0020] The diagnostic reasoning agent is used to analyze the structured data and generate preliminary analysis opinions corresponding to the medical identification task.
[0021] In this embodiment, structured data can be input into a diagnostic reasoning agent to obtain preliminary analysis opinions corresponding to medical assessment tasks. This improves efficiency by allowing the diagnostic reasoning agent to perform medical assessments that would otherwise require manual intervention.
[0022] Step 104: Input the preliminary analysis opinion into the review agent to obtain the credibility label output by the review agent.
[0023] The review agent is used to review the preliminary analysis opinions and generate a credibility label corresponding to the preliminary analysis opinions.
[0024] In this embodiment, an auditing agent can be used to audit the preliminary analysis opinions and generate a corresponding credibility score. For example, the credibility score of the preliminary analysis opinions can be calculated, and credibility labels such as "high / medium / low" can be added to them. For instance, if the calculated credibility score of the preliminary analysis opinions is 65 points, according to the rules, 60-69 points correspond to the label "medium credibility," and a medium credibility label can be added to the preliminary analysis opinions. Through the credibility label, the abstract credibility can be transformed into a specific degree of credibility or risk (low credibility means high risk). The credibility label simplifies the complex evaluation results into operable symbols. In this way, by using simple labels to assist decision-making, key information can be quickly filtered, the risk of misjudgment can be reduced, numerical misleading can be avoided, and the interpretability of the results can be strengthened. The "high / medium / low" label implies "the result is uncertain," which can encourage subsequent comprehensive judgment based on expert opinions.
[0025] It should be noted that in practical applications, when the credibility labels of all parts in the preliminary analysis are medium or low, an overall credibility label (medium or low) can be generated for the preliminary analysis to indicate that the credibility of the preliminary analysis is not high and the quality is poor. When the majority of the credibility labels in the preliminary analysis are high, they can be marked separately, focusing on indicating the content with medium / low credibility.
[0026] Step 105: If the credibility label of the preliminary analysis opinion is inconsistent with the preset label, obtain the expert opinion corresponding to the preliminary analysis opinion.
[0027] In this embodiment, the credibility label can be compared with a preset label. If the credibility label is inconsistent with the preset label, for example, if the credibility label of the preliminary analysis opinion is medium credibility while the preset label is high credibility, it can be determined that the credibility of the preliminary analysis opinion is not high and the risk is high. In this case, it can be sent to experts to obtain expert opinions corresponding to the preliminary analysis opinion. This ensures the accuracy of medical assessment.
[0028] In practical applications, preliminary analysis opinions with an overall credibility rating of low or medium can be directly submitted to experts for collaborative processing. This is simple and convenient. If most of the content in the preliminary analysis opinion is highly credible, the credibility rating of the preliminary analysis opinion can be compared with preset ratings. Preliminary analysis opinions that do not meet the preset ratings can be submitted to experts for collaborative analysis, thus ensuring the accuracy of the opinions. In practical applications, experts can leverage their domain experience and critical thinking to quickly locate logical flaws missed in the preliminary review opinions and supplement medical content not covered by the aforementioned diagnostic reasoning agent.
[0029] Step 106: Adjust the preliminary analysis based on the expert opinions to obtain the identification results corresponding to the medical identification task.
[0030] In this embodiment, expert opinions can be used to adjust the preliminary analysis to obtain the identification results corresponding to the medical identification task. This allows for a comparison between manual review and AI results, ensuring that the identification results are authoritative, credible, and compliant with medical standards and legal requirements.
[0031] It should be noted that when the credibility label indicates that the credibility of the preliminary analysis opinion is higher than the preset threshold, that is, when the credibility of each part of the preliminary analysis opinion is higher than the preset threshold, the preliminary analysis opinion can be directly determined as the final identification result. In this way, time can be saved while ensuring the quality of the identification result. In this embodiment, the original data in the medical assessment information uploaded by the user can be vectorized to obtain target data. Then, the medical assessment task and the target data are input into a reading and extraction agent to obtain structured data output by the agent. This structured data is then input into a diagnostic reasoning agent to obtain preliminary analysis opinions. These preliminary analysis opinions can then be input into a review agent to obtain a credibility label corresponding to them. If the credibility label of the preliminary analysis opinion is inconsistent with a preset label, expert opinions are obtained. Finally, the preliminary analysis opinion is adjusted based on the expert opinions to obtain the assessment result. This method, by generating preliminary analysis opinions through reading and extraction agents and diagnostic reasoning agents, reduces manual intervention. Furthermore, low-credibility cases trigger expert review, preventing the waste of expert resources on simple cases and improving the efficiency of medical assessments. Simultaneously, after review by the review agent, experts can also manually verify low-credibility preliminary analysis opinions, ensuring the final assessment result is authoritative and complies with medical standards and legal requirements.
[0032] The medical identification analysis method proposed in this application is applicable to scenarios such as medical malpractice technical identification, medical dispute mediation, and judicial identification. It provides efficient, accurate, and compliant intelligent assistance to identification institutions, legal departments, and medical institutions, and promotes the transformation of the medical identification industry from "human experience-driven" to "data intelligence-driven".
[0033] In one possible implementation, the process of vectorizing the raw data in the user-uploaded medical assessment information to obtain the target data includes: The original data is subjected to OCR recognition and data cleaning to obtain the first data; The first data is anonymized to obtain the second data; The second data is parsed into structured vectors to obtain the target data.
[0034] In practical applications, the original data consists of multi-source, heterogeneous medical materials. OCR recognition, cleaning, de-identification, and vectorization can be performed on the original data. Specifically, OCR text extraction can be performed on PDF, image, and DICOM formats from the original data, followed by data cleaning, such as removing blank pages, watermarks, and duplicate pages, to obtain the final data. When extracting text, OCR models specifically trained for the characteristics of medical text (such as special fonts, medical symbols, and complex layouts) can be selected. These models are typically trained on large-scale medical document datasets to improve the accuracy of medical text recognition. For example, the model needs to be adaptable to different formats of data, such as handwritten medical records and scanned medical reports. When the original data is input into the OCR model, it can recognize and extract text information, including body text, titles, and table content. During this process, it is crucial to ensure accurate recognition of various fonts, font sizes, and layout formats, minimizing recognition errors to obtain relatively clean and readable text data, providing an accurate foundation for subsequent structured processing. Furthermore, medical data from different sources varies greatly in format; for example, PDFs may have different layouts, and image quality and resolution also differ. OCR cleaning can convert this diverse data into a unified text format, allowing subsequent structured processing to be based on a consistent data format, improving processing efficiency and consistency. This effectively improves data quality, reduces noise interference, and provides clearer and more accurate data for subsequent processing.
[0035] Subsequently, sensitive information such as patient names and ID numbers can be anonymized. This can be done by replacing, masking, or encrypting sensitive information. Even if the data is illegally obtained during transmission, storage, or use, attackers cannot directly access the true and valid sensitive information, thus reducing the security risks associated with data breaches. Furthermore, data anonymization complies with laws and regulations and ensures that patients' sensitive information is effectively protected during data processing, preventing unauthorized access and disclosure of patient privacy.
[0036] Finally, the anonymized second-party data can be parsed into computable structured vectors to obtain the target data. In practical applications, after obtaining the cleaned and anonymized second-party data, a rule engine can be used to extract and organize the core fields from medical records, examination reports, images, medical orders, and legal provisions based on preset rules and templates. The extracted core fields can then be organized according to a specific structure and format to form a structured rule data file. In practical applications, a vectorization engine can be used to construct structured vectors based on this structured information. In this way, the vectorization engine can perform semantic analysis on the text, converting each word, sentence, or even the entire document into vector form. During the vectorization process, appropriate methods can be designed to preserve the original data's layout and table structure information. For example, layout information (such as paragraph position, bolding, etc.) and table structure (row and column relationships, cell content, etc.) can be integrated into the vector using specific encoding methods, or stored as additional metadata along with the text vector. This preserves the semantic information of the original layout, tables, and medical terminology, generating a unified input format.
[0037] In practical applications, this vectorized data format is "understandable" by computers, making it suitable for subsequent reading and extraction by intelligent agents. When processing data, reading and extraction agents can receive this structured vector data, which has undergone vectorization, has a unified format, and contains rich semantic information. This allows them to extract features and learn patterns more efficiently, thereby completing the task.
[0038] In one possible implementation, the reading extraction agent is constructed in the following way: Using a medical-specific OCR model and vectorization engine, sample data is converted into sample vectors; A medical regulations knowledge base is constructed based on authoritative data, including medical systems, regulations, and medical terminology. Based on the sample vectors and the medical regulations knowledge base, the reading extraction agent is constructed; A prompt word template is constructed for the reading extraction agent, and a verification rule is introduced for the reading extraction agent, wherein the verification rule is used to verify the output of the reading extraction agent.
[0039] In practical applications, all potentially useful original materials, such as paper medical records, imaging films, electronic medical records, and legal provisions, can be listed and recorded. Then, clinical experts, legal experts, and information engineers work together to compile a unified rule manual that includes key fields such as "patient name, age, chief complaint, diagnosis code, surgery name, medication dosage, and timeline," as well as the specific triggering conditions for eighteen core systems, such as "preoperative discussion, death case discussion, and three-level ward round system."
[0040] In the above embodiments, a medical-specific OCR model and vectorization engine can be used to convert sample data into sample vectors. For example, paper sample data can be scanned into PDF or images, input into a medical-specific OCR model for text recognition, and then a self-developed error correction model can be used to correct easily misread parts such as drug names and surgical names. Then, the following operations can be performed: ① Remove blank pages, duplicate pages, and watermarks; ② Standardize the dosage, date, and code to the national standard format; ③ Desensitize sensitive information such as names, ID cards, mobile phone numbers, and facial images to protect privacy without affecting subsequent analysis.
[0041] Subsequently, the cleaned text, tables, and images were converted into sample vectors that computers could understand using different vectorization engines.
[0042] At the same time, authoritative information such as the eighteen core systems, national regulations, guidelines for various specialties, disease codes, and surgical names can all be "broken down and stored in the database." First, it is put into a full-text search engine, and then a knowledge graph is built. Diseases, surgeries, drugs, regulations, and levels of responsibility are all connected by relationships such as "cause-treatment-violation-basis," forming a medical regulatory knowledge base like a "medical identification-specific knowledge network."
[0043] In the above embodiments, during the construction of the reading extraction agent, it can be built based on the generated sample vectors and medical regulations to achieve automatic classification of unstructured data (medical records / examinations / images / regulations), key field extraction, and standardized mapping. Simultaneously, editable prompt templates can be designed for the reading extraction agent, clearly defining the task objective as extracting core information such as the patient's chief complaint, past medical history, diagnosis, surgery, and medication, and outputting structured JSON. For example, 128 sets of prompt templates can be pre-written according to "department × case type × document type." For instance, when extracting key information from a medical record of a thoracic surgery postoperative infection dispute, placeholders can be reserved in the template to automatically fill in the patient's name, surgery name, etc., outputting structured results in one go. This allows the agent to have clear goals and output requirements when processing tasks, improving the accuracy and consistency of task execution. Furthermore, the editable prompt templates allow users to flexibly adjust them according to specific needs, meeting personalized requirements in different scenarios and enhancing the agent's applicability. In practical applications, verification rules can also be introduced for the above-mentioned reading extraction agent, such as medical terminology dictionary and numerical range rules, so as to facilitate real-time verification of the extraction results.
[0044] In one possible implementation, the reading extraction agent is used to extract and standardize information from the target data based on the medical identification task, generating structured data, including: Based on the prompt word template corresponding to the medical identification task, determine the task objective; The target data is categorized according to the task template. According to the rules and models corresponding to the categories of the target data, the key fields of the target data are extracted and standardized mapping is performed to generate structured data; During the processing, the key fields are validated according to the validation rules. If an anomaly is found, the key field is highlighted and the user is prompted to perform a manual review.
[0045] In practical applications, task objectives can be determined based on prompt word templates corresponding to medical assessment tasks. For example, different tasks require different extracted content. Then, automatic classification of medical records (medical records / examinations / images / regulations), key field extraction, and standardized mapping can be implemented. Editable prompt word templates control the output content, ensuring the completeness, accuracy, and compliance with medical assessment standards of key information. Furthermore, output can be formatted accordingly. Simultaneously, validation rules, such as medical terminology dictionaries and numerical range rules, are needed to perform real-time validation of the extracted key fields. Outliers (such as date conflicts or dosage exceeding limits) are highlighted (e.g., automatically highlighted in red) and prompt for manual review. This ensures the accuracy of structured data.
[0046] In one possible implementation, the preliminary analysis includes a summary of the diagnosis and treatment, points of contention, and recommendations on the level of liability. The diagnostic reasoning agent is used to analyze the structured data and generate preliminary analysis opinions corresponding to the medical assessment task, including: The structured data is input into a large medical model to obtain the diagnosis and treatment summary output by the large medical model. The structured data is jointly reasoned based on the rule base and the large reasoning model to determine the disputed nodes; The recommended level of responsibility is generated based on the medical records and legal provisions in the structured data.
[0047] In the above embodiments, preliminary analysis opinions may include a treatment summary, disputed points, and recommendations on the level of responsibility. In practical applications, a large-scale medical model can be invoked to convert structured data input into a draft summary of the treatment process that conforms to the assessment criteria. This summary can be automatically generated in a five-stage format: "admission-preoperative-intraoperative-postoperative-discharge." The large-scale medical model supports selecting output templates by department and case type, and can also output a table version with a timeline and a paragraph text version. This large-scale medical model is used to generate corresponding treatment summaries based on structured data; the specific structure and type of the large-scale medical model are not limited. Then, based on a rule base and a large-scale reasoning model, it can be used to automatically mark disputed points (high-risk points) such as missing preoperative information, failure to conduct consultations, and medication errors, and output doubt tags and evidence location. For example, matching can be performed according to the rule base first; for instance, the eighteen core systems can be compared with the medical records item by item. Once risk points such as "missing preoperative discussion" or "antibiotics exceeding 7 days" are found, they are automatically tagged and a risk score is given. For scores exceeding 0.7, a second confirmation can be made by inputting the data into a large-scale reasoning model. The model's output of disputed nodes can then be used to determine the disputed nodes within the entire structured data. Alternatively, both the rule base and the large-scale reasoning model can be used simultaneously to reason about the structured data, and the results can be combined to identify the disputed nodes. There are no specific restrictions on the method. Furthermore, the structure and type of the large-scale reasoning model are not limited and can be determined based on the actual situation. This approach, compared to manual reasoning, can improve speed and uncover areas not addressed by manual methods, preventing omissions.
[0048] Then, based on the medical records and legal provisions included in the structured data of this medical assessment task, a responsibility level recommendation can be given using a three-part structure of "major premise-minor premise-conclusion," with the specific clauses highlighted in the report for easy expert review. For example, based on the "Medical Accident Grading Standards" and legal provisions, a preliminary responsibility level recommendation (full / major / minor / no responsibility) can be generated, citing the corresponding clauses as evidence. This provides assessment experts with efficient, accurate, and traceable intelligent assistance, significantly improving the efficiency of medical assessments and helping experts quickly form preliminary analytical opinions.
[0049] In one possible implementation, the reviewing agent is used to review the preliminary analysis opinion and generate a credibility tag corresponding to the preliminary analysis opinion, including: The preliminary analysis opinions are reviewed for compliance using a vectorized regulatory knowledge base and a preset rule engine. The compliance review includes regulatory clause matching, terminology standardization check, format specification detection, and logical consistency verification. If the preliminary analysis opinion is abnormal, the abnormality of the preliminary analysis opinion will be displayed in a specific format; Based on the regulatory matching degree obtained from the matching of regulatory clauses, the expert's historical correction rate, and the deviation of similar cases, a credibility label corresponding to the preliminary analysis opinion is generated.
[0050] In practical applications, legal provisions can be segmented into natural paragraphs, vectorized, and stored in a search engine. This creates a vectorized legal knowledge base that supports semantic-level fuzzy retrieval, enabling precise clause location. In the above embodiment, when performing compliance checks on the obtained preliminary analysis opinions, the vectorized legal knowledge base can be used to match the preliminary analysis opinions with legal provisions. Simultaneously, each sentence generated in the preliminary analysis opinions is processed using a preset rule engine to determine if its terminology is standard, its format is complete, and its logic is consistent. In practical applications, a three-color "red, yellow, green" alert system can be used to generate a rectification list with a single click when anomalies are detected. In practical applications, a credibility score of 0-100 can be assigned to the preliminary analysis opinions by comprehensively considering indicators such as legal matching degree, expert historical correction rate, and deviation from similar cases, and then this score can be used as the corresponding credibility label. Alternatively, the credibility label can be determined directly based on the indicators. In practical applications, the aforementioned indicators can also include model confidence, such as the confidence of the medical field large model outputting a diagnosis summary in a diagnostic reasoning agent, and the confidence of the reasoning large model determining disputed nodes. This provides a more comprehensive assessment of the credibility of the preliminary analysis. The specific calculation method for credibility is not limited in this embodiment, and the method for dividing credibility labels can be determined based on actual circumstances. In practical applications, the audit results can be visually displayed using a radar chart. This ensures that the conclusions are legal, compliant, and auditable.
[0051] In one possible implementation, before vectorizing the raw data in the user-uploaded medical assessment information, the method further includes: The medical identification task in the medical identification information is decomposed into at least one sub-task using a directed acyclic graph (DAG). Based on the agent's workload, the subtasks are assigned to the agents corresponding to the subtasks, wherein the agents include: a reading extraction agent, a diagnostic reasoning agent, and an auditing agent.
[0052] In practical applications, a DAG workflow can be used to break down the entire medical assessment task into 15 atomic steps, including OCR, extraction, dispute identification, legal retrieval, liability determination, and expert review. Each step has retry, timeout, and priority strategies. Computational resources can be dynamically allocated based on the agent's capabilities, supporting failure retries, timeout circuit breakers, and priority queues to ensure high availability and efficiency in large-scale concurrent scenarios, guaranteeing sub-second response times even during peak periods. For example, in practical applications, the reading and extraction agent can be deployed using a "dual-tower" strategy: a large 72B model handles complex medical records, while a small 7B model handles simple medical records, with a gating network automatically distributing the workload. This saves computational power while maintaining speed. Thus, when the original data for the medical assessment task is complex, the large model can be used; when it is simple, the small model can be used. During the task allocation process, the A2A protocol can be used for allocation. The A2A protocol (Agent-to-Agent Protocol) is an open standard protocol designed to solve the interoperability problem between different AI agents, enabling them to collaborate securely across platforms and frameworks.
[0053] In one possible implementation, obtaining the expert opinion corresponding to the preliminary analysis opinion includes: Provide experts with a Word-like online editor, which supports real-time revisions and comparisons by multiple users; The system receives the expert's modifications to the preliminary analysis opinions in the online editor, thereby obtaining expert opinions.
[0054] In practical applications, a Word-like online collaborative editor can be provided, supporting features such as difference comparison, annotation tracking, version archiving, and mandatory confirmation of required fields. All modification records can be structurally saved for subsequent continuous model learning and quality improvement, achieving a closed-loop fusion of AI generation and expert wisdom.
[0055] In practical applications, the aforementioned online editor supports character-level and semantic-level dual-mode difference comparison, and also features voice-to-text input and a medical symbol shortcut bar to facilitate expert revisions. It also includes mandatory confirmation of fields requiring review, including analysis conclusions and attribution suggestions. In addition, it provides version tracking functionality, automatically recording all modifications and generating an "Expert-Time-Modified Content" log for subsequent model fine-tuning.
[0056] One possible implementation involves a three-tiered deployment mechanism: "model sandbox—grayscale—full deployment." This mechanism utilizes expert revision data, updated regulatory data, and real-world case feedback for LoRA fine-tuning and reinforcement learning, regularly generating model evaluation reports (accuracy, recall, expert satisfaction) to achieve system self-evolution and risk control. Each expert modification is automatically recorded as a "training sample." Monthly model fine-tuning is automatically triggered, running 100 historical tasks in the sandbox. Only when accuracy, recall, and expert satisfaction all meet the standards is the model officially deployed, forming a closed loop of "getting smarter with use."
[0057] In one possible implementation, during data transmission, field-level anonymization, SM4 encryption, and blockchain audit logs can be implemented throughout the entire data transmission, storage, and inference chain to ensure compliance with the Personal Information Protection Law and the Data Security Law.
[0058] In practical applications, this solution not only provides forensic institutions with efficient, accurate, and traceable medical assessment reports, but also offers standardized access interfaces for third-party medical information systems and legal platforms, promoting the high-quality development of the medical assessment industry towards digitalization, intelligence, and ecological sustainability. Third-party access includes standardized RESTful APIs and gRPC interfaces based on the MCP protocol, supporting hospitals, legal institutions, and insurance systems to quickly access and obtain structured assessment results.
[0059] The above are embodiments of the method proposed in this application. Based on the same inventive concept, this application can provide a medical identification and analysis system. Exemplarily, its structure is as follows: Figure 2 As shown, the medical identification and analysis system includes: a multi-source heterogeneous data foundation construction module, an intelligent reading and key point extraction intelligent agent module, a diagnostic reasoning and document generation intelligent agent module, a regulatory matching and compliance review module, an expert collaborative editing and version tracking module, and a multi-agent orchestration and task scheduling module.
[0060] Among them, the multi-source heterogeneous data foundation construction module is responsible for OCR recognition, key field extraction, data cleaning, desensitization and vectorization processing of multimodal identification data such as PDFs, images, videos, and electronic medical records, generating structured rule data files, and building a medical regulatory knowledge base including eighteen core systems, medical regulations, clinical guidelines and terminology standards, providing a unified, reliable and traceable data foundation for subsequent intelligent analysis.
[0061] Intelligent Reading—Focusing on the intelligent agent module, it relies on the structured data and knowledge base output by the data foundation construction module to build an intelligent reading agent, realizing automatic classification, field extraction, terminology standardization and anomaly detection of medical records, and controlling the output content with editable prompt word templates to ensure that key information is complete, accurate and in line with medical identification standards.
[0062] The Diagnostic Reasoning—Document Generation Intelligent Agent Module, based on key extracted results, calls upon a finely tuned large-scale medical model to automatically generate a summary of the patient's diagnosis and treatment process, disputed point tags, liability level suggestions, and compliance risk warnings. It supports multi-department templates, multi-document merging, and dual-format output of summary / full version, and embeds interpretable reasoning in the form of "syllogisms" to assist experts in quickly forming preliminary assessment opinions.
[0063] The regulatory matching and compliance audit module utilizes a vectorized regulatory knowledge base and rule engine to perform real-time regulatory clause matching, terminology standardization checks, format specification checks, and logical consistency verification on the generated diagnosis and treatment summaries and analysis opinions. It provides risk warnings in a "three-color warning" manner and supports "one-click rectification list" and locating the original regulatory text to ensure that the conclusions are legal, compliant, and auditable.
[0064] The expert collaborative editing and version tracking module provides a Word-like online multi-user collaborative editor that supports difference comparison, annotation tracking, version archiving, and mandatory confirmation of required fields. All modification records are saved in a structured manner for continuous model learning and quality improvement, achieving a closed-loop integration of AI generation and expert wisdom.
[0065] The multi-agent orchestration and task scheduling module, based on the DAG workflow engine and A2A protocol, decomposes complex identification tasks into atomic sub-tasks such as OCR recognition, field extraction, dispute recognition, legal retrieval, responsibility determination, and expert review. It dynamically allocates computing resources according to the agent's capability profile, supports failure retry, timeout circuit breaking, and priority queues, and ensures high availability and high efficiency in large-scale concurrent scenarios.
[0066] Based on the same inventive concept as the above-described method embodiments, this application also provides a medical identification and analysis device, the structure of which is as follows: Figure 3 As shown.
[0067] Figure 3 This is a schematic diagram of the internal structure of a medical identification and analysis device provided in an embodiment of this application. Figure 3 As shown, the device includes: At least one processor 301; And a memory 302 that is communicatively connected to at least one processor; The memory 302 stores instructions that can be executed by at least one processor, which are executed by at least one processor 301 to enable at least one processor 301 to perform the above-described medical identification analysis method.
[0068] In one possible implementation, the processor 301 is capable of performing the following: vectorizing the raw data in the medical identification information uploaded by the user to obtain target data, wherein the medical identification information also includes the medical identification task corresponding to the raw data; inputting the target data and the medical identification task into a reading extraction agent to obtain structured data output by the reading extraction agent, wherein the reading extraction agent is used to extract and standardize the target data based on the medical identification task to generate structured data; inputting the structured data into a diagnostic reasoning agent to obtain preliminary analysis opinions output by the diagnostic reasoning agent, wherein the diagnostic reasoning agent is used to analyze the structured data to generate preliminary analysis opinions corresponding to the medical identification task; inputting the preliminary analysis opinions into a review agent to obtain credibility tags output by the review agent, wherein the review agent is used to review the preliminary analysis opinions and generate credibility tags corresponding to the preliminary analysis opinions; if the credibility tags of the preliminary analysis opinions are inconsistent with preset tags, obtaining expert opinions corresponding to the preliminary analysis opinions; adjusting the preliminary analysis opinions based on the expert opinions to obtain the identification result corresponding to the medical identification task.
[0069] Some embodiments of this application provide corresponding to Figure 1 A non-volatile computer storage medium stores computer-executable instructions, which are configured to execute the aforementioned medical identification and analysis method.
[0070] In one possible implementation, the computer-executable instructions are configured to perform the following: vectorization processing of raw data in user-uploaded medical assessment information to obtain target data, wherein the medical assessment information also includes a medical assessment task corresponding to the raw data; inputting the target data and the medical assessment task into a reading extraction agent to obtain structured data output by the reading extraction agent, wherein the reading extraction agent is used to extract and standardize information from the target data based on the medical assessment task to generate structured data; inputting the structured data into a diagnostic reasoning agent to obtain preliminary analysis opinions output by the diagnostic reasoning agent, wherein the diagnostic reasoning agent is used to analyze the structured data to generate preliminary analysis opinions corresponding to the medical assessment task; inputting the preliminary analysis opinions into a review agent to obtain credibility tags output by the review agent, wherein the review agent is used to review the preliminary analysis opinions and generate credibility tags corresponding to the preliminary analysis opinions; if the credibility tags of the preliminary analysis opinions are inconsistent with preset tags, obtaining expert opinions corresponding to the preliminary analysis opinions; adjusting the preliminary analysis opinions based on the expert opinions to obtain the assessment result corresponding to the medical assessment task.
[0071] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for IoT devices and media are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0072] The systems, media, and methods provided in this application are one-to-one correspondences. Therefore, the systems and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the systems and media will not be repeated here.
[0073] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0074] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0075] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0076] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0077] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0078] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0079] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0080] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0081] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A medical identification and analysis method, characterized in that, The method includes: The original data in the medical assessment information uploaded by the user is vectorized to obtain the target data, wherein the medical assessment information also includes the medical assessment task corresponding to the original data; The target data and the medical identification task are input into the reading and extraction agent to obtain the structured data output by the reading and extraction agent. The reading and extraction agent is used to extract and standardize the target data based on the medical identification task to generate structured data. The structured data is input into the diagnostic reasoning agent to obtain the preliminary analysis opinion output by the diagnostic reasoning agent, wherein the diagnostic reasoning agent is used to analyze the structured data and generate the preliminary analysis opinion corresponding to the medical identification task. The preliminary analysis opinion is input into the review agent to obtain the credibility label output by the review agent. The review agent is used to review the preliminary analysis opinion and generate the credibility label corresponding to the preliminary analysis opinion. If the credibility label of the preliminary analysis opinion is inconsistent with the preset label, the expert opinion corresponding to the preliminary analysis opinion is obtained. The preliminary analysis was adjusted based on the expert opinions to obtain the assessment results corresponding to the medical assessment task.
2. The method according to claim 1, characterized in that, The process of vectorizing the original data in the medical assessment information uploaded by the user to obtain the target data includes: The original data is subjected to OCR recognition and data cleaning to obtain the first data; The first data is anonymized to obtain the second data; The second data is parsed into structured vectors to obtain the target data.
3. The method according to claim 1, characterized in that, The reading extraction agent is constructed in the following way: Using a medical-specific OCR model and vectorization engine, sample data is converted into sample vectors; A medical regulations knowledge base is constructed based on authoritative data, including medical systems, regulations, and medical terminology. Based on the sample vectors and the medical regulations knowledge base, the reading extraction agent is constructed; A prompt word template is constructed for the reading extraction agent, and a verification rule is introduced for the reading extraction agent, wherein the verification rule is used to verify the output of the reading extraction agent.
4. The method according to claim 3, characterized in that, The reading extraction agent is used to extract and standardize information from the target data based on the medical identification task, generating structured data, including: Based on the prompt word template corresponding to the medical identification task, determine the task objective; The target data is categorized according to the task template. According to the rules and models corresponding to the categories of the target data, the key fields of the target data are extracted and standardized mapping is performed to generate structured data; During the processing, the key fields are validated according to the validation rules. If an anomaly is found, the key field is highlighted and the user is prompted to perform a manual review.
5. The method according to claim 1, characterized in that, The preliminary analysis includes a summary of the diagnosis and treatment, points of contention, and recommendations on the level of liability. The diagnostic reasoning agent is used to analyze the structured data and generate preliminary analysis opinions corresponding to the medical assessment task, including: The structured data is input into a large medical model to obtain the diagnosis and treatment summary output by the large medical model. The structured data is jointly reasoned based on the rule base and the large reasoning model to determine the disputed nodes; The recommended level of responsibility is generated based on the medical records and legal provisions in the structured data.
6. The method according to claim 1, characterized in that, The reviewing agent is used to review the preliminary analysis opinions and generate credibility tags corresponding to the preliminary analysis opinions, including: The preliminary analysis opinions are reviewed for compliance using a vectorized regulatory knowledge base and a preset rule engine. The compliance review includes regulatory clause matching, terminology standardization check, format specification detection, and logical consistency verification. If the preliminary analysis opinion is abnormal, the abnormality of the preliminary analysis opinion will be displayed in a specific format; Based on the regulatory matching degree obtained from the matching of regulatory clauses, the expert's historical correction rate, and the deviation of similar cases, a credibility label corresponding to the preliminary analysis opinion is generated.
7. The method according to claim 1, characterized in that, Before performing vectorization processing on the original data in the medical assessment information uploaded by the user, the method further includes: The medical identification task in the medical identification information is decomposed into at least one sub-task using a directed acyclic graph (DAG). Based on the agent's workload, the subtasks are assigned to the agents corresponding to the subtasks, wherein the agents include: a reading extraction agent, a diagnostic reasoning agent, and an auditing agent.
8. The method according to claim 1, characterized in that, The expert opinions obtained corresponding to the preliminary analysis opinions include: Provide experts with a Word-like online editor, which supports real-time revisions and comparisons by multiple users; The system receives the expert's modifications to the preliminary analysis opinions in the online editor, thereby obtaining expert opinions.
9. A medical identification and analysis device, characterized in that, The device includes: At least one processor; And, a memory communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform a medical identification analysis method as described in any one of claims 1-8.
10. A computer storage medium storing computer-executable instructions, characterized in that, When the computer-executable instructions are executed, they implement a medical identification and analysis method as described in any one of claims 1-8.