A gout patient multi-modal risk stratification system based on cue engineering

By constructing a multimodal risk stratification system based on prompting engineering, the problem of unstable risk stratification for gout patients in existing technologies is solved. This enables the effective use of unstructured medical record texts and improves the accuracy of risk stratification, thereby enhancing the interpretability and dynamic adaptability of the system.

CN122638166APending Publication Date: 2026-08-25SUZHOU INST OF BIOMEDICAL ENG & TECH CHINESE ACADEMY OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611106493.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-24
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing risk stratification technologies for gout patients rely on structured indicators or manually entered information, making it difficult to fully utilize unstructured medical record texts such as progress notes, doctor's orders, and chief complaints. This results in unstable risk stratification results and insufficient interpretive basis, especially in long-term management scenarios where information such as changes in treatment adherence and lifestyle triggers are difficult to align and verify on the same timeline.

Method used

A multimodal risk stratification system based on prompting engineering is constructed. By acquiring structured and unstructured data, a patient time-series diagnosis and treatment dataset is generated. Semantic labels are extracted using a large language model, and conflict verification and confidence correction are performed to generate patient time-series multimodal feature vectors to achieve risk stratification.

Benefits of technology

It improved the utilization rate of unstructured medical record text, enhanced the credibility of semantic tags, improved the accuracy of risk stratification and the interpretability of the system, and strengthened dynamic adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122638166A_ABST
    Figure CN122638166A_ABST
Patent Text Reader

Abstract

This invention discloses a multimodal risk stratification system for gout patients based on prompting engineering, belonging to the field of medical informatics and AI-assisted chronic disease management technology. The system acquires and integrates structured diagnostic and treatment data and unstructured medical record text of gout patients through a construction unit, generating a patient time-series diagnostic and treatment dataset; a parsing unit extracts key semantic tags based on a gout diagnostic and treatment knowledge base and a dynamic prompting template-driven large language model; a calibration unit performs directional conflict verification between semantic tags and laboratory indicators, diagnostic codes, medication records, and adjacent medical records, generating confidence correction values; a stratification unit performs time axis alignment and weighted coding to obtain risk stratification results; and a feedback unit generates explanatory information and risk evolution records, and updates the prompting template in reverse. This system can improve the accuracy, interpretability, and dynamic adaptability of risk stratification for gout patients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical data processing technology, and in particular to a multimodal risk stratification system for gout patients based on cue engineering. Background Technology

[0002] Gout is a chronic metabolic disease closely related to hyperuricemia. Its clinical management typically relies on a variety of information, including the patient's serum uric acid levels, renal function indicators, joint symptoms, frequency of previous attacks, comorbidities, medication use, and lifestyle. Existing medical information systems usually record patient data through electronic medical records, laboratory information systems, prescription systems, and follow-up systems. In clinical decision support scenarios, rule-based judgment, statistical scoring, or machine learning models are used to assess the risk of recurrence, disease progression, or complications. In recent years, with the development of natural language processing and large language models, some systems have begun to extract clinical semantic information from free text in medical records, such as chief complaint, present medical history, treatment adherence, dietary triggers, descriptions of joint pain, and records of tophi. This textual information is then combined with structured indicators for chronic disease patient status identification, risk prediction, and individualized management.

[0003] However, existing risk stratification technologies for gout patients still primarily rely on structured indicators or manually entered information, making it difficult to fully utilize the implicit clinical semantic content in medical records, doctor's orders, chief complaints, and present medical history. Even when natural language processing or large language models are introduced, problems such as insufficient disease constraints, lack of original text evidence to support semantic extraction results, unclear temporal correspondences between data from different sources, and difficulty in automatically calibrating when there are conflicts between textual semantics and laboratory indicators or medication records often exist. Especially in the context of long-term gout management, changes in treatment adherence, lifestyle triggers, joint attack status, tophi progression, and renal function risk often exhibit significant temporal evolution characteristics. If the system cannot align, verify, and process the credibility of multimodal information on the same timeline, it can easily lead to unstable risk stratification results, insufficient interpretive basis, and affect the targeting and continuity of subsequent clinical interventions.

[0004] Therefore, it is necessary to propose a multimodal risk stratification technology that is more suitable for the long-term management of gout patients. Summary of the Invention

[0005] This application provides a multimodal risk stratification system for gout patients based on prompting engineering, in order to improve the accuracy, interpretability and dynamic adaptability of risk stratification for gout patients.

[0006] This application provides a multimodal risk stratification system for gout patients based on prompting engineering, including: The building unit is used to acquire structured diagnosis and treatment data and unstructured medical record text of gout patients, standardize the structured diagnosis and treatment data and integrate the consultation time to generate a patient time-series diagnosis and treatment dataset; The parsing unit is used to generate disease-specific constraint prompts based on the gout diagnosis and treatment knowledge base, the patient time-series diagnosis and treatment dataset, and the dynamic prompt template. This drives the large language model to extract treatment compliance, lifestyle triggers, joint attack status, description of tophi, and description of renal function risk from unstructured medical record text, and generate a set of structured semantic tags carrying original text evidence fragments and consultation time. The calibration unit is used to perform directional conflict verification between the structured semantic tag set and laboratory indicators, diagnostic codes, medication records and adjacent medical records, and generate confidence correction values ​​based on the conflict source and original evidence fragments to form a calibration semantic tag set. The hierarchical unit is used to align the patient time-series diagnosis and treatment dataset, calibration semantic label set, laboratory indicators, diagnostic codes and medication records along the same time axis, and perform weighted encoding based on confidence correction values ​​to generate patient time-series multimodal feature vectors, and obtain risk stratification results of low risk, medium risk or high risk; The feedback unit is used to generate feature contribution explanation information and risk evolution records aggregated by semantic category based on risk stratification results and patient temporal multimodal feature vectors, and feeds back conflict sources and feature contribution explanation information to the parsing unit. It also updates the extraction constraints of dynamic prompt templates for semantic categories where both contribution and conflict frequency exceed preset thresholds.

[0007] This application has the following beneficial technical effects: (1) Improve the utilization rate of unstructured medical record text. (2) Enhance the credibility of semantic tags through directional conflict verification. (3) Improve the accuracy of risk stratification based on confidence-weighted coding. (4) Enhance the interpretability and dynamic adaptability of the system by updating the prompt template through explanation feedback. Attached Figure Description

[0008] Figure 1 This is a schematic diagram of a multimodal risk stratification system for gout patients based on prompting engineering, provided in the first embodiment of this application. Detailed Implementation

[0009] Many specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application; therefore, this application is not limited to the specific embodiments disclosed below.

[0010] The first embodiment of this application provides a multimodal risk stratification system for gout patients based on prompting engineering. Please refer to... Figure 1This figure is a schematic diagram of the first embodiment of this application. The following is in conjunction with... Figure 1 The first embodiment of this application provides a multimodal risk stratification system for gout patients based on prompting engineering.

[0011] The gout patient multimodal risk stratification system based on prompting engineering includes a construction unit 101, a parsing unit 102, a calibration unit 103, a stratification unit 104, and a feedback unit 105.

[0012] Unit 101 is used to acquire structured diagnosis and treatment data and unstructured medical record text of gout patients, standardize the structured diagnosis and treatment data and integrate the consultation time to generate a patient time-series diagnosis and treatment dataset.

[0013] Building unit 101 is used to generate a patient time-series medical data set for subsequent analysis, calibration, and stratification before the system begins risk stratification of gout patients. The structured medical data referred to here is data stored in the form of fields in hospital information systems, electronic medical record systems, laboratory information systems, prescription systems, or follow-up systems, which can be directly read by field name and field value. This data may include at least patient anonymity identifiers, gender, age, date of visit, diagnosis code, serum uric acid, creatinine, estimated glomerular filtration rate, blood urea nitrogen, C-reactive protein, complete blood count indicators, number of previous gout attacks, presence of hypertension or diabetes, medication name, dosage, frequency of administration, prescription start and end dates, and follow-up visit records. The unstructured medical record text referred to here is text content written in natural language that cannot directly express clinical meaning using fixed fields. It may include chief complaint, present illness, past medical history, progress notes, physical examination records, doctor's orders, discharge summary, follow-up records, and doctors' descriptions of diet, alcohol consumption, pain location, joint swelling, tophi, and compliance. When building unit 101 acquires the above data, it can do so through hospital data interfaces, database views, batch file import, or authorized data exchange interfaces. It uses the patient's unique identifier or an anonymized patient ID as the association key to link data of the same patient in different systems to the same patient record.

[0014] To meet the consistency requirements of subsequent processing, construction unit 101 needs to standardize the structured medical data. Standardization refers to processing data from different sources, with different field names, and different units of measurement into an expression form that the system can uniformly recognize. For example, the blood uric acid index may be recorded as "UA" or "uric acid" in different systems; construction unit 101 uniformly maps it to the blood uric acid field. The blood uric acid unit may be μmol / L or mg / dL; the system uniformly converts it to a preset unit. The consultation date, test date, and prescription date may use different date formats; the system uniformly converts them to a sortable time format of year, month, day, and time (minutes and seconds). For missing values, construction unit 101 can set missing markers according to the field type, instead of simply deleting the entire record. For example, if there is no C-reactive protein test result for a certain consultation, the field is marked as "not tested" so that the subsequent stratification unit 104 can identify the source of the missing data. For obviously unreasonable outliers, the construction unit 101 can mark them according to the preset medical range or data quality rules. For example, a negative blood uric acid value, a consultation date later than the data export date, or a prescription end time earlier than the prescription start time can all be marked as pending verification.

[0015] Construction unit 101 also integrates the structured medical data by visit time. This visit time integration refers to merging data from different sources into corresponding visit or follow-up nodes based on the patient's actual visit time. A visit node can be understood as a data set formed during a patient's outpatient, emergency, hospitalization, or follow-up visit. Its time can be determined by the registration time, medical record creation time, prescription issuance time, or discharge time. When the same patient has laboratory, prescription, and medical history records within a short period, construction unit 101 can merge them into the same visit node according to a preset time window. For example, if a patient has an outpatient visit on March 1, 2026, receives a febuxostat prescription on the same day, and completes blood uric acid and creatinine tests on March 2, 2026, and the system sets the outpatient laboratory test merging window to one to three days before and after the visit date, then the test result can be assigned to the visit node corresponding to March 1, 2026. If the patient has a follow-up visit on April 15, 2026, then construction unit 101 will create a new visit node for this visit and arrange it after the previous node according to the time sequence. In this way, the system can form a continuous record of the patient's disease progression over time, rather than an isolated data table.

[0016] The patient time-series medical record dataset is the output of construction unit 101 and the foundational data used by subsequent parsing unit 102, calibration unit 103, and hierarchical unit 104. This dataset includes at least the patient's anonymous ID, multiple visit nodes arranged chronologically, structured indicators for each visit node, diagnostic information, medication records, and an index of the unstructured medical record text associated with that node. To facilitate subsequent large language model parsing, construction unit 101 can retain the original text content, source, creation time, and associated visit node of the unstructured medical record text in the dataset, but does not make medical semantic judgments at this stage. This approach aims to enable parsing unit 102 to generate disease-specific constraint prompts within a clear patient timeline and medical context, and to avoid mixing and interpreting text content from different visit times. For example, if a patient records "redness, swelling and pain in the left first metatarsophalangeal joint for the past week" in the initial consultation text and "no recurrence recently" in the follow-up consultation text three months later, the construction unit 101 will classify the two texts into different consultation nodes through time integration, so that the subsequent system can recognize that the gout activity status changes over time, and will not mistakenly treat the two descriptions as the same time status.

[0017] In a specific implementation process, the construction unit 101 can first read a patient's three medical records. The first visit includes a serum uric acid level of 560 μmol / L, a diagnosis of gout, a prescription of colchicine and nonsteroidal anti-inflammatory drugs, and is associated with the chief complaint text "redness, swelling and pain in the left big toe joint for three days"; the second visit includes a serum uric acid level of 430 μmol / L, a prescription of febuxostat, and is associated with the disease progress text "pain relieved, occasional alcohol consumption"; the third visit includes a serum uric acid level of 360 μmol / L, an elevated creatinine record, and a follow-up visit text "regular medication, no obvious joint swelling or pain". After uniformly processing the serum uric acid, drug name, date format, and diagnosis code, the construction unit 101 establishes three consecutive nodes according to the three visit dates, and stores the corresponding indicators, diagnosis, medication, and text index under each node, ultimately generating the patient's time-series medical data set. The subsequent parsing unit 102 can identify lifestyle triggers and treatment adherence at different time points based on the dataset, the calibration unit 103 can verify the text semantics with laboratory indicators and medication records of the same or adjacent nodes, and the stratification unit 104 can form multimodal feature vectors based on a continuous time axis. Therefore, the construction unit 101 not only completes data collection, but also provides a unified, traceable, and computable time-series data foundation for the entire risk stratification process.

[0018] The parsing unit 102 is used to generate disease-specific constraint prompts based on the gout diagnosis and treatment knowledge base, the patient time-series diagnosis and treatment dataset, and the dynamic prompt template. This drives the large language model to extract treatment adherence, lifestyle triggers, joint attack status, description of tophi, and description of renal function risk from unstructured medical record text, and generate a set of structured semantic tags carrying original text evidence fragments and consultation time.

[0019] The parsing unit 102, based on the patient time-series diagnosis and treatment dataset already formed by the construction unit 101, performs semantic extraction on unstructured medical record text for gout risk stratification, and outputs a set of structured semantic labels that can be directly used by the subsequent calibration unit 103 and stratification unit 104. The cueing engineering referred to here means that before calling the large language model to process the medical record text, cue words containing task objectives, medical concept boundaries, output format, prohibited inferences, evidence citation requirements, and time attribution requirements are constructed according to preset rules, enabling the large language model to complete the text parsing task for a specific disease within a constrained scope. The large language model refers to a language model capable of receiving natural language text input and outputting natural language or structured results. It can be deployed on a local server, a hospital private cloud, or a compliant and authorized computing environment. This invention does not limit the specific model name, as long as it can identify semantic information from the medical record text based on the cue words and output it according to the specified format.

[0020] The gout diagnosis and treatment knowledge base used by parsing unit 102 is a collection of disease knowledge established to limit the scope of semantic extraction. It can include gout-related symptom words, sign words, disease course description words, medication names, lifestyle triggers, complication-related descriptions, and clinical expressions related to renal function risk. For example, joint attack status can be associated with expressions such as "redness, swelling, heat, and pain," "first metatarsophalangeal joint pain," "sudden onset of pain at night," "acute attack," "pain relief," and "no recurrence"; lifestyle triggers can be associated with expressions such as "alcohol consumption," "high-purine diet," "seafood," "animal organs," "hot pot," "staying up late," and "dehydration after exercise"; treatment adherence can be associated with expressions such as "regular medication," "discontinuing medication on one's own," "intermittent medication," "failure to follow up with doctor's orders," "missed dose," and "refusal to lower uric acid treatment"; descriptions of tophi can be associated with expressions such as "auricular nodules," "periarticular induration," "white crystalline substances," and "tophi ulceration"; and descriptions of renal function risk can be associated with expressions such as "elevated creatinine," "abnormal renal function," "chronic kidney disease," "positive urine protein," and "decreased eGFR." The aforementioned knowledge base can be compiled from medical guidelines, hospital treatment standards, expert rules, and historical medical record annotation samples, and can be stored in the form of a thesaurus, a thesaurus, a category tree, or a rule table.

[0021] The dynamic prompt template is the basic template for the parsing unit 102 to generate disease-specific constraint prompts. It is not a fixed, unchanging question, but a task template containing multiple replaceable fields. Replaceable fields can include the time of the patient's current visit node, the structured medical summary of that node, the brief status of adjacent visit nodes, the semantic category to be extracted, the decision boundary for each semantic category, the output field name, the requirements for extracting evidence fragments, and the restriction that content not appearing cannot be supplemented based on common sense. The dynamic prompt template is called dynamic because it can adjust the prompt constraints based on the specific content in the patient's time-series medical dataset and the conflict sources and feature contribution explanations returned by the feedback unit 105. For example, if subsequent feedback shows that the "treatment adherence" category frequently conflicts due to phrases like "discontinuing medication after pain relief" in the text, the parsing unit 102 can add the restriction "If the text contains descriptions of discontinuing medication, intermittent medication, or self-reducing dosage, it should be prioritized as adherence risk, and should not be marked as good adherence solely because of symptom relief" to the extraction constraints of the treatment adherence category in the next prompt generation.

[0022] The disease-specific constraint prompt is generated by the parsing unit 102 based on the gout diagnosis and treatment knowledge base, the patient time-series diagnosis and treatment dataset, and the dynamic prompt template. This prompt must include at least the following meanings: the large language model must extract semantics only for risk stratification of gout patients; each extraction result must be assigned its semantic category, standardized label value, consultation time, original text evidence fragment, and initial confidence level; text descriptions from different consultation nodes cannot be mixed into the same time state; and for content without direct textual evidence, the output must be "not mentioned" or "uncertain," and inferences based on common sense are not allowed. The original text evidence fragment refers to the smallest continuous text segment or short sentence in the original medical record that can support a certain semantic label. Its function is to enable the subsequent calibration unit 103 to trace the source of the semantic label. For example, if the medical record states "swelling and pain in the left big toe joint after drinking alcohol multiple times in the past half month," then "multiple drinking alcohol" and "swelling and pain in the left big toe joint" can be used as original text evidence fragments for lifestyle triggers and joint attack status, respectively. The consultation time refers to the consultation node time corresponding to the medical record text to which the evidence fragment belongs, and can be directly inherited from the time attribute established by the construction unit 101 for this text.

[0023] During execution, parsing unit 102 first reads the unstructured medical record text corresponding to a specific visit node from the patient's time-series medical data set, and simultaneously reads the structured summaries of that node and necessary adjacent nodes. The adjacent node summaries are used only for time constraints and contextual cues, and are not used to replace evidence in the current text. Subsequently, parsing unit 102 selects terms, synonyms, and negations related to five semantic targets from the gout diagnosis and treatment knowledge base and fills them into a dynamic prompt template, forming disease-specific constraint prompts for the current visit node. The large language model receives these prompts and the text to be parsed, and outputs a structured result. To facilitate subsequent processing, parsing unit 102 can require the output to be in table, key-value pair, or JSON format, including at least the semantic category, label value, evidence fragment, visit time, and initial confidence level. The initial confidence level can be high, medium, or low, or it can be a numerical level. If high, medium, and low confidence levels are used, high indicates that there is clear and direct evidence in the text, such as "the patient discontinued febuxostat on their own for two weeks"; medium indicates that there is indirect but relatively clear evidence in the text, such as "not taking medication regularly"; and low indicates that the text description is vague or requires further calibration, such as "recent control is average". This initial confidence level is only a preliminary evaluation of the clarity of the text during the parsing phase. The subsequent confidence level correction value is generated by the calibration unit 103 in combination with laboratory indicators, diagnostic codes, medication records, and adjacent medical records.

[0024] To avoid unfounded expansions in the large language model, the parsing unit 102 preferably sets negative and uncertain recognition rules. Negative recognition means that when expressions such as "no tophi," "no joint swelling," "denies drinking alcohol," or "no recurrence" appear in the text, the parsing unit 102 should label the corresponding category as negative or low risk, rather than extracting it as positive risk. Uncertain recognition means that when expressions such as "possible," "suspected," "to be investigated," or "patient's self-report is unclear" appear in the text, the parsing unit 102 should label the value as uncertain and lower the initial confidence level. The above processing can be explicitly requested from prompt words to require the large language model to recognize negative words, time words, and uncertain words, or the parsing unit 102 can perform a secondary check using rules after the model output.

[0025] For example, for a patient's chief complaint and medical history text on March 1, 2026: "Redness, swelling, and pain in the left first metatarsophalangeal joint for three days; alcohol consumption for two consecutive nights prior to the attack; irregular febuxostat use; no identifiable tophi on physical examination," the structured semantic tag set generated by parsing unit 102 can include: joint attack status as "acute attack," original evidence fragment as "redness, swelling, and pain in the left first metatarsophalangeal joint for three days," and consultation date as March 1, 2026; lifestyle trigger as "alcohol consumption," original evidence fragment as "drinking alcohol for two consecutive nights"; treatment adherence as "irregular medication use," original evidence fragment as "irregular febuxostat use"; and tophi description as "no identifiable tophi found," original evidence fragment as "no identifiable tophi found." If the same text does not contain any description related to renal function, the renal function risk description output is "not mentioned," and no positive risk tag is generated. This structured semantic tag set preserves the implicit clinical information in the medical record text without making excessive judgments based on the original evidence.

[0026] For example, given the text "The patient has been receiving regular uric acid-lowering treatment and has not experienced significant joint swelling or pain in the past three months, but this follow-up examination indicates that creatinine has increased compared to before," the parsing unit 102 can generate semantic tags for treatment adherence as "regular treatment," joint flare status as "no recent flare-ups," and renal function risk description as "creatinine has increased compared to before," each carrying the corresponding original text evidence fragment and the follow-up examination time. Since "creatinine has increased compared to before" involves comparison with previous indicators, the parsing unit 102 can first output renal function risk-related tags based on the text evidence. Whether it is consistent with laboratory indicators is then handled by the calibration unit 103, which performs directional conflict verification by combining the actual creatinine test value with adjacent medical records.

[0027] The structured semantic label set output by parsing unit 102 should maintain a correlation with the patient's time-series medical record dataset. Each semantic label can include the patient's anonymity ID, visit node ID, visit time, text source, semantic category, label value, original text evidence fragment, and initial confidence level. Through this data structure, calibration unit 103 can determine which visit and which medical record text each semantic label comes from, hierarchical unit 104 can co-encode the calibrated semantic labels with laboratory indicators, diagnostic codes, and medication records on the same timeline, and feedback unit 105 can update the cue constraints for specific semantic categories based on conflict frequency and feature contribution.

[0028] Furthermore, the parsing unit is specifically used for: Based on the patient's time-series medical records, extract the unstructured medical record text, consultation time, blood uric acid status, uric acid-lowering drug record, acute attack-related diagnosis, and state summary of the previous and next adjacent consultation nodes from the current consultation node, and generate a node parsing input package. Based on the template version number returned by the gout diagnosis and treatment knowledge base and feedback unit, positive expression, negative expression, uncertain expression, time attribution rules, cross-node inference prohibition rules and original text evidence extraction rules are configured for treatment compliance, lifestyle triggers, joint attack status, description of tophi and description of renal function risk, respectively, and a semantic category constraint table is generated. The node parsing input package, semantic category constraint table and dynamic prompt template are combined into disease constraint prompt words, so that the large language model can generate candidate semantic labels only based on the current consultation node text or the original text content with a clear time reference. Each candidate semantic label is required to carry the label value, consultation time, original text evidence fragment and initial confidence level. The candidate semantic tags are verified for evidence integrity and time attribution. Tags that do not carry original text evidence fragments and do not belong to the unmentioned state are deleted, and tags whose evidence fragment time points are inconsistent with the current medical treatment node are marked as cross-node suspected tags. Candidate semantic tags that have not been deleted, suspected cross-node tags, and their corresponding original text evidence fragments are written into a structured semantic tag set according to the visit node, and suspected cross-node tags are provided to the calibration unit as priority verification objects for directional conflict verification.

[0029] The parsing unit, after the construction unit has formed a patient time-series diagnosis and treatment dataset, converts the unstructured medical record text in each visit node into structured semantic labels with clear semantic categories, time attribution, and evidence sources. A visit node, as referred to here, is a data unit formed according to the patient's actual visit, follow-up visit, hospitalization, or follow-up time; each visit node corresponds to one diagnosis and treatment process and its related data. The parsing unit first extracts the unstructured medical record text, visit time, blood uric acid status, uric acid-lowering drug record, acute attack-related diagnosis, and status summary of the previous and next adjacent visit nodes for the current visit node, generating a node parsing input package. The blood uric acid status can include states such as reached target, not reached target, increased compared to before, decreased compared to before, or not detected; the uric acid-lowering drug record can include prescriptions, renewals, interruptions, or discontinuations of drugs such as febuxostat, allopurinol, and benzbromarone; the acute attack-related diagnosis can include diagnostic codes or names such as acute gouty arthritis and acute gout attack. The state summary of adjacent visit nodes is only used to indicate the disease background of the current text in the model. For example, the previous node is "medication treatment after acute attack" and the next node is "the blood uric acid decreased during follow-up visit". However, it cannot replace the current node text as direct evidence of semantic label.

[0030] The node parsing input package is a structured input provided by the parsing unit to the subsequent prompt word generation process. It includes at least the current visit node number, visit time, current node medical record text, current node structured state, and adjacent node summaries. This setup aims to ensure that the large language model, when parsing the medical record text, understands the time location and clinical context of the current text, without incorrectly interpreting the states of adjacent nodes as the current node's diagnosis. For example, if the current node text states "no obvious joint pain today," while the previous node summary states "redness, swelling, and pain in the first metatarsophalangeal joint of the left foot a week ago," the parsing unit will require the model to generate the current node's semantic label based solely on the current node text in subsequent prompt words. It should not directly label the current node as an acute attack simply because the previous node has an attack record.

[0031] The parsing unit further generates a semantic category constraint table based on the gout diagnosis and treatment knowledge base and the template version number returned by the feedback unit. The template version number is a version identifier generated by the feedback unit after updating and verifying the dynamic prompt template. The parsing unit calls the corresponding version of the prompt template based on this version number to ensure the traceability of the subsequent parsing process. The semantic category constraint table is used to limit the recognition boundaries of various gout-related semantics for the large language model, including at least five categories: treatment adherence, lifestyle triggers, joint attack status, description of tophi, and description of renal function risk. For each semantic category, the semantic category constraint table is configured with affirmative expressions, negative expressions, uncertain expressions, time attribution rules, rules prohibiting cross-node inference, and rules for extracting original text evidence. Affirmative expressions are textual expressions that directly support the validity of a certain label, such as "regular medication," "discontinued medication on one's own," "attack after drinking alcohol," "joint redness, swelling, and pain," "tophi found," and "elevated creatinine." Negative expressions are textual expressions that directly deny the validity of a certain label, such as "denies drinking alcohol," "no tophi found," "no recurrence," and "no obvious joint swelling or pain." Uncertain expressions refer to textual expressions whose label status cannot be directly determined, such as "medication is okay", "occasionally uncomfortable", "suspected gout stones", "kidney function needs to be re-examined", etc.

[0032] The time-based attribution rule determines which visit node a semantic description in the medical record text should be assigned to. If there is no specific time indication in the text, the semantic content belongs to the current visit node. If expressions such as "last attack," "three months ago," or "previously present" appear in the text, the parsing unit requires the large language model to label them as historical descriptions and not directly as the current state of the current node. The rule prohibiting cross-node inference means that the large language model cannot generate the semantic label of the current node solely based on the state summary of adjacent visit nodes, nor can it directly replace the pain description of the previous node, the follow-up results of the next node, or medication information at other time points with the conclusion of the current node. The original text evidence extraction rule requires that each candidate semantic label must carry an original text fragment that supports the label, and this fragment should preferably be the smallest continuous short sentence in the medical record text, such as "stopped medication on its own for two weeks," "left foot red, swollen, and painful after drinking alcohol," or "no tophi seen," rather than the entire medical record text.

[0033] After forming the node parsing input package and the semantic category constraint table, the parsing unit combines them with the dynamic prompt template to form disease-specific constraint prompts. These prompts explicitly define the task scope, usable textual basis, output format, and prohibited behaviors of the large language model, ensuring that the model generates candidate semantic labels based solely on the current consultation node text or original text content with a clear time reference. For each candidate semantic label, the parsing unit requires it to simultaneously include a label value, consultation time, original text evidence fragment, and initial confidence level. The label value represents the specific semantic outcome; for example, treatment adherence is defined as "regular medication," "intermittent medication," "discontinued medication on one's own," or "not mentioned"; lifestyle triggers are defined as "alcohol consumption," "high-purine diet," "not mentioned," or "uncertain"; and joint attack status is defined as "acute attack," "symptom relief," "no recent attack," or "uncertain." The initial confidence level indicates the clarity with which the model arrives at the label based on the original text evidence fragment, and can be high, medium, or low. For example, "the patient stopped taking febuxostat on his own for two weeks" corresponds to a high confidence level, while "the patient's medication is working well" is a vague statement and corresponds to a low or moderate confidence level.

[0034] After obtaining candidate semantic tags, the parsing unit does not directly write them all into the structured semantic tag set. Instead, it performs evidence integrity checks and time attribution checks. Evidence integrity checks involve checking whether each candidate semantic tag carries original text evidence fragments and whether the corresponding content of those original text evidence fragments can be found in the current medical record text. If a candidate semantic tag indicates the presence or absence of a certain risk factor but does not provide original text evidence fragments, and the tag is not in the "not mentioned" state, the parsing unit deletes the tag. For example, if the large language model outputs "lifestyle triggers are alcohol consumption," but there are no original text fragments about alcohol consumption in the current text, the tag is deleted. If the output is "lifestyle triggers not mentioned," since it indicates that there is no corresponding content in the text, it is not required to carry positive evidence fragments, but the category can still be recorded as "not mentioned."

[0035] Time attribution verification refers to checking whether the time indicated by the original evidence fragment is consistent with the current medical visit node. If the evidence fragment clearly belongs to the current medical visit process, such as "visited for left foot pain" or "attacked recently after drinking alcohol," then the tag is assigned to the current node. If the evidence fragment indicates that it belongs to other times, such as "took medication for gout attack last month," "discovered tophi three years ago," or "elevated creatinine during the last follow-up visit," then the tag is not directly used as the definitive tag for the current node, but is marked as a suspected cross-node tag. The suspected cross-node tag refers to a semantic tag that appears in the current medical record text but whose time attribution is inconsistent with the current medical visit node and may affect subsequent calibration judgments. Suspected cross-node tags are not simply deleted, as they may be of reference value for judging the course of the disease, but need to be prioritized for verification by the calibration unit in conjunction with adjacent medical records, laboratory indicators, and medication records.

[0036] For example, a patient's current visit text reads, "Follow-up visit today; no significant joint redness, swelling, or pain in the past two weeks; had an attack last month after drinking alcohol; currently taking febuxostat regularly." The parsing unit can use "no significant joint redness, swelling, or pain in the past two weeks" as evidence of the current node's joint attack status, generating a candidate semantic label of "no recent attacks or stable symptoms"; it can use "currently taking febuxostat regularly" as evidence of the current node's treatment adherence, generating a label of "regular medication use"; however, "had an attack last month after drinking alcohol," although containing both lifestyle triggers and attack status, clearly points to last month. Therefore, the parsing unit marks it as a suspected cross-node label and provides it to the calibration unit for priority verification to determine whether it corresponds to the previous adjacent visit node or a historical attack record.

[0037] Finally, the parsing unit writes the candidate semantic tags that were not deleted, the suspected cross-node tags, and their corresponding original text evidence fragments into a structured semantic tag set according to the visit node. This structured semantic tag set includes at least the patient anonymity identifier, visit node, semantic category, tag value, visit time, original text evidence fragment, initial confidence level, and a marker indicating whether it is a suspected cross-node tag. For suspected cross-node tags, the parsing unit also provides them to the calibration unit as priority verification objects for directional conflict checking, enabling the calibration unit to prioritize determining whether there is a time mismatch or directional conflict between the tag and laboratory indicators, diagnostic codes, medication records, and adjacent visit records. Through the above processing, the parsing unit can introduce gout disease constraints, time node constraints, evidence integrity constraints, and feedback template version constraints while parsing medical record text using a large language model, thereby transforming unstructured medical record text into traceable, verifiable, and usable structured semantic data for subsequent risk stratification.

[0038] The calibration unit 103 is used to perform directional conflict verification between the structured semantic tag set and laboratory indicators, diagnostic codes, medication records and adjacent medical records, and generate confidence correction values ​​based on the conflict source and original evidence fragments to form a calibration semantic tag set.

[0039] The calibration unit 103 is used to correct the credibility of the structured semantic tag set generated by the parsing unit 102, ensuring that the semantic information extracted from the unstructured medical record text is consistent with the patient's objective medical data, medication records, and pre- and post-visit status before entering the stratification unit 104. Because unstructured medical record texts may contain brief doctor records, incomplete patient descriptions, inconsistencies between the time of disease progression descriptions and examination results, and biases in the understanding of semantic boundaries by the large language model, the semantic tags obtained by the parsing unit 102 cannot be directly equated with the final credible tags. Therefore, the calibration unit 103, while retaining the original text evidence fragment, performs directional conflict verification on the consistency between the semantic tags and other medical evidence, and generates confidence correction values ​​accordingly, forming a calibrated semantic tag set for subsequent risk stratification.

[0040] The directional conflict referred to here means that the direction of risk change, disease state, or behavioral state expressed by the structured semantic label is inconsistent with the direction reflected by laboratory indicators, diagnostic codes, medication records, or adjacent medical records. This conflict is not simply a matter of whether two fields are equal, but rather a matter of whether different pieces of evidence point to the same clinical meaning. For example, if the structured semantic label shows "regular medication," but the prescription record for the same period shows that long-term uric acid-lowering drugs were not renewed, and subsequent blood uric acid levels continued to rise, then there is a directional conflict of adherence between the semantic label and the medication record and laboratory indicators. Similarly, if the semantic label shows "no recent attacks," but the diagnostic code for the same medical visit includes acute gouty arthritis, and the prescription record shows short-term treatment with colchicine or nonsteroidal anti-inflammatory drugs, then there is a directional conflict of joint attack state between the semantic label and the diagnostic code and medication record. Furthermore, if the semantic label shows "reduced risk of renal function," but creatinine levels are elevated or the estimated glomerular filtration rate is decreased, then there is a directional conflict of renal function risk between the semantic label and laboratory indicators. By defining directional conflict, calibration unit 103 is able to handle the problem of inconsistent representations of different types of data, rather than relying on exactly the same text or fields.

[0041] The calibration unit 103 receives inputs including at least a structured semantic label set, laboratory indicators from the patient's time-series medical data set, diagnostic codes, medication records, and adjacent medical visit records. Laboratory indicators may include serum uric acid, creatinine, estimated glomerular filtration rate, blood urea nitrogen, C-reactive protein, white blood cell count, and other indicators related to gout activity, inflammation, and renal function. Diagnostic codes may include diagnostic information such as gout, hyperuricemia, gouty arthritis, tophi, and chronic kidney disease. Medication records may include uric acid-lowering drugs, anti-inflammatory analgesics, glucocorticoids, urine alkalizing drugs, and related start and end dates of medication. Adjacent medical visit records refer to one or more medical visit records within a preset time range before or after the current visit, used to determine whether the state represented by the semantic label conforms to the continuous changes in the patient's disease course. The preset time range can be set according to the hospital's follow-up visit cycle and gout management needs. For example, in an outpatient follow-up visit scenario, it can be set to the most recent one or two medical visit records within ninety days before or after the current visit date; in an inpatient scenario, it can be set to the medical course records before and after the same hospitalization period.

[0042] In practice, calibration unit 103 can first determine the semantic category of each structured semantic tag and configure corresponding sources of verification evidence for different semantic categories. For example, the treatment adherence tag is mainly verified against prescription renewal status, medication notes, long-term changes in blood uric acid, and medication descriptions in adjacent medical records; the lifestyle trigger tag is mainly verified against original evidence fragments, descriptions of alcohol consumption and diet in medical records, and short-term attack records; the joint attack status tag is mainly verified against diagnostic codes, anti-inflammatory and analgesic drug prescriptions, inflammatory markers, physical examination descriptions, and adjacent medical records; the tophi description tag is mainly verified against physical examination records, imaging or ultrasound-related records, diagnostic codes, and preceding and following text descriptions; and the renal function risk description tag is mainly verified against creatinine, estimated glomerular filtration rate, proteinuria, chronic kidney disease diagnosis, and related medication adjustment records. By configuring different verification evidence for different semantic categories, the calibration process can better reflect the actual meaning of gout diagnosis and treatment data.

[0043] The confidence correction value is a numerical value or level obtained by the calibration unit 103 after adjusting the confidence level of the semantic tag based on the directional conflict verification result, and is used for weighted encoding by the subsequent layering unit 104. In this embodiment, the confidence correction value is represented by a value between 0 and 1, and is used for weighted encoding by the subsequent layering unit 104. For ease of implementation, the system converts the initial confidence level output by the parsing unit 102 into a base value, where high confidence corresponds to 1.0, medium confidence corresponds to 0.7, and low confidence corresponds to 0.4. The calibration unit 103 corrects the base value based on the direction of the candidate verification evidence: when there is no contrary evidence and at least one supporting piece of evidence, the confidence correction value increases by 0.1, with a maximum of 1.0; when there is one piece of contrary evidence, the confidence correction value decreases by 0.2; when there are two or more pieces of contrary evidence, the confidence correction value decreases by 0.4; when the original evidence fragment contains uncertain words such as "possible," "suspected," "acceptable," "unknown," or "to be investigated," the value is further decreased by 0.1 based on the above correction results; when the original evidence fragment contains negative words such as "not seen," "denied," "no obvious," or "no recurrence," the calibration unit 103 redetermines the risk direction of the label according to the semantic direction pointed to by the negative expression. A corrected value below 0 is taken as 0, and a value above 1 is taken as 1.

[0044] To illustrate how the confidence correction value is generated, an example can be given. If parsing unit 102 extracts treatment adherence as "regular medication" from a medical record text stating "the patient regularly takes febuxostat and has not missed a dose in the past two months," the original text evidence is clear, and the initial confidence level is high. Calibration unit 103 reads prescription records from the same time period and finds that febuxostat was continuously prescribed, and that the serum uric acid level decreased from 560 μmol / L to 380 μmol / L at the next follow-up visit, while subsequent medical records do not mention any discontinuation of medication, then all evidence directions support "regular medication." Calibration unit 103 can maintain the confidence correction value of this label at a high value and form a calibration semantic label. Conversely, if the medical record only states "the patient reported that medication was adequate," and the parsing unit 102 extracts treatment adherence as "basic pattern," but the prescription record shows that the uric acid-lowering medication has not been renewed for more than two months, and adjacent records show "relapse after self-discontinuation of medication," with a higher blood uric acid level than before, then this semantic label has a directional conflict with the medication record, adjacent medical records, and laboratory indicators. In this case, the calibration unit 103 can lower the confidence correction value and mark the source of the conflict as "missing medication record for renewal, adjacent medical records indicating self-discontinuation of medication, and elevated blood uric acid." This processing does not delete the original semantic label, but rather passes it as a low-confidence label to the stratification unit 104, so that subsequent risk stratification will not overly rely on this textual statement.

[0045] A similar calibration method can be used for joint flare-up status. For example, if the current text contains "no obvious joint swelling or pain recently," parsing unit 102 generates a "no recent flare-up" label. However, if the diagnosis code for the same visit node is acute gouty arthritis, the prescription record includes a short course of colchicine and nonsteroidal anti-inflammatory drugs (NSAIDs), and the physical examination record also describes "mild redness and swelling of the right ankle," calibration unit 103 will identify a directional conflict between this label and the diagnosis code, medication record, and physical examination record, and reduce the confidence of the "no recent flare-up" label based on the source of the conflict. If the original evidence fragment is only "no obvious joint swelling or pain," where "obvious" itself has degree uncertainty, calibration unit 103 can further reduce the weight of this label. Conversely, if the current text, diagnosis code, medication record, and adjacent records do not indicate an acute flare-up, and no new anti-inflammatory analgesics are added, the "no recent flare-up" label can maintain a high confidence level.

[0046] For descriptions of renal function risk, calibration unit 103 should pay particular attention to the directional consistency between the text description and the actual test values. For example, if parsing unit 102 extracts the label "stable renal function" from the medical record text, but laboratory indicators show that creatinine increased from 85 μmol / L to 130 μmol / L, the estimated glomerular filtration rate decreased from 80 to 55, and an adjacent medical record contains a doctor's order to "pay attention to changes in renal function," then calibration unit 103 should determine that there is a directional conflict between the label and the laboratory indicators and the adjacent medical record, reduce its confidence correction value, and record the source of the conflict. If the text clearly states "creatinine has increased compared to before, consider increased risk of renal function," and the laboratory indicators also show the same change, then the renal function risk label can obtain a higher confidence level.

[0047] The calibration semantic tag set formed by calibration unit 103 should add a calibration result field to the original structured semantic tag set. Each calibration semantic tag may include the patient's anonymity number, visit node number, semantic category, original tag value, original evidence fragment, visit time, initial confidence level, confidence level correction value, presence of directional conflict, source of conflict, and the calibrated tag status. The calibrated tag status may include credible, partially credible, low confidence, or pending review, etc., which are used by hierarchical unit 104 to determine the weight of the tag in the patient's temporal multimodal feature vector. The source of conflict needs to specify which type of evidence it comes from, for example, it can be marked as "conflict with laboratory indicators", "conflict with medication records", "conflict with adjacent visit records", or a combination of multiple sources. In this way, feedback unit 105 can subsequently count the frequency of conflicts in different semantic categories and use it for updating the dynamic prompt template.

[0048] Through the above processing, calibration unit 103 achieves a reliable conversion between unstructured text extraction results and computable clinical semantic input.

[0049] Furthermore, the calibration unit is specifically used for: Based on the semantic categories and visit times in the structured semantic tag set, laboratory indicators, diagnostic codes, and medication records for the same visit node and its adjacent visit nodes are extracted from the patient's time-series medical records. Candidate validation evidence sets are generated for each semantic tag. Gout risk direction determination rules are configured for the candidate validation evidence sets according to the semantic categories. Specifically, for the treatment adherence tag, the continuity of uric acid-lowering drug prescriptions, discontinuation descriptions, and changes in serum uric acid are used to determine the support, opposition, or irrelevant status of each candidate validation evidence. For the joint attack status tag, the acute gout diagnostic code, short-term anti-inflammatory analgesic drug prescriptions, and pain descriptions are used to determine the support, opposition, or irrelevant status of each candidate validation evidence. For the renal function risk description tag, changes in creatinine are used to determine the support, opposition, or irrelevant status of each candidate validation evidence. The process involves estimating changes in glomerular filtration rate and related diagnoses of renal function to determine the supporting, contradictory, or irrelevant status of each candidate validation evidence, thus forming an evidence direction record. This record is then compared for consistency with the original evidence fragments corresponding to the semantic tags. When the semantic direction of the original evidence fragment is opposite to that of at least two candidate validation evidence fragments, a conflict source record is generated, containing the conflicting data source, conflicting consultation node, and conflicting semantic category. Based on the conflict source record, the number of supporting evidence fragments, the number of contradictory evidence fragments, and whether the original evidence fragment contains negative or uncertain words, a confidence correction value for the corresponding semantic tag is generated. Finally, the confidence correction value and the conflict source record are written into the corresponding structured semantic tags to form a calibration semantic tag set for weighted encoding by the hierarchical unit.

[0050] The calibration unit is used to reconfirm the structured semantic labels output by the parsing unit, ensuring that the semantic information extracted from the medical record text can be corroborated with the patient's test results, diagnostic information, medication history, and changes in disease progression before entering subsequent hierarchical processing. The candidate verification evidence set referred to here is a group of diagnostic and treatment evidence extracted from the patient's time-series medical data set around a specific semantic label, which can be used to determine the credibility of that semantic label. Each semantic label has a semantic category and a consultation time. The calibration unit extracts laboratory indicators, diagnostic codes, and medication records from the same consultation node and adjacent consultation nodes before and after that time, centered on the consultation node corresponding to that consultation time. For example, if a semantic label is "poor treatment adherence," and the consultation time is March 10th of a certain year, the calibration unit can extract the blood uric acid, creatinine, diagnostic code, and uric acid-lowering drug prescription records for that consultation on March 10th. It can also extract information such as whether medication was discontinued, missed, not renewed, or blood uric acid increased during the previous and next follow-up visits, thereby forming a candidate verification evidence set corresponding to that label.

[0051] The gout risk direction determination rule refers to the rule used to determine whether candidate validation evidence supports, contradicts, or is irrelevant to a certain semantic label. A supporting state indicates that the candidate validation evidence is consistent with the clinical direction expressed by the semantic label; a contradictory state indicates that the candidate validation evidence is contrary to the clinical direction expressed by the semantic label; and an irrelevant state indicates that the candidate validation evidence cannot effectively prove or refute the semantic label. For example, for the treatment adherence label, if the semantic label is "regularly taking uric acid-lowering drugs," and the medication record shows continuous renewal of febuxostat or allopurinol prescriptions, no mention of discontinuation in adjacent visit records, and a decrease in serum uric acid, then this evidence can be determined as supporting. If the medication record shows a long period without renewal, adjacent visit records show "self-discontinuation of medication" or "irregular medication," and an increase in serum uric acid, then this evidence can be determined as contradictory. If a certain laboratory indicator is not directly related to medication continuity, it can be determined as irrelevant.

[0052] For the joint attack status label, the calibration unit combines the acute gout diagnostic code, short-term prescriptions for anti-inflammatory analgesics, and pain descriptions to determine the direction of evidence. For example, if the semantic label is "no recent attack," but there is an acute gouty arthritis diagnostic code at the same visit point, and the prescription includes colchicine, nonsteroidal anti-inflammatory drugs, or short-term glucocorticoids, while adjacent medical record texts describe "joint redness, swelling, and pain," then these candidate validation pieces of evidence are contrary to the semantic direction of "no recent attack." Conversely, if the diagnostic code does not indicate an acute attack, the medication record does not include any new anti-inflammatory analgesics, and adjacent medical record records all state "no joint redness, swelling, or pain" or "no recurrence," then this evidence supports the "no recent attack" label.

[0053] For the renal function risk description label, the calibration unit combines changes in creatinine, estimated changes in glomerular filtration rate (GFR), and renal function-related diagnoses to determine the direction of evidence. For example, if the semantic label is "increased risk of renal function," and creatinine has increased since the previous visit, the estimated GFR has decreased, and the diagnosis code or medical record contains information such as chronic kidney disease or renal function abnormalities, then the candidate validation evidence is in a supporting state. If creatinine is stable or decreasing, the estimated GFR is stable or increasing, and there is no diagnosis related to renal function abnormalities, then the candidate validation evidence is contrary to or at least does not support the label. The above judgment results are compiled into an evidence direction record, which includes at least the semantic label number, the source of the candidate validation evidence, the visit node where the candidate validation evidence is located, the content of the candidate validation evidence, and its supporting, contrary, or irrelevant status.

[0054] The calibration unit further compares the evidence direction record with the original evidence fragments carried by the semantic tags. The semantic direction of the original evidence fragments can be determined based on affirmative, negative, and uncertain expressions within the fragment. For example, "patient stopped taking medication on their own for two weeks" corresponds to poor treatment adherence, "taking medication regularly and without missing doses" corresponds to good treatment adherence, and "possibly not taking medication regularly" belongs to an uncertain direction. If the semantic direction of the original evidence fragment is opposite to the direction of at least two candidate verification evidences, the calibration unit generates a conflict source record. This record specifies which data source the conflict originates from, which visit node it corresponds to, and which semantic category it belongs to. For example, "The treatment adherence tag conflicts with the medication record and adjacent visit records; the conflict nodes are the current node and the next follow-up visit node; the conflict semantic category is treatment adherence."

[0055] Subsequently, the calibration unit generates a confidence correction value based on the conflict source record, the number of supporting evidence, the number of contradictory evidence, and whether the original evidence fragment contains negative or uncertain words. The confidence correction value indicates the credibility of the semantic tag in subsequent hierarchical coding and can be set to high, medium, or low levels, or a value between zero and one. For example, if a tag has three supporting pieces of evidence, no contradictory evidence, and the original evidence fragment is clear, the confidence correction value can be set to high; if a tag has one supporting piece of evidence, two contradictory pieces of evidence, and the original evidence fragment contains uncertain words such as "possible," "acceptable," or "unknown," the confidence correction value should be lowered; if the original evidence fragment is clear but there are more than two contradictory pieces of evidence among the candidate verification evidence, the tag is not deleted but marked as low credibility and the conflict source is retained so that subsequent hierarchical units can reduce its coding weight.

[0056] After completing the above processing, the calibration unit writes the confidence correction value and conflict source record into the corresponding structured semantic label to form a calibration semantic label set. This calibration semantic label set retains the original semantic category, label value, original evidence fragment, and consultation time, while adding evidence direction record, conflict source record, and confidence correction value, enabling the hierarchical unit to perform weighted coding based on the label confidence level.

[0057] The hierarchical unit 104 is used to align the patient time-series diagnosis and treatment dataset, calibration semantic label set, laboratory indicators, diagnostic codes and medication records along the same time axis, and perform weighted encoding based on confidence correction values ​​to generate patient time-series multimodal feature vectors, thereby obtaining risk stratification results of low risk, medium risk or high risk.

[0058] The stratification unit 104, after the construction unit 101 forms the patient time-series diagnosis and treatment dataset, the parsing unit 102 generates the structured semantic label set, and the calibration unit 103 forms the calibration semantic label set, transforms the diagnosis and treatment information of the same patient from different sources, at different times, and in different data formats into a unified multimodal feature expression, and obtains low-risk, medium-risk, or high-risk risk stratification results accordingly. Here, "multimodal" means that the data participating in risk stratification is not a single type of data, but simultaneously includes numerical laboratory indicators, categorical diagnostic codes, time-based medication records, semantic labels obtained from parsing medical record text, and the patient's medical record sequence formed over time. The "patient time-series multimodal feature vector" refers to a set of feature values ​​that can be recognized and calculated by the risk stratification algorithm after the above-mentioned different types of data are organized according to a unified time axis. These feature values ​​can be stored in either a tabular vector form or a sequence vector form arranged by the visit node.

[0059] Hierarchical unit 104 first aligns the patient's time-series medical records, calibration semantic label set, laboratory indicators, diagnostic codes, and medication records along the same time axis. This alignment means using the consultation nodes already established by construction unit 101 as a benchmark, assigning each type of data to its corresponding consultation node or a time window with a clear temporal relationship to that node, thus preventing data generated at different times from being incorrectly treated as the same condition. For example, if a patient describes joint pain in their outpatient record on March 1st, completes a blood uric acid test on March 2nd, and receives a colchicine prescription on March 1st, and if the preset outpatient laboratory merging window is three days before and after the consultation date, then the aforementioned text, blood uric acid test results, and prescription records can be aligned to the March 1st consultation node. If the patient returns for a follow-up visit on June 1st with new blood uric acid results and a new medication description, then this data is aligned to the June 1st consultation node, without being mixed with the March 1st data. Through this alignment, hierarchical unit 104 can reflect the changes in gout patients' condition, medication, and behavioral risks over time.

[0060] After timeline alignment, the stratification unit 104 encodes different types of data. For laboratory indicators, the stratification unit 104 can retain their standardized values ​​or convert them into status characteristics such as normal, high, significantly high, decreasing, and increasing based on medical thresholds. For example, serum uric acid can be simultaneously encoded as the current value, whether it has decreased since the previous visit, and whether it has reached the preset control target; creatinine and estimated glomerular filtration rate can be encoded as the current renal function status, whether it has worsened since then, and whether there are persistent abnormalities. For diagnostic coding, the stratification unit 104 can use a category coding method to convert diagnostic information such as gout, hyperuricemia, acute gouty arthritis, tophi, and chronic kidney disease into corresponding features. For medication records, the stratification unit 104 can encode the drug category, whether uric acid-lowering treatment is present, whether new anti-inflammatory and analgesic treatment has been added, whether the prescriptions are consecutive, and whether the prescription time covers the current visit. The above coding methods are all data coding methods that can be implemented by those skilled in the art. The key point of this invention is that these codes are all organized around the same patient's timeline and enter the risk stratification process together with the calibrated semantic tags.

[0061] For the calibration semantic label set, the hierarchical unit 104 does not simply use the labels as presence or absence features, but rather performs weighted encoding based on the confidence correction values ​​generated by the calibration unit 103. Specifically, the hierarchical unit 104 first encodes the risk direction of the semantic labels into three states: increased risk, decreased risk, and neutral risk. Increased risk is encoded as 1, decreased risk as -1, and neutral or not mentioned as 0. Then, this risk direction encoding value is multiplied by the corresponding semantic label's confidence correction value to obtain the semantic risk encoding value. If a semantic label has conflict source records, and the number of contrary evidence is not less than the number of supporting evidence, the hierarchical unit 104 generates a conflict suppression marker and multiplies the semantic risk encoding value by 0.2 before writing it into the node feature vector. Therefore, semantic labels with high confidence and not subject to conflict suppression have a greater effect on risk hierarchical development, while semantic labels with low confidence or subject to conflict suppression have a reduced effect on risk hierarchical development. For example, parsing unit 102 extracts the "regular medication" tag from the medical record text, but calibration unit 103 finds that both the medication record and adjacent medical records support this tag and sets its confidence correction value to a high value. In this case, stratification unit 104 assigns a high positive weight to the treatment adherence-related features. Conversely, if the text only vaguely states "medication is acceptable," but the prescription record shows that the prescription has not been renewed for a long time, and adjacent records indicate that the medication was discontinued on its own, calibration unit 103 lowers the confidence correction value of this tag. In this case, stratification unit 104 treats it only as low-confidence information during encoding, and may even retain the "inconsistency risk" feature at the same time, so that risk stratification is not misled by a single textual statement.

[0062] To facilitate understanding, an example can be used to illustrate the specific implementation of weighted encoding. Suppose that at a certain patient visit, the calibrated semantic labels include "alcohol-related triggers" being positive with a high confidence correction value, "regular medication use" being positive but with a low confidence correction value, and "no recent flare-ups" being positive with a moderate confidence correction value. The stratification unit 104 can encode "alcohol-related triggers" as a clearly present lifestyle risk feature, "regular medication use" as a low-confidence adherence improvement feature, and "no recent flare-ups" as a moderately confident disease activity reduction feature. If, at the same visit, serum uric acid remains significantly elevated, and the prescription record shows discontinuous uric acid-lowering treatment, the final patient temporal multimodal feature vector will simultaneously reflect information such as "lifestyle risk exists," "poor uric acid control," "insufficient evidence of adherence improvement," and "recent symptom relief may be possible but with limited confidence." Thus, risk stratification is based on the comprehensive state of multi-source evidence, rather than a single indicator or a single text label.

[0063] The patient's temporal multimodal feature vector can be organized according to the patient dimension and the visit node dimension. For each visit node, the hierarchical unit 104 can generate corresponding node features, including laboratory indicator features, diagnostic features, medication features, calibration semantic label features, adjacent node change features, and data conflict features. For the entire patient, the hierarchical unit 104 can arrange multiple node features in chronological order to form a patient pathology sequence. If the system uses traditional machine learning methods, the features of the most recent visit node, the trend features of the recent visits, and the historical maximum risk feature can be merged into a fixed-length vector; if the system uses a sequence model, the node vector arranged in chronological order can be directly input. Regardless of the specific algorithm used, it should be ensured that the input features come from multimodal data aligned to the same time axis, and the influence of confidence correction values ​​on semantic labels should be preserved.

[0064] In this embodiment, the stratification unit 104 employs a deterministic classification process based on fixed field encoding and three types of risk scoring rules to achieve risk stratification. This deterministic classification process does not rely on statistical models or machine learning models with undisclosed training processes. Instead, it takes the patient's temporal multimodal feature vector as input, calculates the recurrence risk score, the tophi or chronic progression risk score, and the renal function complication risk score, respectively, and outputs the risk stratification result based on the comparison results of the three types of risk scores with the corresponding score thresholds.

[0065] Specifically, the hierarchical unit 104 first uses the patient's time-series medical records as the basic processing object, generating a node feature vector for each visit node. The node feature vector, arranged in a fixed field order, includes: serum uric acid target status, direction of uric acid change, direction of creatinine change, estimated glomerular filtration rate change, continuity of uric acid-lowering drug prescriptions, short-term addition of anti-inflammatory drugs, acute gout diagnosis status, treatment adherence semantic features, lifestyle trigger semantic features, joint attack status semantic features, tophi semantic features, renal function risk semantic features, number of conflict sources, and number of conflict suppression markers. The encoding method for each field is pre-fixed in the system to ensure that feature vectors are generated according to the same rules across different visit nodes for the same patient and between different patients.

[0066] The status of serum uric acid levels is determined based on the serum uric acid test value at the current consultation point. For patients without recorded tophi, a serum uric acid level less than 360 μmol / L is coded as 0, indicating that the target has been met; a level greater than or equal to 360 μmol / L is coded as 1, indicating that the target has not been met. For patients with a description of tophi or a tophi-related diagnosis, a serum uric acid level less than 300 μmol / L is coded as 0, and a level greater than or equal to 300 μmol / L is coded as 1. If there is no serum uric acid test value at the current consultation point, this field is coded as 0, and serum uric acid deficiency is recorded in the missing marker to avoid misclassifying a lack of testing as meeting the target. The direction of change in serum uric acid is determined by comparing the serum uric acid test value at the current consultation point with that at the previous consultation point. When the current value increases by more than 30 μmol / L compared with the previous value, it is coded as 1; when it decreases by more than 30 μmol / L, it is coded as -1; when the change does not exceed 30 μmol / L, it is coded as 0. If the serum uric acid value at the previous point or any point is missing, it is coded as 0 and recorded as a missing trend.

[0067] The direction of change in creatinine and the direction of change in estimated glomerular filtration rate (GFR) are used to indicate changes in renal function. If creatinine increases by at least 26.5 μmol / L or 30% compared to the previous visit, the direction of change is coded as 1; otherwise, it is coded as 0. If no comparable creatinine value is available, it is coded as 0 and recorded as missing. If the estimated GFR decreases by at least 15% compared to the previous visit, the direction of change in estimated GFR is coded as 1; otherwise, it is coded as 0. If no comparable estimated GFR is available, it is coded as 0 and recorded as missing. The continuity of uric acid-lowering medication prescriptions is determined based on the start and end dates of prescriptions for uric acid-lowering drugs such as febuxostat, allopurinol, and benzbromarone. An interval of no more than 30 days between two consecutive prescriptions is coded as 1; an interval of more than 30 days is coded as 0. If the patient has never been prescribed uric acid-lowering medication, it is coded as 0. The status of newly added anti-inflammatory drugs is determined based on whether colchicine, nonsteroidal anti-inflammatory drugs, or short-term glucocorticoids have been added within 14 days before or after the current visit. A new addition is coded as 1, and no new addition is coded as 0. The status of acute gout diagnosis is determined based on whether acute gouty arthritis, acute gout attack, or a synonymous diagnosis is coded at the current visit. A coded 1 is coded if present, and 0 is coded if absent.

[0068] For the five semantic features—treatment adherence, lifestyle triggers, joint attack status, description of tophi, and description of renal function risk—the hierarchical unit 104 first performs basic encoding based on the risk direction of the corresponding semantic label in the calibration semantic label set. If the semantic label indicates an increased risk, such as self-discontinuation of medication, intermittent medication, alcohol-induced attack, acute attack, presence of tophi, or elevated creatinine, the risk direction is encoded as 1. If the semantic label indicates a decreased risk, such as regular medication, denial of alcohol consumption, no recent attacks, no tophi, or stable renal function, the risk direction is encoded as -1. If the semantic label is not mentioned, uncertain, or not directly related to the current risk, the risk direction is encoded as 0. Subsequently, the hierarchical unit 104 multiplies the risk direction encoding value by the confidence correction value of the semantic label to obtain the corresponding semantic risk encoding value. For example, the risk direction code for "discontinuing medication on one's own" is 1, and the confidence correction value is 0.8, so the semantic feature of treatment adherence is 0.8; the risk direction code for "regular medication" is -1, and the confidence correction value is 0.9, so the semantic feature of treatment adherence is -0.9; the risk direction code for "lifestyle triggers not mentioned" is 0, so the semantic feature of lifestyle triggers is 0.

[0069] If multiple semantic labels exist for the same semantic category at the same medical visit node, the hierarchical unit 104 prioritizes the semantic risk encoding value with the largest absolute value as the node semantic feature for that semantic category. When there is a risk-increasing label and a risk-reducing label with the same absolute value, the label with the conflict source record is prioritized, and the number of conflict sources for that semantic category is increased by one. If a semantic label has a conflict source record, and the number of contrary evidence is not less than the number of supporting evidence, the hierarchical unit 104 generates a conflict suppression label for that semantic label and multiplies the semantic risk encoding value by 0.2 before writing it into the node feature vector. For example, the parsing unit extracts "regular medication," whose risk direction encoding is -1 and confidence correction value is 0.8. However, the calibration unit confirms the existence of a record of not renewing prescriptions and a record of "self-discontinuation of medication" in the adjacent node, and the number of contrary evidence is not less than the number of supporting evidence. In this case, the hierarchical unit 104 generates a conflict suppression label and adjusts the semantic feature from -0.8 to -0.16. Thus, text semantics with obvious conflicts will not directly have an excessive impact on the risk level.

[0070] After generating the node feature vector for each medical visit node, the hierarchical unit 104 selects the node feature vectors of the three most recent medical visit nodes and concatenates them in chronological order to form a patient temporal multimodal feature vector. If a patient has only one or two medical visit nodes, the fields of the missing nodes are filled with 0, and the number of missing nodes is recorded. If a patient has more than three medical visit nodes, the three medical visit nodes closest to the current risk stratification time are selected first. Thus, the patient temporal multimodal feature vector includes not only the current medical visit status but also changes in blood uric acid, medication continuity, acute exacerbation status, and semantic risk between adjacent medical visit nodes, enabling the hierarchical unit 104 to make risk assessments based on the continuous course of the disease.

[0071] Hierarchical unit 104 calculates recurrence risk scores, tophi or chronic progression risk scores, and renal function complication risk scores based on the patient's temporal multimodal feature vectors. The recurrence risk score is calculated according to the following rules: 2 points for serum uric acid not reaching the target level at the current node; 1 point for consecutive increases in serum uric acid in the last three nodes; 3 points for an acute gout diagnosis at the current node; 2 points for a short-term new anti-inflammatory medication status at the current node; 2 points for a joint attack semantic feature greater than 0.6; 1 point for a lifestyle trigger semantic feature greater than 0.6; and 1 point for treatment adherence semantic feature greater than 0.6. A recurrence risk score of 6 or higher is considered high risk, 3 to 5 points is considered medium risk, and below 3 points is considered low risk.

[0072] The risk score for tophi or chronic progression is calculated according to the following rules: 3 points for a semantic feature of tophi greater than 0.6; 3 points for a tophi-related diagnosis at the current node or adjacent nodes; 2 points for serum uric acid levels not reaching the target in at least two of the most recent three nodes; 2 points for a joint attack state semantic feature greater than 0.6 or an acute gout diagnosis in at least two of the most recent three nodes; 1 point for a uric acid-lowering drug prescription continuity score of 0 at least once in the most recent three nodes. A risk score of 5 or higher for tophi or chronic progression is considered high risk for chronic progression, 3 to 4 is considered medium risk, and below 3 is considered low risk.

[0073] The risk score for renal function complications is calculated according to the following rules: 2 points for a change in creatinine direction of 1; 2 points for a change in estimated glomerular filtration rate direction of 1; 3 points for the presence of chronic kidney disease, abnormal renal function, or a synonymous diagnostic code at the current node; 2 points for a renal function risk semantic feature greater than 0.6; 1 point for two consecutive renal function risk semantic features greater than 0.6 in the most recent three nodes. A renal function complication risk score of 5 or higher is considered high risk, 3 to 4 is considered medium risk, and below 3 is considered low risk.

[0074] The final risk stratification result is determined by the combined state of the three risk categories. When any one of the risks of recurrence, tophi or chronic progression, and renal complications is high, and the main semantic feature supporting this high-risk judgment is not suppressed by conflict, stratification unit 104 outputs high risk. When there is no high risk, but any one of the three risk categories is medium risk, or there is any persistent abnormality such as two consecutive nodes of serum uric acid not meeting the target, discontinuous uric acid-lowering drug prescription, persistently elevated creatinine, or persistently declining estimated glomerular filtration rate, stratification unit 104 outputs medium risk. In other cases, stratification unit 104 outputs low risk. If a high-risk judgment mainly relies on conflict-suppressed semantic features, but the corresponding structured risk feature does not meet the high-risk condition, then this high-risk judgment is not directly used as the final high-risk basis. Stratification unit 104 downgrades it to a medium-risk candidate and provides the conflict-suppressed semantic category, conflict source, and corresponding medical visit node to feedback unit 105.

[0075] For example, a patient's serum uric acid levels at their three most recent medical visits were 520 μmol / L, 540 μmol / L, and 560 μmol / L, respectively, all below the target level and continuously rising. At the current visit, the patient was diagnosed with acute gouty arthritis, and a colchicine prescription was added within 14 days before and after this visit. The parsing unit extracted two semantic tags: "redness, swelling, and pain in the left foot after drinking alcohol" and "discontinuation of medication on one's own." The confidence correction values ​​generated by the calibration unit were 0.9 and 0.8, respectively, and neither generated a conflict suppression marker. When the stratification unit 104 calculated the recurrence risk score, 2 points were awarded for serum uric acid below the target level, 1 point for continuously rising serum uric acid, 3 points for an acute gout diagnosis, 2 points for a short-term addition of anti-inflammatory drugs, 2 points for a joint attack semantic feature greater than 0.6, 1 point for a lifestyle trigger semantic feature greater than 0.6, and 1 point for treatment adherence semantic feature greater than 0.6. The recurrence risk score was 12 points, therefore, the patient was determined to be at high risk of recurrence, and the final high risk was output.

[0076] For example, a patient's serum uric acid levels at their two most recent medical visits were 390 μmol / L and 400 μmol / L, respectively, both below target but not significantly elevated. There was no diagnosis of acute gout, no short-term addition of anti-inflammatory medication, and the uric acid-lowering medication was continuously administered. The medical record only mentions "occasional discomfort." After calibration, the semantic feature of joint attack status is 0.3. When calculating the recurrence risk score in stratified unit 104, only 2 points are awarded because the current serum uric acid level is below target, failing to reach the recurrence risk threshold. However, because the serum uric acid level is below target for two consecutive visits, constituting a persistent abnormal state, the final output is medium risk. For another example, a patient's current serum uric acid level is within target range, the uric acid-lowering medication is continuously administered, there is no diagnosis of acute gout, no short-term addition of anti-inflammatory medication, and the treatment adherence semantic feature is -0.9. The risks of lifestyle triggers, tophi, and renal function are all 0 or decreasing. Therefore, the risks of recurrence, chronic progression, and renal complications are all low risk, resulting in a low risk output.

[0077] In this embodiment, the aforementioned scoring items, scores, and thresholds are stored in a rule configuration table as system preset rules. The rule configuration table includes at least a rule number, applicable risk type, input field name, trigger condition, score, risk level threshold, and activation status. The values ​​in the rule configuration table can be written during system initialization or adjusted by an authorized administrator according to the gout management goals adopted by the medical institution; the adjusted rule version number is saved together with the risk stratification results. Regardless of whether it is adjusted, the stratification unit 104 performs the same data reading, field encoding, score comparison, and level output process according to the currently activated rule version.

[0078] For example, a patient's serum uric acid level decreased from 560 to 480 during their last three visits, then rose again to 530 micromoles per liter. Prescription records show discontinuation of uric acid-lowering medication. The calibrated semantic labels indicate "repeated alcohol consumption" and "self-discontinuation of medication" as highly reliable. The most recent text also mentions "recurrence of redness, swelling, and pain in the left big toe joint." After aligning the timeline, stratification unit 104 will encode poor uric acid control, discontinuous medication use, highly reliable lifestyle triggers, highly reliable adherence risk, and recent attack status, forming a high-risk feature combination and outputting a high-risk result. As another example, if a patient's serum uric acid level has continuously decreased to near the target level, the last two medical records indicate no recurrence, the prescription records are continuous, the calibrated semantic label "regular medication use" is highly reliable, and there is no evidence related to tophi or abnormal kidney function, then stratification unit 104 can output a low-risk result. If a patient's serum uric acid level has decreased but not reached the target, with occasional alcohol consumption, no recent acute attacks, but unstable adherence evidence, then a medium-risk result can be output.

[0079] The risk stratification results output by the stratification unit 104 not only include low-risk, medium-risk, or high-risk categories, but can also simultaneously output the patient temporal multimodal feature vector or its index used to generate the results, so that the feedback unit 105 can further generate feature contribution explanation information and risk evolution records. The stratification unit 104 can also save the intermediate features corresponding to each visit node, such as the current blood uric acid status, renal function status, medication continuity status, semantic label confidence, and conflict markers, so that it can trace whether a change in risk level was caused by changes in laboratory indicators, medication, text semantics, or calibration confidence.

[0080] Furthermore, the hierarchical unit is specifically used for: establishing a patient node sequence according to the visit time in the patient time-series medical data set; assigning the calibration semantic label set, laboratory indicators, diagnostic codes, and medication records to the corresponding visit nodes to generate node-level aligned data; extracting from the node-level aligned data the status of serum uric acid reaching the target level, the direction of continuous change in serum uric acid, the direction of change in creatinine and estimated glomerular filtration rate, the continuity of uric acid-lowering drug prescriptions, the status of short-term addition of anti-inflammatory drugs, and the status of acute gout diagnosis to generate structured risk feature groups; converting treatment adherence, lifestyle triggers, joint attack status, description of tophi, and description of renal function risk into semantic risk features, and correcting them according to the corresponding confidence levels. The encoding weights are adjusted; for semantic risk features with conflicting sources and where the number of contrary evidences is not less than the number of supporting evidences, conflict suppression markers are generated; based on the structured risk feature group, semantic risk features, conflict suppression markers, and changes in blood uric acid, medication continuity, and attack status between adjacent visit nodes, a patient temporal multimodal feature vector is generated; based on the patient temporal multimodal feature vector, the risks of recurrence, tophi or chronic progression, and renal complications are determined respectively, and a high risk is output when any risk reaches the high risk condition and the corresponding semantic risk feature is not suppressed by conflict, a medium risk is output when there is only a single persistent abnormality that does not reach the high risk condition, and a low risk is output otherwise.

[0081] The hierarchical unit, after the calibration unit has formed a calibration semantic label set, transforms the structured medical information and calibrated textual semantic information of the same patient at different time points into feature expressions that can be directly used for risk assessment, and outputs low-risk, medium-risk, or high-risk hierarchical results. The patient node sequence referred to here is the result of arranging the medical data generated from each outpatient visit, hospitalization, follow-up visit, or revisit of the same patient into consecutive nodes, based on the patient's consultation time. Each consultation node corresponds to a specific time or time interval and is used to accommodate laboratory indicators, diagnostic codes, medication records, and calibration semantic labels related to that consultation. When generating node-level aligned data, the hierarchical unit uses the consultation time in the patient's time-series medical data set as a benchmark, grouping data generated within the same consultation time or a preset time window into the same node. For example, if a patient consults on March 1st, completes a blood uric acid test on March 2nd, and obtains a prescription for uric acid-lowering drugs on March 1st, then the above data can be grouped into the consultation node corresponding to March 1st; if a follow-up visit is conducted on April 10th and new test results and medical record text are generated, then the next consultation node is formed.

[0082] Within each node-level aligned data set, the hierarchical unit further extracts structured risk feature groups. These structured risk feature groups are a set of features derived from fields in laboratory indicators, diagnostic codes, and medication records that directly reflect gout risk status. The target uric acid level can be determined based on the medical institution's pre-set control goals; for example, reaching the target range is marked as achieving the target, and exceeding the target range is marked as not achieving the target. The direction of continuous change in uric acid refers to whether the current node's uric acid level is increasing, decreasing, or relatively stable relative to the previous node. For example, if the previous node's uric acid level was 520 micromoles per liter, and the current node's is 420 micromoles per liter, it is marked as decreasing. The direction of change in creatinine and estimated glomerular filtration rate (GFR) reflects changes in renal function; an increase in creatinine or a decrease in estimated GFR usually indicates an increased risk of renal dysfunction. The continuity of uric acid-lowering drug prescriptions indicates whether the patient continues to receive prescriptions for uric acid-lowering drugs such as febuxostat and allopurinol between adjacent nodes; the short-term addition status of anti-inflammatory drugs indicates whether the current node has added colchicine, nonsteroidal anti-inflammatory drugs, or short-term glucocorticoids for the treatment of acute attacks; the acute gout diagnosis status indicates whether the current node has a diagnostic code such as acute gouty arthritis. These features collectively form a structured risk feature set, which is used to subsequently generate patient temporal multimodal feature vectors.

[0083] For calibrating the semantic label set, the hierarchical unit converts treatment adherence, lifestyle triggers, joint attack status, tophi descriptions, and renal function risk descriptions into semantic risk features. Semantic risk features, as referred to here, are characteristics obtained from parsed and calibrated medical record text that reflect the patient's behavioral risk, symptom status, or complication risk. For example, treatment adherence labels can be converted to features such as regular medication, intermittent medication, self-discontinuation, or uncertain; lifestyle triggers can be converted to features such as alcohol consumption, high-purine diet, dehydration, or not mentioned; joint attack status can be converted to features such as acute attack, symptom relief, recent absence, or uncertain; tophi descriptions can be converted to presence, suspected, absent, or unclear progression; and renal function risk descriptions can be converted to increased risk, stable risk, not mentioned, or uncertain. The hierarchical unit does not treat all semantic labels as equally reliable inputs, but adjusts their encoding weights according to the confidence correction values ​​generated by the calibration unit. If a semantic label has a high confidence level, its corresponding semantic risk feature has a higher role in subsequent risk assessment; if the confidence level is low, its role is reduced.

[0084] The conflict suppression marker is used to prevent semantic risk features with obvious conflicts from directly affecting the risk level. Specifically, when a semantic risk feature has conflict source records, and the number of contrary evidence identified during the calibration process is not less than the number of supporting evidence, the stratification unit generates a conflict suppression marker for that semantic risk feature. The number of supporting evidence and the number of contrary evidence here come from the calibration unit's judgment result on candidate verification evidence. For example, if the medical record text states "the patient claims to take medication regularly," but the medication record shows that the prescription was not renewed for two months, and the adjacent medical record records "relapse after self-discontinuation of medication," and blood uric acid is also higher than before, then the semantic risk feature "regular medication" has multiple contrary pieces of evidence. The stratification unit generates a conflict suppression marker for it, so that this semantic risk feature is no longer a valid basis for reducing the risk level. Conversely, if the text records "self-discontinuation of medication," the medication record also shows that the prescription was not renewed, and blood uric acid is elevated, then this semantic risk feature is consistent with other evidence, does not generate a conflict suppression marker, and can be used as an important basis for improving risk assessment.

[0085] The hierarchical unit then generates a patient temporal multimodal feature vector based on structured risk feature groups, semantic risk features, conflict suppression markers, and changes between adjacent visit nodes. This feature vector can be understood as a set of risk states arranged in chronological order. It includes structured features such as whether the current node's uric acid level is within target range, whether renal function has worsened, whether an acute gout diagnosis has been made, and whether new anti-inflammatory drugs have been added. It also includes semantic features such as treatment adherence, lifestyle triggers, and joint attack status, as well as the changing trends between adjacent nodes. For example, a decrease in uric acid level indicates improved control, while a continuous increase in uric acid level indicates worsening control; a change from continuous medication to interrupted medication indicates increased adherence risk; and a change from no attack to acute attack indicates increased recurrence risk. By incorporating this information into the patient temporal multimodal feature vector, the hierarchical unit can express the patient's current status and disease progression, rather than just the static result of a single visit.

[0086] Based on the patient's temporal multimodal feature vector, stratification units are used to determine the risks of recurrence, tophi or chronic progression, and renal complications. Recurrence risk is primarily determined based on recent joint attack status, short-term new use of anti-inflammatory drugs, acute gout diagnosis, uncontrolled uric acid levels, lifestyle triggers, and treatment adherence. Tophi or chronic progression risk is primarily determined based on tophi description, long-term uncontrolled uric acid levels, recurrent attack history, and persistently abnormal disease progression. Renal complications risk is primarily determined based on the direction of creatinine changes, the estimated direction of glomerular filtration rate changes, renal function-related diagnoses, and renal function risk descriptions. These three risk categories can be internally assessed as high, moderate, low, normal, abnormal, or severely abnormal, which are then used for final risk stratification.

[0087] In the final output, if any of the risks—recurrence risk, tophi or chronic progression risk, or renal complication risk—meets the high-risk criteria, and the corresponding semantic risk feature supporting this high-risk judgment is not suppressed by conflict, then the hierarchical unit outputs high risk. For example, if a patient's serum uric acid levels have consistently remained below target, they have recently started taking anti-inflammatory medication, the diagnostic code suggests acute gouty arthritis, and the high-confidence semantic label shows "recurrence after alcohol consumption," and neither the lifestyle trigger nor the attack state is suppressed by conflict, then high risk can be output. If only a single persistent abnormality does not meet the high-risk criteria, such as consistently elevated serum uric acid without recent acute attacks, tophi progression, or significant renal function deterioration, then medium risk is output. In other cases, such as serum uric acid reaching or continuously improving, no recent attacks, continuous medication use, and no clear tophi progression or renal function risk, low risk is output.

[0088] For example, a patient's serum uric acid levels at three consecutive nodes are 560, 520, and 540 micromoles per liter, all below the target range. The most recent node shows a new colchicine prescription and a diagnosis of acute gouty arthritis. The calibration semantic label indicates "redness, swelling, and pain in the left foot joint after drinking alcohol," with high confidence and no conflict suppression marker. The hierarchical unit generates a patient temporal multimodal feature vector that simultaneously includes information such as persistently abnormal serum uric acid, elevated attack status, clear lifestyle triggers, and new anti-inflammatory drugs. The recurrence risk meets the high-risk criteria, therefore, a high-risk output is given. As another example, a patient's serum uric acid drops from 500 to 420 but still falls below the target. Medication records are continuous, there have been no recent acute attacks, and the text occasionally mentions alcohol consumption with low confidence. The system can then classify this as a single persistent abnormality or mild behavioral risk, outputting a medium-risk output. If the patient's serum uric acid consistently reaches the target, there is no diagnosis of acute gout, no new anti-inflammatory drugs, and the calibration semantic label shows regular medication use without conflict suppression, then a low-risk output is given.

[0089] Feedback unit 105 is used to generate feature contribution explanation information and risk evolution record aggregated by semantic category based on risk stratification results and patient temporal multimodal feature vectors, and feed back conflict source and feature contribution explanation information to parsing unit, and update dynamic prompt template extraction constraints for semantic category where both contribution and conflict frequency exceed preset thresholds.

[0090] Feedback unit 105 processes the reasons for the formation of the risk stratification results, their changes over time, and the categories prone to conflict during semantic parsing after the risk stratification results are obtained by the stratification unit 104. The processing results are then fed back to the parsing unit 102, enabling the parsing unit 102 to generate more explicit extraction constraints when subsequently generating disease-specific constraint prompts. Feedback unit 105 does not simply display the risk level; instead, it combines the risk stratification results, the patient's temporal multimodal feature vector, the sources of conflict identified by the calibration unit 103, and the confidence correction of each semantic label to form feedback information that can be understood by doctors, recorded by the system, and used to update the dynamic prompt template.

[0091] The feature contribution explanation information referred to here is information used to explain which features mainly influence the risk stratification result of a patient. This contribution can be directly output by the risk stratification algorithm used by the stratification unit 104, or it can be calculated and generated by the feedback unit 105 based on the feature change direction, feature weight, confidence correction value, and risk level change. If the stratification unit 104 adopts a rule-based stratification method, the feedback unit 105 can use the features that trigger the stratification rules as the main contributing features, such as persistently high uric acid levels above the target range, recent acute joint attacks, discontinuous uric acid-lowering drug prescriptions, low confidence in treatment adherence labels, or high confidence in alcohol consumption trigger labels. If the stratification unit 104 adopts a machine learning classification model, the feedback unit 105 can read the feature importance, local explanation results, or the positive and negative impact of each feature on the classification result output by the model and convert them into an expression that doctors can understand. The "contribution" here is not required to be limited to a specific mathematical algorithm, as long as it can reflect the magnitude and direction of the influence of a feature on the current low-risk, medium-risk, or high-risk result, and can be generated according to consistent rules within the same system.

[0092] Semantic category aggregation means that feedback unit 105 does not simply display explanatory information as scattered fields, but rather summarizes the contribution information related to semantic categories such as treatment adherence, lifestyle triggers, joint attack status, description of tophi, and description of renal function risk. For example, if a patient is classified as high-risk, the contribution sources of their characteristics include the highly reliable semantic label of "frequent alcohol consumption in the past three months," the highly reliable semantic label of "self-discontinuation of febuxostat," the laboratory indicator feature of "persistently elevated blood uric acid," and the diagnostic code of "acute gouty arthritis." Feedback unit 105 can categorize "frequent alcohol consumption" into the lifestyle trigger category, "self-discontinuation of febuxostat" into the treatment adherence category, and "acute gouty arthritis" into the joint attack status category, and record the corresponding evidence, contribution direction, and contribution intensity under each category. In this way, doctors or subsequent systems can intuitively see that high risk is not generated by a single indicator, but is supported by multiple semantic categories and structured indicators.

[0093] The risk evolution record refers to the record generated by the feedback unit 105 based on the risk stratification results of different patient visit nodes and the patient's temporal multimodal feature vector, documenting the changes in the patient's risk status over time. This record may include the risk level at each visit node, key risk factors, reasons for increased or decreased risk, explanations of changes between adjacent visit nodes, and relevant evidence sources. For example, a patient might be at high risk during their first visit, primarily due to significantly elevated uric acid, an acute joint attack, and alcohol consumption; at their second visit, they might be at medium risk, primarily due to decreased uric acid and symptom relief, but still unstable adherence; and at their third visit, they might be at low risk, primarily due to uric acid levels approaching the target level, no recurrence, and continuous medication use. The feedback unit 105 can organize these changes into a continuous risk evolution record, enabling clinicians to understand the patient's risk changes and facilitating subsequent follow-up management.

[0094] Feedback unit 105 also feeds back conflict source and feature contribution explanation information to parsing unit 102. This feedback does not require manual re-entry; instead, the system writes the already generated conflict statistics and contribution results into the dynamic prompt template management data, which parsing unit 102 can then access when processing the same patient's or similar patient's medical records. Conflict sources originate from calibration unit 103, specifically indicating which type of evidence a semantic label directionally conflicts with, such as conflicting with laboratory indicators, diagnostic codes, medication records, adjacent medical records, or multiple sources simultaneously. Feature contribution explanation information indicates whether the semantic category has a high impact on the risk stratification results. Feedback unit 105 considers both conflict sources and contribution to avoid frequently adjusting the prompt template for low-impact, occasional conflict semantic categories, and also to avoid ignoring high-frequency conflict categories that have a significant impact on risk stratification.

[0095] Both the contribution level and the conflict frequency exceeding preset thresholds are conditions for feedback unit 105 to trigger a dynamic prompt template update. These preset thresholds can be set by system administrators or medical institutions based on data scale, model performance, and clinical requirements. The contribution level threshold is used to determine whether a semantic category has sufficient impact on risk stratification, and the conflict frequency threshold is used to determine whether the semantic category repeatedly exhibits inconsistencies between parsing and calibration. For example, the system can set a condition that if the average contribution level of a semantic category ranks among the top few of all semantic categories in the most recent 100 patient samples or the most recent month's processed records, and the number of directional conflicts in that category exceeds a preset number or the conflict ratio exceeds a preset ratio, then the semantic category is considered to need an update prompt constraint. For ease of implementation, simpler rules can also be set, such as triggering a template update if a semantic category is listed as a major contributing factor more than five times in twenty consecutive risk stratifications, and there are directional conflicts with medication records or laboratory indicators more than three times. The above thresholds are not fixed values; those skilled in the art can set them according to the actual data volume and hospital quality control requirements, but the triggering conditions should be clearly stored in the system and traceable.

[0096] The dynamic prompt template extraction constraint update refers to the feedback unit 105 modifying the category definition, judgment boundary, negation recognition rule, time attribution rule, or output requirements used by the parsing unit 102 when generating disease-specific constraint prompts based on the triggered semantic category and its conflict source. For example, if the treatment adherence category has multiple instances of high contribution and high conflict frequency, and the main conflict source is medication records showing no prescription renewal, but the text contains "the patient stated that medication is adequate" which is parsed as good adherence, then the feedback unit 105 can update the prompt constraint of the treatment adherence category, adding the following content: "When the text is expressed as vague expressions such as 'medication is adequate,' 'medication is average,' or 'occasionally missed doses,' it should not be directly marked as regular medication, but should be marked as 'adherence uncertain,' and the output should be a fragment of original text evidence supporting this judgment." If the conflict source mainly comes from adjacent medical records prompting self-discontinuation of medication, the constraint "When the current text does not explicitly state continuous medication, but adjacent medical records contain descriptions of discontinuation or missed doses, the model should be prompted to output adherence labels that need calibration" can be further added.

[0097] Feedback unit 105 can adopt a version-based management approach when updating dynamic prompt templates. Each update saves the template content before the update, the template content after the update, the semantic category that triggered the update, the corresponding contribution, the frequency of conflicts, the main source of conflicts, and the update time. For medical scenarios, a manual review mechanism can also be set up, whereby feedback unit 105 first generates template update suggestions, which are then confirmed by authorized doctors, data administrators, or system administrators before being activated.

[0098] In one specific embodiment, a medical institution processed the medical records of multiple gout patients over a period of time. Feedback unit 105 statistically analyzed the data and found that the treatment adherence category contributed significantly to the risk stratification of several high-risk patients, while frequently conflicting with medication records. Specifically, the medical record text often contained descriptions such as "the patient reported that the medication was acceptable" or "the patient adjusted their medication after intermittent use." Parsing unit 102 sometimes extracted these as indicating good adherence, but calibration unit 103 determined a directional conflict based on discontinuous prescriptions, elevated blood uric acid, and "discontinued medication on their own" in adjacent records. Based on this, feedback unit 105 marked the treatment adherence category as needing updating and added processing requirements for vague or risky expressions such as "acceptable," "intermittent," and "adjusted on their own" to the dynamic prompt template. This required parsing unit 102 to categorize these expressions as uncertain or risky when generating subsequent prompts, and required the output of original text evidence fragments. After this update, parsing unit 102 can reduce the misjudgment of vague medication descriptions as good adherence when processing similar text in the future.

[0099] The information output by feedback unit 105 can be stored in the patient risk management record or displayed on the clinical decision support interface. For a single patient, the displayed content may include the current risk level, risk level change trend, main contributing semantic categories, key original text evidence fragments, changes in relevant laboratory indicators, and the presence of high-conflict, low-confidence labels. At the system level, feedback unit 105 can output contribution statistics and conflict frequency statistics for each semantic category over a period of time, used to manage the updates of dynamic prompt templates. In this way, feedback unit 105 serves both the risk interpretation of a single patient and the continuous optimization of parsing unit 102.

[0100] Furthermore, the feedback unit is specifically used for: Based on the risk stratification results, the feature items that trigger changes in risk level are read from the patient's temporal multimodal feature vector and categorized into semantic categories corresponding to treatment adherence, lifestyle triggers, joint attack status, tophi description, and renal function risk description, generating a level trigger contribution record. This level trigger contribution record is then associated with conflict source records in the calibration semantic label set to determine the main conflict source, conflict occurrence node, and conflict-suppressed semantic risk features for each semantic category, generating a semantic error attribution record. When the number of level trigger contributions, conflict frequency, and conflict suppression times for the same semantic category all reach the corresponding preset thresholds, a targeted template correction rule is generated based on the main conflict source. Specifically, conflicts in medication records correspond to additional constraints on medication continuity and discontinuation expression extraction; conflicts in laboratory indicators correspond to additional constraints on citing evidence of serum uric acid or renal function indicators; and conflicts in adjacent medical records correspond to additional constraints on attribution of medical time and prohibition of cross-node inference. The dynamic prompt template is updated using targeted template correction rules, and the consistency of parsing before and after the update is verified using historical medical texts with generated semantic error attribution records. If the frequency of conflicts in the corresponding semantic category decreases after the update, an enabled template version is generated. The version number of the enabled template version, the triggering semantic category, the main source of conflict, the targeted template correction rules, and the verification results are fed back to the parsing unit for subsequent generation of disease-specific constraint prompts.

[0101] The feedback unit is used to trace the key features that lead to changes in risk level after the risk stratification unit completes risk stratification, and uses the traceability results to improve the dynamic prompt templates used by the parsing unit in the future. The features that trigger changes in risk level refer to those features in the patient's temporal multimodal feature vector that directly affect the judgment of low, medium, or high risk, such as persistent failure to reach target uric acid levels, discontinuation of uric acid-lowering medication, recent introduction of anti-inflammatory drugs, acute gout diagnosis, poor treatment adherence, clear alcohol consumption triggers, tophi progression, or increased risk of renal function impairment. After reading the risk stratification results, the feedback unit first extracts these features involved in the level judgment from the patient's temporal multimodal feature vector and categorizes them according to their clinical meaning into semantic categories such as treatment adherence, lifestyle triggers, joint attack status, tophi description, and renal function risk description. For example, "febuxostat not renewed for two months" and "discontinued medication on one's own" are categorized into the treatment adherence category, "attack after drinking alcohol" into the lifestyle trigger category, and "acute gout diagnosis" and "joint redness, swelling, and pain" into the joint attack status category. The feedback unit records the above classification results as a level-triggered contribution record. This record includes at least the patient's anonymity identifier, the visit node, the risk level, the feature item that triggers the level change, the semantic category to which it belongs, and the direction of the effect of the feature item on the risk increase or decrease.

[0102] After generating a graded trigger contribution record, the feedback unit associates it with conflict source records in the calibration semantic label set. Semantic error attribution, as referred to here, means determining whether a semantic category has a source of bias in risk assessment due to inconsistencies between text parsing and other diagnostic evidence. The feedback unit matches patient anonymity identifiers, visit nodes, and semantic categories to determine whether a semantic category triggering a change in risk grade also has conflict source records, and further reads whether the conflict source originates from medication records, laboratory indicators, or adjacent visit records. A conflict occurrence node refers to the specific visit node that generates the conflict, such as the current node, the previous follow-up visit node, or the next follow-up visit node. Conflict-suppressed semantic risk features refer to semantic features whose effect has been reduced or masked by the stratification unit because contradictory evidence is no less than supporting evidence. For example, if the medical record text contains "regular medication," but the medication record shows no prescription renewal, and an adjacent visit record shows "discontinued medication on one's own," then "regular medication" can be identified as a conflict-suppressed semantic risk feature, with the main conflict sources being the medication record and adjacent visit records. The feedback unit generates a semantic error attribution record based on this information, so as to determine whether the dynamic prompt template needs to be modified.

[0103] The feedback unit generates a targeted template correction rule only when the number of rank trigger contributions, conflict frequency, and conflict suppression times for the same semantic category within a preset statistical period all reach the corresponding preset thresholds. The preset statistical period can be the most recent month, the most recent 100 parsing tasks, or the medical records of several recent patients; the number of rank trigger contributions indicates the number of times the semantic category actually affects the risk level judgment; the conflict frequency indicates the number of times the parsing results of the semantic category conflict directionally with the calibration evidence; and the number of conflict suppression times indicates the number of times the semantic risk characteristics of the semantic category are reduced or masked by the stratification unit due to conflict. When all three conditions are met simultaneously, it indicates that the semantic category is both important to the stratification results and exhibits repetitive parsing bias, making it suitable for targeted modification of the prompt template. For example, the system can be set to trigger the template correction rule only when the treatment compliance category triggers risk level changes at least five times, conflicts occur at least three times, and conflicts are suppressed at least twice within a statistical period, to avoid frequent template modifications due to occasional errors.

[0104] The targeted template correction rules are generated based on the primary source of conflict and directly apply to the corresponding semantic category constraints in the dynamic prompt template. If the primary source of conflict is medication records, it indicates that the parsing unit's identification of expressions related to medication continuity or discontinuation is not accurate enough. The feedback unit adds extraction constraints for expressions such as prescription continuity, discontinuation, missed doses, intermittent medication, self-reduction, and self-discontinuation in the treatment adherence category, and requires the large language model to cite explicit original text evidence when outputting good adherence. If the primary source of conflict is laboratory indicators, it indicates that the text semantics are inconsistent with the direction of blood uric acid or renal function indicators. The feedback unit adds evidence citation constraints in the blood uric acid control or renal function risk-related categories, requiring the parsing unit not to make judgments based solely on vague terms such as "stable" or "improved," but to simultaneously output original text evidence fragments supporting the judgment. If the primary source of conflict is adjacent medical records, it indicates that the model may be mixing interpretations of disease descriptions from different medical times. The feedback unit adds medical time attribution constraints and prohibits cross-node inference constraints, requiring each semantic label to be generated only based on the current medical time node or text with explicitly labeled time.

[0105] After updating the dynamic prompt template using the aforementioned targeted template correction rules, the feedback unit does not immediately use it for formal parsing. Instead, it uses historical medical records that have already generated semantic error attribution records to verify the consistency of parsing before and after the update. This parsing consistency verification refers to parsing the same batch of historical medical records using both the pre-update and post-update templates, and comparing whether the frequency of conflicts in the corresponding semantic categories has decreased, whether the original text evidence fragments are clearer, and whether the time attribution is more consistent. For example, for multiple historical texts that were previously incorrectly parsed as having good adherence due to "medication is acceptable," if the updated template can output "adherence is uncertain" or "needs to be calibrated in conjunction with medication records," and the number of conflicts with medication records is reduced, then the update is considered effective. If the frequency of conflicts in the corresponding semantic category decreases after the update, the feedback unit generates an enabled template version; if the frequency of conflicts does not decrease or other significant conflicts are added, then this template version is not enabled, and the original dynamic prompt template is retained or awaits manual review.

[0106] After generating the enabled template version, the feedback unit sends the template version number, triggering semantic category, main source of conflict, targeted template correction rules, and verification results back to the parsing unit. The template version number identifies which version of the dynamic prompt template the parsing unit subsequently calls, facilitating the traceability of subsequent semantic tags. The triggering semantic category specifies which category of treatment adherence, lifestyle triggers, joint attack status, description of tophi, or description of renal function risk is being addressed in this update. The main source of conflict indicates that the template modification is based on medication records, laboratory indicators, or adjacent visit records. The verification results show that the frequency of conflicts in historical visit texts for this template version has decreased. Through this process, the feedback unit can transform repetitive semantic biases discovered during risk stratification into executable prompt template constraints, making the subsequent generation of disease-specific constraint prompts by the parsing unit more consistent with the data characteristics of long-term gout patient management scenarios. This improves the stability of semantic parsing, the interpretability of risk stratification, and the adaptability of the system in continuous application.

[0107] In one optional implementation, to enable direct implementation of the calibration and stratification units, the system employs the following quantification rules to complete confidence correction, weighted coding, feature vector construction, and risk stratification. When the initial confidence level output by the parsing unit is set to high, medium, or low, it is converted to 1.0, 0.7, and 0.4, respectively. The calibration unit counts the number of supporting and opposing pieces of evidence for each semantic label. Supporting evidence refers to laboratory indicators, diagnostic codes, medication records, or adjacent medical records that align with the risk direction of the semantic label; opposing evidence refers to the aforementioned evidence that contradicts the risk direction of the semantic label. If there is no contrary evidence and at least one supporting piece of evidence exists, the confidence level is adjusted by increasing the initial confidence level by 0.1, but not exceeding 1.0. If one piece of contrary evidence exists, the confidence level is adjusted by decreasing the initial confidence level by 0.2. If two or more pieces of contrary evidence exist, the confidence level is adjusted by decreasing the initial confidence level by 0.4. If the original evidence contains uncertain words such as "possible," "suspected," "reasonable," "unknown," or "to be investigated," the confidence level is further reduced by 0.1. If the original evidence contains negative words such as "not seen," "denied," "no obvious," or "no recurrence," the risk direction of the label is redefined according to the negative semantic direction. A value below 0 is taken as 0, and a value above 1 is taken as 1. Therefore, "moderately reduced" specifically means a reduction of 0.2, and "significantly reduced" specifically means a reduction of 0.4 or more.

[0108] In one example, the parsing unit extracts the "regular medication" tag from the medical record text with an initial confidence level of 1.0. However, the calibration unit finds that the uric acid-lowering medication has not been renewed for more than 30 days, and a subsequent adjacent medical record states "relapse after self-discontinuation of medication." Both of these contradict the semantic direction of "regular medication," resulting in two pieces of contradictory evidence and a confidence level correction of 0.6. If the original text fragment is "the patient stated that medication was acceptable," and also contains the uncertain word "acceptable," the confidence level correction further decreases to 0.5. This value is written into the calibration semantic tag set and used by the hierarchical unit during weighted encoding.

[0109] The hierarchical unit constructs a node feature vector at each visit node. Each node feature vector includes, in a fixed field order, the following semantic features: serum uric acid target status, direction of serum uric acid change, direction of creatinine change, direction of estimated glomerular filtration rate change, continuity of uric acid-lowering drug prescription, short-term addition of anti-inflammatory drugs, acute gout diagnosis status, treatment adherence semantic features, lifestyle trigger semantic features, joint attack status semantic features, tophi semantic features, renal function risk semantic features, number of conflict sources, and number of conflict suppression markers. In the serum uric acid target status, patients with ordinary gout have a serum uric acid level less than 360 μmol / L, while patients with tophi have a level less than 300 μmol / L. The target status is coded as 0, and the non-target status is coded as 1. In the direction of serum uric acid change, an increase compared to the previous node is coded as 1, a decrease as -1, and a change of less than 30 μmol / L is coded as 0. When creatinine increases by at least 26.5 μmol / L or by 30% compared to the previous node, the direction of change is recorded as 1; otherwise, it is recorded as 0. Similarly, when the estimated glomerular filtration rate (GFR) decreases by 15% compared to the previous node, the direction of change is recorded as 1; otherwise, it is recorded as 0. For uric acid-lowering medication prescription continuity, a prescription interval of no more than 30 days between two consecutive visits is recorded as 1; an interval exceeding 30 days is recorded as 0. For short-term additions of anti-inflammatory drugs, a new colchicine, nonsteroidal anti-inflammatory drug (NSAID), or short-term glucocorticoid is added within 14 days before or after the current node, recorded as 1; otherwise, it is recorded as 0. For acute gout diagnosis, an acute gouty arthritis or acute gout attack diagnosis code is present at the current node, recorded as 1; otherwise, it is recorded as 0.

[0110] For semantic features, the hierarchical unit first encodes the risk direction of the label into three states: increased risk, decreased risk, or neutral risk, denoted as 1, -1, and 0 respectively. This is then multiplied by the corresponding confidence correction value to obtain the semantic risk encoding value. For example, the risk direction of "discontinuing medication on one's own" is increased risk, with a confidence correction value of 0.9, so the treatment adherence semantic feature is 0.9; the risk direction of "regular medication use" is decreased risk, with a confidence correction value of 0.8, so the treatment adherence semantic feature is -0.8; and the risk direction of "not mentioned" is neutral, with a semantic feature of 0. If a semantic label has conflict source records, and the number of contradictory evidences is not less than the number of supporting evidences, a conflict suppression label is generated. This semantic risk encoding value is then multiplied by 0.2 before being written into the node feature vector to reduce the impact of obviously conflicting text labels on the risk stratification results.

[0111] The stratification unit concatenates the node feature vectors of the three most recent visit nodes according to the visit time sequence to form a patient temporal multimodal feature vector. If a patient has fewer than three visit nodes, the fields of the missing nodes are filled with 0, and a missing marker is set. This patient temporal multimodal feature vector serves as the input to the risk stratification model. In this embodiment, the risk stratification model adopts a deterministic rule classification model, which can be implemented without relying on additional training data. Its input is the aforementioned patient temporal multimodal feature vector, and its output includes recurrence risk, tophi or chronic progression risk, renal function complication risk, and final risk level.

[0112] Specifically, recurrence risk is determined according to the following rules: 2 points for current serum uric acid not reaching the target level; 1 point for two consecutive periods of elevated serum uric acid; 3 points for the current period of acute gout diagnosis; 2 points for the current period of short-term new anti-inflammatory medication use; 2 points for a joint attack semantic feature greater than 0.6; 1 point for a lifestyle trigger semantic feature greater than 0.6; and 1 point for treatment adherence semantic feature greater than 0.6. A total recurrence risk score of 6 or above indicates high risk of recurrence, 3 to 5 points indicates medium risk, and below 3 points indicates low risk. The risk of tophi or chronic progression is determined according to the following rules: 3 points for a tophi semantic feature greater than 0.6; 2 points for two or more consecutive periods of uncontrolled serum uric acid; 2 points for two or more recent periods of joint attack risk; and 2 points for a gout course or diagnostic record indicating chronic gout. A total score of 5 or above indicates high risk of chronic progression, 3 to 4 points indicates medium risk, and below 3 points indicates low risk. The risk of renal complications is determined according to the following rules: 1 point for the direction of change in creatinine, 2 points for the direction of change in estimated glomerular filtration rate, 3 points for the presence of renal function-related diagnoses, and 2 points for a semantic feature of renal function risk greater than 0.6. A total score of 5 or above indicates a high risk of renal complications, 3 to 4 points indicates a medium risk, and below 3 points indicates a low risk.

[0113] The final risk level is determined based on the combination of three risk categories. A high risk is output when any one of the following risks—recurrence risk, tophi or chronic progression risk, or renal complication risk—is high, and the semantic risk feature supporting this high-risk judgment is not suppressed by conflict. A medium risk is output when no high risk exists, but any one of the following risks is medium, or when there is any persistent abnormality in any of the following: serum uric acid levels failing to reach the target for two consecutive periods, discontinuous uric acid-lowering medication regimen, or persistent abnormal renal function indicators. All other risks are output as low risk. If a high-risk judgment relies entirely on semantic risk features suppressed by conflict, this high risk is not directly used as the final high-risk basis but is downgraded to a medium-risk candidate, and the corresponding conflict source is recorded by the feedback unit.

[0114] For example, a patient's serum uric acid levels at their three most recent medical visits were 520 μmol / L, 540 μmol / L, and 560 μmol / L, all below target and continuously rising. At the current visit, an acute gouty arthritis diagnosis has been made, and a new colchicine prescription has been added. The parsing unit extracts two semantic tags: "redness, swelling, and pain in the left foot after drinking alcohol" and "discontinuation of medication on one's own," with confidence correction values ​​of 0.9 and 0.8, respectively, and neither is subject to conflict suppression. When the stratified unit calculates the recurrence risk based on this, 2 points are awarded for below-target serum uric acid, 1 point for continuously rising levels, 3 points for the acute gout diagnosis, 2 points for the newly added anti-inflammatory medication, 2 points for the semantic features of joint attacks, 1 point for lifestyle triggers, and 1 point for treatment adherence risk. The total recurrence risk score is 12 points, therefore the recurrence risk is high, and the final output is "high risk." For example, if another patient's serum uric acid is slightly higher than the target value twice consecutively, but there is no diagnosis of acute exacerbation, no new anti-inflammatory drugs are added, medication is continued, and the text only mentions "occasional discomfort" which has been corrected to low confidence, then this only constitutes a persistent abnormal state, and the output is medium risk. If the patient's serum uric acid is within the target range, there are no acute exacerbations, medication is continued, and there is no evidence of tophi or renal function risk, then the output is low risk.

[0115] The preset statistical period in the feedback unit defaults to the 100 most recent stratified patient visit nodes; the threshold for the number of contribution triggers is 5, the threshold for the frequency of conflicts is 3, and the threshold for the number of times conflicts are suppressed is 2. For a certain semantic category, a targeted template correction rule is generated only when all three thresholds are reached simultaneously within the same statistical period. The consistency verification of parsing before and after the update is performed using historical patient visit texts that have already generated semantic error attribution records. If the updated template reduces the frequency of conflicts for that semantic category by at least 20% compared to before the update, an enabled template version is generated; if the reduction is less than 20%, the template version is not enabled. Through the above specific rules, those skilled in the art can clearly implement the multimodal risk stratification process of this application based on the input data, coding rules, confidence correction methods, and risk scoring thresholds.

[0116] Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims of this application.

Claims

1. A multimodal risk stratification system for gout patients based on cue engineering, characterized in that, include: The building unit is used to acquire structured diagnosis and treatment data and unstructured medical record text of gout patients, standardize the structured diagnosis and treatment data and integrate the consultation time to generate a patient time-series diagnosis and treatment dataset; The parsing unit is used to generate disease-specific constraint prompts based on the gout diagnosis and treatment knowledge base, the patient time-series diagnosis and treatment dataset, and the dynamic prompt template. This drives the large language model to extract treatment compliance, lifestyle triggers, joint attack status, description of tophi, and description of renal function risk from unstructured medical record text, and generate a set of structured semantic tags carrying original text evidence fragments and consultation time. The calibration unit is used to perform directional conflict verification between the structured semantic tag set and laboratory indicators, diagnostic codes, medication records and adjacent medical records, and generate confidence correction values ​​based on the conflict source and original evidence fragments to form a calibration semantic tag set. The hierarchical unit is used to align the patient time-series diagnosis and treatment dataset, calibration semantic label set, laboratory indicators, diagnostic codes and medication records along the same time axis, and perform weighted encoding based on confidence correction values ​​to generate patient time-series multimodal feature vectors, and obtain risk stratification results of low risk, medium risk or high risk; The feedback unit is used to generate feature contribution explanation information and risk evolution records aggregated by semantic category based on risk stratification results and patient temporal multimodal feature vectors, and feeds back conflict sources and feature contribution explanation information to the parsing unit. It also updates the extraction constraints of dynamic prompt templates for semantic categories where both contribution and conflict frequency exceed preset thresholds.

2. The multimodal risk stratification system for gout patients based on cue engineering according to claim 1, characterized in that, The calibration unit is specifically used for: Based on the semantic categories and consultation times in the structured semantic tag set, laboratory indicators, diagnostic codes, and medication records of the same consultation node and its one adjacent consultation node before and after are extracted from the patient time-series diagnosis and treatment dataset, and candidate verification evidence groups corresponding to each semantic tag are generated. Based on semantic categories, rules for determining the direction of gout risk are configured for candidate validation evidence groups. Specifically, the treatment adherence label is combined with the continuity of uric acid-lowering drug prescriptions, discontinuation descriptions, and changes in blood uric acid to determine the support, opposition, or irrelevant status of each candidate validation evidence. The joint attack status label is combined with acute gout diagnosis codes, short-term prescriptions for anti-inflammatory and analgesic drugs, and pain descriptions to determine the support, opposition, or irrelevant status of each candidate validation evidence. The renal function risk description label is combined with changes in creatinine, estimated changes in glomerular filtration rate, and renal function-related diagnoses to determine the support, opposition, or irrelevant status of each candidate validation evidence, thus forming an evidence direction record. The consistency of the evidence direction record is compared with the original evidence fragment with the corresponding semantic tag. When the semantic direction of the original evidence fragment is opposite to the direction of at least two candidate verification evidence, a conflict source record containing conflict data source, conflict treatment node and conflict semantic category is generated. Based on the conflict source record, the number of supporting evidence, the number of contradictory evidence, and whether the original evidence fragment contains negative or uncertain words, generate the confidence correction value of the corresponding semantic label; The confidence correction value and conflict source record are written into the corresponding structured semantic label to form a set of calibration semantic labels for weighted encoding by the hierarchical unit.

3. The multimodal risk stratification system for gout patients based on cue engineering according to claim 1, characterized in that, The hierarchical unit is specifically used for: Based on the patient's time of visit in the patient's time-series diagnosis and treatment dataset, a patient node sequence is established, and the calibration semantic label set, laboratory indicators, diagnostic codes and medication records are assigned to the corresponding visit nodes to generate node-level aligned data. From the node-level aligned data, extract the following data: serum uric acid target status, direction of continuous change in serum uric acid, direction of change in creatinine and estimated glomerular filtration rate, continuity of uric acid-lowering drug prescriptions, short-term new anti-inflammatory drug status, and acute gout diagnosis status, and generate structured risk feature groups. Treatment adherence, lifestyle triggers, joint attack status, description of tophi, and description of renal function risk are converted into semantic risk features, and the coding weights are adjusted according to the corresponding confidence correction values. For semantic risk features with conflicting source records and the number of contrary evidence is not less than the number of supporting evidence, conflict suppression labels are generated. Based on the structured risk feature group, semantic risk features, conflict suppression markers, and changes in blood uric acid, medication continuity, and seizure status between adjacent visit nodes, a patient temporal multimodal feature vector is generated. Based on the patient's temporal multimodal feature vector, the risk of recurrence, the risk of tophi or chronic progression, and the risk of renal complications are determined respectively. When any risk reaches the high-risk condition and the corresponding semantic risk feature is not suppressed by conflict, a high risk is output. When there is only a single persistent abnormality that does not reach the high-risk condition, a medium risk is output. The rest are output as low risk.

4. The multimodal risk stratification system for gout patients based on cue engineering according to claim 1, characterized in that, The feedback unit is specifically used for: Based on the risk stratification results, the feature items that trigger changes in risk level are read from the patient’s temporal multimodal feature vector and classified into the semantic categories corresponding to treatment adherence, lifestyle triggers, joint attack status, tophi description and renal function risk description, and a level trigger contribution record is generated. The level-triggered contribution record is associated with the conflict source record in the calibration semantic label set to determine the main conflict source, conflict occurrence node and conflict-suppressed semantic risk characteristics of each semantic category, and generate semantic error attribution record; When the number of times the level triggers contribution, the frequency of conflict, and the number of times the conflict is suppressed all reach the corresponding preset thresholds for the same semantic category, a targeted template correction rule is generated based on the main source of conflict. Among them, for medication record conflicts, constraints on medication continuity and discontinuation expression extraction are added; for laboratory indicator conflicts, constraints on citing evidence of blood uric acid or renal function indicators are added; and for conflicts between adjacent medical records, constraints on the attribution of medical time and prohibition of cross-node inference are added. The dynamic prompt template is updated using targeted template correction rules, and the consistency of parsing before and after the update is verified using historical medical texts with generated semantic error attribution records. If the frequency of conflict in the corresponding semantic category decreases after the update, an enabled template version is generated. The version number of the enabled template, the trigger semantic category, the main source of conflict, the targeted template correction rules, and the verification results are fed back to the parsing unit for subsequent generation of disease constraint prompt words.

5. The multimodal risk stratification system for gout patients based on prompting engineering according to claim 1, characterized in that, The parsing unit is specifically used for: Based on the patient's time-series medical records, extract the unstructured medical record text, consultation time, blood uric acid status, uric acid-lowering drug record, acute attack-related diagnosis, and state summary of the previous and next adjacent consultation nodes from the current consultation node, and generate a node parsing input package. Based on the template version number returned by the gout diagnosis and treatment knowledge base and feedback unit, positive expression, negative expression, uncertain expression, time attribution rules, cross-node inference prohibition rules and original text evidence extraction rules are configured for treatment compliance, lifestyle triggers, joint attack status, description of tophi and description of renal function risk, respectively, and a semantic category constraint table is generated. The node parsing input package, semantic category constraint table and dynamic prompt template are combined into disease constraint prompt words, so that the large language model can generate candidate semantic labels only based on the current consultation node text or the original text content with a clear time reference. Each candidate semantic label is required to carry the label value, consultation time, original text evidence fragment and initial confidence level. The candidate semantic tags are verified for evidence integrity and time attribution. Tags that do not carry original text evidence fragments and do not belong to the unmentioned state are deleted, and tags whose evidence fragment time points are inconsistent with the current medical treatment node are marked as cross-node suspected tags. Candidate semantic tags that have not been deleted, suspected cross-node tags, and their corresponding original text evidence fragments are written into a structured semantic tag set according to the visit node, and suspected cross-node tags are provided to the calibration unit as priority verification objects for directional conflict verification.