Animal epidemic disease sub-diagnosis method and system based on multi-modal large model architecture
Patent Information
- Application Number
- CN202611156512.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-31
- Publication Date
- 2026-09-22
AI Technical Summary
[0004]本发明旨在提供一种基于多模态大模型架构的动物疫病分诊断方法及系统,有效解决了多模态数据语义断层与诊疗流程环节脱节的问题,显著提升了分诊断的可解释性、资源调度的适配精度及区域预警的时效性
本发明通过将症状描述文本数据及病灶部位图像数据构成的多模态原始数据统一映射至知识图谱本体层生成病例知识子图,使原本语义割裂的多源信息获得结构化、可推理的共同表示基础,从而支撑后续语义检索与规则推理生成兼具数据驱动与知识约束的候选疫病集合及置信度;在此基础上,综合风险评分融合病例特征与区域历史疫情背景因子,任务派发实例依据风险等级在专家能力子图中动态匹配处置主体,现场核验数据经冲突仲裁后增量更新病例知识子图并触发状态重评估,最终汇聚状态更新事件中的确诊病例时空分布执行空间推理生成区域预警报告,实现了从个体感知到区域态势的完整知识价值链闭环,有效解决了多模态数据语义断层与诊疗流程环节脱节的问题,显著提升了分诊断的可解释性、资源调度的适配精度及区域预警的时效性。
Smart Images

Figure CN122800199A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of knowledge graph technology, specifically to a method and system for the sub-diagnosis of animal diseases based on a multimodal large model architecture. Background Technology
[0002] Current animal disease triage and diagnosis systems primarily rely on primary-level veterinarians or farmers submitting single or limited modal data, such as symptom descriptions and lesion images, via mobile devices. These data are then initially classified and screened for risk by a back-end expert system or traditional machine learning model. This approach typically processes each modality of data independently before directly inputting it into the classifier, lacking a unified semantic representation framework. This prevents the knowledge-level association and fusion of key entities in unstructured text and visual features in images. Furthermore, the diagnostic reasoning process is fragmented from subsequent resource allocation, on-site verification, and regional early warning systems, failing to form a dynamic closed loop centered on case knowledge. Consequently, diagnostic results struggle to support accurate tiered treatment and situation assessment.
[0003] In practical applications, the core technical problem of semantic fragmentation of multimodal data and disconnection from the diagnosis and treatment process has been exposed. Specifically, after heterogeneous processing, the inherent correlation between different modalities of data is lost, causing the generation of candidate diseases to rely only on shallow feature matching, and making it impossible to use the multidimensional relationships between diseases, symptoms and environment accumulated in the knowledge graph for deep reasoning; and due to the lack of a unified knowledge carrier throughout the entire process, new evidence from on-site verification cannot automatically backtrack and correct previous diagnostic conclusions, and regional early warning can only be based on discrete case statistics and ignore spatiotemporal clustering patterns, ultimately resulting in limited accuracy of subdiagnosis, frequent resource misallocation, and delayed early warning. Summary of the Invention
[0004] This invention aims to provide a method and system for sub-diagnosis of animal diseases based on a multimodal large model architecture, which effectively solves the problem of semantic fragmentation of multimodal data and disconnection from diagnosis and treatment process links, and significantly improves the interpretability of sub-diagnosis, the adaptability of resource scheduling, and the timeliness of regional early warning.
[0005] To achieve the above objectives, the technical solution adopted by this invention is: a method for sub-diagnosis of animal diseases based on a multimodal large model architecture, comprising: Multimodal raw data, consisting of symptom description text data and lesion site image data, is mapped to the ontology layer of the knowledge graph to generate a case knowledge subgraph. Semantic retrieval and rule-based reasoning are performed based on this subgraph to generate a candidate disease set and its confidence level. The candidate disease set and its confidence level are then fused with case characteristics to quantify the comprehensive risk score and classify risk levels. Based on the risk level, a treatment entity is matched in the expert capability subgraph of the knowledge graph to generate a task assignment instance. On-site verification data is received to incrementally update the case knowledge subgraph, triggering a status reassessment and generating a status update event. The spatiotemporal distribution of confirmed cases in the status update event is aggregated, and spatial reasoning is performed to generate a regional early warning report.
[0006] Preferably, the multimodal raw data consisting of symptom description text data and lesion site image data is mapped to the knowledge graph ontology layer to generate a case knowledge subgraph, including: Modal separation and metadata extraction are performed on the symptom description text data and lesion site image data to generate a modal data stream; Based on the modal data stream, text semantic segments, image semantic segments, and [other segments] are extracted respectively to form a set of each modal semantic segment; Align the sets of semantic fragments of each modality with the knowledge graph ontology mode layer to generate a set of normalized triples; Based on the normalized triple set, a case knowledge subgraph is generated and the root node Uniform Resource Identifier is returned.
[0007] Preferably, semantic retrieval and rule-based reasoning are performed based on the case knowledge subgraph to generate a candidate disease set and confidence level, including: Starting from the root node Uniform Resource Identifier, traverse the knowledge graph to obtain an initial candidate disease pool; For each disease in the initial candidate disease pool, a semantic similarity score with the case knowledge subgraph is calculated; The case data credibility correction factor and the knowledge timeliness decay factor are superimposed on the semantic similarity score to generate a comprehensive confidence score; Based on the comprehensive confidence score, a set of candidate diseases is extracted and accompanied by a list of key identification points and suggested verification items.
[0008] Preferably, the candidate disease set and confidence level are integrated with case characteristics to quantify the comprehensive risk score and classify the risk level, including: The inherent risk attributes are extracted from the candidate disease set and combined with the confidence level to calculate the weighted disease risk baseline value; Based on the aforementioned case knowledge subgraph, real-time feature indicators are extracted to form a case-specific risk vector; The weighted disease risk baseline value is non-linearly fused with the case-specific risk vector and superimposed with regional historical epidemic background factors to generate a comprehensive risk score; The risk levels are classified based on the comprehensive risk score and packaged into a risk decision package.
[0009] Preferably, based on the risk level, a handling entity is matched in the expert capability subgraph of the knowledge graph to generate a task dispatch instance, including: Activate the corresponding set of candidate disposal subjects based on the risk level labels in the risk decision package; For each entity in the candidate disposal entity set, a multi-dimensional matching score is calculated, consisting of professional suitability, geographical proximity, language compatibility, and load balancing. The multi-dimensional matching scores are weighted and fused to generate the final matching score, and the first and second subjects are selected to create a task dispatch instance. The task dispatch instance and associated information are encapsulated into a structured task instruction and pushed to the target task's main terminal.
[0010] Preferably, receiving on-site verification data to incrementally update the case knowledge subgraph triggers a state reassessment and generates a state update event, including: Receive on-site verification data and perform structured parsing and source credibility labeling to generate incremental knowledge fragments; Align the incremental knowledge fragments with the current version of the case knowledge subgraph, perform conflict arbitration based on source credibility tags and data type priority, merge them to generate a new version, and record the change log. Based on the new version of the case knowledge subgraph, the confidence scores and comprehensive risk scores of candidate diseases are recalculated and significant changes are identified. The status label and handling protocol guidelines are updated based on the significant changes and encapsulated as a status update event.
[0011] Preferably, the spatiotemporal distribution of confirmed cases in the status update events is aggregated, and spatial reasoning is performed to generate regional early warning reports, including: Filter out confirmed or high-risk events from the status update events and write them into the regional epidemic time series subgraph as observation points; Perform spatiotemporal neighborhood retrieval and calculate the spatiotemporal clustering index centered on the observation point; The spatiotemporal clustering index is fused with auxiliary indicators such as the laboratory positive rate and the case growth rate to generate a regional early warning score and classify early warning levels. Extract response suggestion templates based on the aforementioned warning level and fill them in to generate a regional warning report.
[0012] Preferably, a comprehensive confidence score is generated by superimposing a case data credibility correction factor and a knowledge timeliness decay factor onto the semantic similarity score, including: The confidence correction factor for case data was determined based on the confidence-weighted average of text semantic segments and image semantic segments; The knowledge timeliness decay factor is determined by normalizing the historical occurrence frequency of diseases in the current spatiotemporal context. The semantic similarity score, the case data credibility correction factor, and the knowledge timeliness decay factor are linearly combined according to preset weights to generate a comprehensive confidence score.
[0013] Preferably, the weighted fusion of the multidimensional matching scores to generate the final matching score includes: The weighting coefficients of professional fit, geographical proximity, language compatibility, and load balancing are dynamically adjusted based on the comprehensive risk score. The final matching score is generated by multiplying the scores of each dimension in the multidimensional matching score with the corresponding weight coefficients and then summing the results.
[0014] On the other hand, this invention proposes an animal disease triage and diagnosis system based on a multimodal large model architecture, comprising: The data mapping unit is used to map multimodal raw data, consisting of symptom description text data and lesion site image data, to the knowledge graph ontology layer to generate a case knowledge subgraph. The semantic retrieval unit is used to perform semantic retrieval and rule reasoning based on the case knowledge subgraph to generate a candidate disease set and confidence level. The risk quantification unit is used to integrate the candidate disease set, confidence level, and case characteristics to quantify the comprehensive risk score and classify the risk level. The resource matching unit is used to match the handling subject in the knowledge graph expert capability subgraph according to the risk level and generate a task dispatch instance; The incremental update unit is used to receive on-site verification data to incrementally update the case knowledge subgraph, trigger state reassessment, and generate a state update event. The early warning generation unit is used to aggregate the spatiotemporal distribution of confirmed cases in the status update events and perform spatial reasoning to generate regional early warning reports.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention maps multimodal raw data, consisting of symptom description text data and lesion site image data, onto a knowledge graph ontology layer to generate a case knowledge subgraph. This provides a structured and reasonable common representation basis for previously semantically fragmented multi-source information, supporting subsequent semantic retrieval and rule-based reasoning to generate a candidate disease set and confidence level that is both data-driven and knowledge-constrained. Based on this, risk scoring is integrated with case characteristics and regional historical epidemic background factors. Task assignment instances are dynamically matched with the handling subject in the expert capability subgraph according to the risk level. On-site verification data is incrementally updated in the case knowledge subgraph after conflict arbitration, triggering a state reassessment. Finally, the spatiotemporal distribution of confirmed cases in the state update event is aggregated, and spatial reasoning is performed to generate a regional early warning report. This achieves a complete closed-loop knowledge value chain from individual perception to regional situation, effectively solving the problem of semantic fragmentation of multimodal data and disconnection from the diagnosis and treatment process. It significantly improves the interpretability of subdiagnosis, the accuracy of resource scheduling, and the timeliness of regional early warning. Attached Figure Description
[0016] Figure 1 This is a flowchart of the animal disease sub-diagnosis method based on a multimodal large model architecture according to the present invention; Figure 2 This is a block diagram of the animal disease sub-diagnosis system based on a multimodal large model architecture according to the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0018] It should be noted that if the embodiments of the present invention involve directional indicators (such as up, down, left, right, front, back, etc.), the directional indicators are only used to explain the relative positional relationship and movement of the components in a specific posture. If the specific posture changes, the directional indicators will also change accordingly.
[0019] Furthermore, if the embodiments of this invention involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the use of "and / or" or "and / or" throughout the text includes three parallel solutions. For example, "A and / or B" includes solution A, solution B, or a solution where both A and B are satisfied simultaneously. Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.
[0020] Please see Figure 1 This invention provides a method for the sub-diagnosis of animal diseases based on a multimodal large model architecture, comprising: Multimodal raw data, consisting of symptom description text data and lesion site image data, is mapped to the knowledge graph ontology layer to generate a case knowledge subgraph; specifically including: Modal separation and metadata extraction are performed on symptom description text data and lesion site image data to generate modal data streams; text semantic fragments and image semantic fragments are extracted from the modal data streams to form sets of semantic fragments for each modality; the sets of semantic fragments for each modality are aligned with the ontology modality layer of the knowledge graph to generate a set of normalized triples; the case knowledge subgraph is generated based on the set of normalized triples and the root node Uniform Resource Identifier is returned.
[0021] By separating modal data into independent semantic fragments and aligning them with the ontology pattern layer to generate normalized triples, the transformation of multi-source raw information into a unified knowledge representation is realized. This enables the case knowledge subgraph to possess the characteristics of being structured, traceable, and cross-modal, providing a semantically consistent and machine-understandable data foundation for subsequent semantic retrieval and rule-based reasoning. It effectively avoids the problems of semantic loss and knowledge fragmentation caused by modal heterogeneity.
[0022] Based on the case knowledge subgraph, semantic retrieval and rule-based reasoning are performed to generate a candidate disease set and confidence level; specifically including: Starting from the root node Uniform Resource Identifier, the knowledge graph is traversed to obtain an initial candidate disease pool. For each disease in the initial candidate disease pool, the semantic similarity score with the case knowledge subgraph is calculated. The case data credibility correction factor and knowledge timeliness decay factor are superimposed on the semantic similarity score to generate a comprehensive confidence score. The candidate disease set is extracted based on the comprehensive confidence score and includes identification tips and a list of suggested verification items.
[0023] The process involves superimposing the case data credibility correction factor and the knowledge timeliness decay factor onto the semantic similarity score to generate a comprehensive confidence score. This includes: determining the case data credibility correction factor based on the weighted average confidence of text semantic segments and image semantic segments; determining the knowledge timeliness decay factor based on the normalization of the historical occurrence frequency of the disease in the current spatiotemporal context; and linearly combining the semantic similarity score with the case data credibility correction factor and the knowledge timeliness decay factor according to preset weights to generate a comprehensive confidence score.
[0024] By introducing a case data credibility correction factor and a knowledge timeliness decay factor to perform multidimensional calibration of semantic similarity scores, the interference of low-quality multimodal input and outdated knowledge on diagnostic results is effectively suppressed, making the ranking of candidate diseases more consistent with the current case status and regional epidemic patterns, thereby improving the robustness and clinical applicability of diagnostic conclusions.
[0025] By integrating candidate disease sets, confidence levels, and case characteristics, a comprehensive risk score is quantified and risk levels are classified; specifically including: The inherent risk attributes of the candidate disease set are extracted and weighted disease risk baseline values are calculated by combining confidence scores; real-time feature indicators are extracted based on the case knowledge subgraph to form case-specific risk vectors; the weighted disease risk baseline values and case-specific risk vectors are nonlinearly fused and superimposed with regional historical epidemic background factors to generate a comprehensive risk score; risk levels are classified according to the comprehensive risk score and packaged into a risk decision package.
[0026] By nonlinearly fusing and quantifying the inherent risk of the disease, real-time characteristics of cases, and historical epidemic background of the region, the risk score reflects both the immediate severity of individual cases and the dynamic impact of the regional epidemic situation. This avoids misjudgment or omission caused by relying on only a single dimension, and improves the scientific nature of risk level classification and the pertinence of treatment decisions.
[0027] Based on the risk level, the relevant entity is matched with the expert capability subgraph of the knowledge graph to generate a task assignment instance; specifically including: Based on the risk level labels in the risk decision package, activate the corresponding candidate disposal subject set; calculate a multi-dimensional matching score for each subject in the candidate disposal subject set, composed of professional suitability, geographical proximity, language compatibility and load balancing; weight and fuse the multi-dimensional matching scores to generate the final matching score and select the first subject to create a task dispatch instance; encapsulate the task dispatch instance and related information into a structured task instruction and push it to the target task subject terminal.
[0028] The final matching score is generated by weighting and fusing the multi-dimensional matching scores. This includes: dynamically adjusting the weight coefficients of professional fit, geographical proximity, language compatibility and load balance based on the comprehensive risk score; and summing the scores of each dimension of the multi-dimensional matching score after multiplying them by their corresponding weight coefficients to generate the final matching score.
[0029] By dynamically adjusting multi-dimensional matching weights based on risk levels and comprehensively assessing the professional, geographical, linguistic, and workload status of the handling entities, adaptive matching of emergency resources and case needs is achieved. This avoids the problems of insufficient professional capabilities in high-risk scenarios or waste of high-quality resources in low-risk scenarios, and significantly improves the response efficiency and suitability of task assignment.
[0030] The system receives on-site verification data and incrementally updates the case knowledge subgraph, triggering a state reassessment and generating a state update event; specifically including: Receive on-site verification data and perform structured parsing and source credibility labeling to generate incremental knowledge fragments; align the incremental knowledge fragments with the current version of the case knowledge subgraph, and merge them to generate a new version after conflict arbitration based on source credibility labels and data type priority, and record the change log; recalculate the candidate disease confidence and comprehensive risk score based on the new version of the case knowledge subgraph and identify significant changes; update the status label and treatment protocol guidelines based on significant changes and encapsulate them as status update events.
[0031] By introducing conflict arbitration based on source credibility and data type priority and recording change logs, the accuracy and traceability of on-site verification data when integrated into the knowledge graph are ensured. This allows the case status to be dynamically corrected with new evidence rather than statically fixed, effectively avoiding the continuation of misjudgments due to information lag or errors, and improving the real-time adaptability and closed-loop error correction capability of diagnosis and treatment decisions.
[0032] The system aggregates the spatiotemporal distribution of confirmed cases in status update events and performs spatial inference to generate regional early warning reports; specifically including: The system filters confirmed or high-risk events from status update events and writes them into the regional epidemic time series sub-graph as observation points; it performs spatiotemporal neighborhood retrieval and calculates spatiotemporal clustering index with the observation points as the center; it integrates the spatiotemporal clustering index with auxiliary indicators composed of laboratory positive rate and case growth rate to generate regional early warning score and classify early warning level; it extracts response suggestion template based on early warning level and fills it to generate regional early warning report.
[0033] By integrating spatiotemporal clustering with multi-dimensional indicators such as the proportion of positive laboratory results and the rate of case growth, regional early warning scoring was achieved, enabling early identification and dynamic grading of the epidemic spread trend, thus avoiding delayed response or over-warning caused by relying solely on the number of cases.
[0034] On the other hand, this embodiment proposes an animal disease triage and diagnosis system based on a multimodal large model architecture, such as... Figure 2 As shown, it includes: The data mapping unit is used to map multimodal raw data, consisting of symptom description text data and lesion site image data, to the knowledge graph ontology layer to generate a case knowledge subgraph. The semantic retrieval unit is used to perform semantic retrieval and rule reasoning based on the case knowledge subgraph to generate a candidate disease set and confidence level. The risk quantification unit is used to integrate the candidate disease set, confidence level, and case characteristics to quantify the comprehensive risk score and classify the risk level. The resource matching unit is used to match the handling subject in the knowledge graph expert capability subgraph according to the risk level and generate a task dispatch instance; The incremental update unit is used to receive on-site verification data to incrementally update the case knowledge subgraph, trigger state reassessment, and generate a state update event. The early warning generation unit is used to aggregate the spatiotemporal distribution of confirmed cases in the status update events and perform spatial reasoning to generate regional early warning reports.
[0035] Furthermore, the aforementioned units, during execution, are also used to implement other steps of the aforementioned animal disease sub-diagnosis method based on a multimodal large model architecture, as follows: Step 1: Standardized mapping and instantiation of multimodal raw data to the knowledge graph ontology layer; This step, as the initial stage of the entire animal disease triage process, bears the core responsibility of transforming unstructured and semi-structured multimodal raw information submitted by farmers or grassroots veterinarians into computer-understandable and reasonable knowledge graph instances. This process is not a simple data format conversion, but rather, based on predefined veterinary domain ontology specifications, it performs deep semantic analysis and structured reorganization of various modalities such as text and images, ultimately generating a case knowledge subgraph strictly aligned with the global knowledge graph's schema layer. This provides a standardized data foundation for subsequent disease inference and risk assessment, specifically including: Step 1.1: Perform modal separation and basic metadata extraction on the user-uploaded multimodal raw data packets to generate a modal data stream with timestamps and source identifiers. In this step, the raw data packets received by the system typically contain symptom description text, photos of lesion sites, and basic aquaculture record information.
[0036] The processing unit first unpacks the data packets based on file extensions and MIME types, separating data of different modalities into independent processing queues. For each separated modal data, the system automatically appends basic metadata such as collection timestamp, uploading device identifier, geographic location coordinates (if available), and user identity hash, forming a modal data stream with contextual information. This modal data stream not only preserves the integrity of the original content but also provides necessary traceability evidence for subsequent cross-modal correlation and credibility assessment through the injection of metadata.
[0037] Step 1.2: Based on the modal data stream, call the corresponding semantic parsing unit to extract entities, attributes and preliminary relationships, forming a set of independent semantic fragments for each modality.
[0038] For text modalities, the natural language understanding component extracts entities and their attribute values such as animal species, age, body temperature, feed intake, excrement characteristics, and duration of illness from symptom descriptions based on veterinary domain dictionaries and syntactic analysis rules. It also identifies the modification and causal relationships between entities. For example, lameness in the left hind leg with joint swelling is parsed into three related semantic units: location (left hind leg), symptom (lameness), and accompanying symptom (joint swelling).
[0039] For image modalities, the visual feature extraction unit outputs structured labels such as lesion type, anatomical location, area ratio, and color distribution through a pre-trained lesion recognition network, and organizes these labels into image semantic fragments isomorphic to the text semantic unit.
[0040] Step 1.3: Align and resolve conflicts between the multimodal semantic fragment set and the predefined veterinary knowledge graph ontology pattern layer to generate a set of normalized triples that conform to ontology constraints.
[0041] The ontology schema layer of the veterinary knowledge graph predefines the core concept system in the field of animal diseases, including entity types such as animal species, breeds, physiological stages, symptoms and signs, pathological changes, laboratory indicators, veterinary drugs, vaccines, diseases, transmission routes, hosts, and environmental factors, as well as relationship types such as manifestation, occurrence, induction, treatment, prevention, and transmission. Value range constraints, cardinality constraints, and logical axioms are set for each entity and relationship.
[0042] During the alignment process, the system maps entities and attributes in each modal semantic fragment to corresponding classes and data attributes in the ontology. For expressions that cannot be directly matched, the system searches for the closest ontology element through thesaurus, hypernym reasoning, or fuzzy matching.
[0043] When semantic fragments from different modalities give contradictory descriptions of the same attribute, the system arbitrates according to preset modality priority rules and confidence thresholds, giving priority to the results of objective measurement modalities, and retaining the conflict record as an uncertainty marker in the annotation attribute of the triple.
[0044] Step 1.4: Based on the normalized set of triples, assemble a unique case knowledge subgraph according to the case instantiation template, persist it to the graph database, and return the root node URI of the subgraph as the reference anchor for subsequent steps.
[0045] The case instantiation template defines the minimum set of triples required to constitute a complete case instance and optional extended structures, including the required animal individual instance, symptom observation event instance, and reporter instance, as well as optional immunization record instance, medication record instance, environmental exposure instance, etc.
[0046] Based on template rules, the system replaces blank nodes in the normalized triples with globally unique URIs, establishes object attribute links between instances, and forms a star or network subgraph structure centered on the root node of the case and radiating to connect various related entities.
[0047] Before persistent storage, the system performs an integrity check to ensure that all required attributes have been assigned values, all relationships point to valid targets, and all data types conform to the XSD specification. Upon successful storage, the system generates a unique identifier for the knowledge subgraph of this case, which is the root node URI.
[0048] Step 2: Generation and confidence quantification of candidate diseases based on knowledge graph semantic retrieval and rule reasoning; After constructing and persisting the case knowledge subgraph, this step takes the root node URI of the subgraph and uses the disease-symptom association rules, historical case similarity measures, and epidemiological prior knowledge contained in the knowledge graph to systematically screen and rank the types of diseases that the current case may involve. This process goes beyond the limitations of traditional keyword matching. Through path traversal on the graph, subgraph isomorphism detection, and semantic similarity calculation, it uncovers implicit associations beyond explicit descriptions, thereby generating a candidate disease list that covers both typical manifestations and atypical variations. Each candidate disease is assigned a confidence score that comprehensively reflects the strength of evidence support and knowledge consistency, specifically including: Step 2.1: Starting from the root node URI of the case knowledge subgraph, perform forward traversal and backward tracing based on the symptom-disease association edge in the knowledge graph to obtain the set of all disease entities that are directly or indirectly related to the current case symptom set, forming the initial candidate disease pool.
[0049] Specifically, the system starts from each symptom instance in the case subgraph and searches for all disease class instances pointing to that symptom in the reverse direction of the relationship; at the same time, for risk factors such as immune deficiency, stress events, and environmental exposure recorded in the case, it searches for related diseases along the relationship of inducing or being susceptible.
[0050] To avoid missing atypical presentations, the traversal depth can be extended to two or three hops, allowing diseases indirectly related through intermediate symptoms or comorbidities to enter the candidate pool.
[0051] In addition, the system will also search the knowledge graph for historical case instances that are highly similar to the current case in terms of animal species, breed, age, region, etc., extract the disease tags of the final diagnosis of these historical cases, and add them to the candidate pool.
[0052] Step 2.2: For each disease in the initial candidate disease pool, calculate its semantic similarity score with the current case knowledge subgraph, as the basic component of the candidate disease confidence score.
[0053] The calculation of semantic similarity is not a simple Jaccard coefficient or cosine similarity, but a weighted metric based on the knowledge graph ontology hierarchy and path distance. For each candidate disease, the system first extracts its set of typical symptoms, set of high-risk factors, and set of identification points defined in the knowledge graph, and then compares these sets with the actual observation set in the current case subgraph.
[0054] During the comparison process, symptoms that are deeper and more specific in the ontological hierarchy are given higher weight; if a case contains exclusionary symptoms listed in the identification criteria of a certain disease, the similarity score of that disease will be significantly reduced.
[0055] At the same time, the system will also calculate the maximum common subgraph size between the case subgraph and the disease definition subgraph, and normalize it as a supplement to the structural similarity.
[0056] Step 2.3: Based on the semantic similarity score, add the case data credibility correction factor and the knowledge timeliness decay factor to generate a comprehensive confidence score for each candidate disease.
[0057] The case data confidence correction factor is derived from the confidence scores of each modality semantic segment in step one and their conflict markers during ontology alignment. If the supporting symptoms of a candidate disease come from a low-confidence modality, or if the symptom is marked as conflict unresolved during cross-modal alignment, its corresponding similarity score will be multiplied by a penalty coefficient less than 1.
[0058] The knowledge timeliness decay factor is used to reflect the dynamic changes in disease prevalence characteristics: if the historical incidence of a disease is extremely low in the current season or region, or if its strain information is marked as eradicated or unseen for many years in the knowledge graph, then even if the symptom match is high, its confidence score will be moderately lowered. The formula for calculating the overall confidence score is as follows: ; in, This represents the overall confidence level of candidate disease d, with a value range of [0,1] and is dimensionless. The semantic similarity score is dimensionless. The confidence correction factor for case data is derived from the weighted average of the confidence levels of each modality in step one, and is dimensionless. It is a knowledge timeliness decay factor, derived by normalizing the historical occurrence frequency of diseases in the current spatiotemporal context, and is dimensionless. and Let be the weighting coefficient, satisfying ,and The confidence level is used to balance the relative importance of data quality and prior knowledge; both are dimensionless constants. This formula ensures that the confidence level not only reflects static knowledge matching but also dynamically adapts to fluctuations in data quality and changes in epidemiological context, making the ranking results closer to real-world clinical situations.
[0059] Step 2.4: Sort the candidate disease list in descending order according to the comprehensive confidence score, select the top N diseases with confidence scores higher than the preset threshold as the final candidate disease set, and generate a list of identification points and suggested verification items for each selected disease.
[0060] The preset threshold balances recall and precision, using an empirical value between 0.3 and 0.5, which can also be dynamically adjusted according to animal species and season.
[0061] For each selected candidate disease, the system extracts differential diagnosis rules between it and other high-confidence candidate diseases from the knowledge graph and generates key identification points in natural language form, such as noting the difference between foot-and-mouth disease and vesicular stomatitis, where the former has more significant lesions in the feet and the latter has more prominent oral lesions.
[0062] At the same time, the system will also list suggested on-site verification items and laboratory testing items based on the diagnosis path definition of the disease in the knowledge graph. This information will directly guide subsequent risk assessment and resource allocation.
[0063] Step 3: Quantify and classify the comprehensive risk score by integrating knowledge graph rules with real-time case characteristics; After outputting the candidate disease set and its confidence level in step two, this step further combines the disease hazard level, transmission characteristics, population impact patterns, and real-time characteristics of current cases stored in the knowledge graph to conduct a multi-dimensional quantitative assessment of the overall risk level of the cases. This assessment is no longer limited to the diagnostic probability of a single disease, but combines diagnostic uncertainty with public health consequences to generate a comprehensive risk score that can directly drive tiered response and resource allocation. The calculation of this score deeply integrates the structured rules in the knowledge graph with the dynamic attributes of case instances, ensuring that the risk assessment is both knowledge-based and realistically relevant, specifically including: Step 3.1: Extract the inherent risk attributes of each disease from the candidate disease set, including the strength of infectivity, mortality rate, zoonotic probability, statutory reporting level, etc., and calculate the weighted disease risk baseline value by combining it with the corresponding confidence score.
[0064] The knowledge graph predefines multi-dimensional risk attribute labels for each disease. These labels are derived from animal disease prevention and control technical specifications and OIE standards, ensuring their authority and stability. For each candidate disease, the system categorizes its infectivity into four levels (low, medium, high, and extremely high) and maps them to values of 0.2, 0.4, 0.7, and 1.0, respectively. The mortality rate is normalized to the 0-1 range. The presence or absence of zoonotic markers is mapped to 1.0 and 0.0, respectively. The legally notifiable disease level (Category I, Category II, and Category III) is mapped to 1.0, 0.6, and 0.3, respectively, along with their confidence levels. The values are multiplied and summed to obtain the weighted baseline disease risk. This baseline reflects the average level of harm that a case may potentially carry given the current diagnostic uncertainty.
[0065] If the candidate set contains multiple high-confidence high-risk diseases, the baseline value will naturally be higher; if only one low-risk disease has a high confidence level, the baseline value will be correspondingly lower. This step's output represents the first-ever fusion of static disease-level knowledge with dynamic confidence levels at the diagnostic level, laying the foundation for subsequent overlay of case-specific risk factors.
[0066] Step 3.2: Based on the case knowledge subgraph, extract real-time feature indicators that reflect the severity of the current case and its impact on the population, including individual symptom severity scores, population incidence rate, population mortality rate, and disease progression rate, to form a case-specific risk vector.
[0067] Individual symptom severity scores are calculated based on the symptom-severity mapping rules in the knowledge graph. For example, the severity of a combination of high fever, neurological symptoms, and recumbency is much higher than that of mild diarrhea. Population morbidity and mortality rates are extracted directly from breeding record instances or reported texts in the case subgraph. If these are missing, they are conservatively estimated based on the default range for similar diseases in the knowledge graph.
[0068] The rate of disease progression was assessed by comparing the interval between the time of first symptom onset and the time of reporting, and by normalizing this assessment in conjunction with the typical incubation period and acute phase duration of the disease as defined in the knowledge graph. All these indicators were normalized to a 0-1 range, forming a multidimensional case-specific risk vector. This vector captures the unique risk profile of the current case that distinguishes it from the general situation; for example, even for a disease of moderate harm, if it breaks out rapidly in a population and has a high mortality rate, its actual risk should be adjusted upwards.
[0069] Step 3.3: The weighted disease risk baseline value and the extracted case-specific risk vector are nonlinearly fused, and regional historical epidemic background factors are superimposed to generate a comprehensive risk score. The fusion process uses a combination of weighted geometric mean and linear compensation to avoid extreme values in a single dimension dominating the overall score, while retaining sensitivity to high-risk signals.
[0070] The regional historical epidemic background factor is obtained from the regional epidemic time series subgraph of the knowledge graph, reflecting the frequency and trend of similar or related epidemics in the current case's location recently. If the epidemic is in an upward phase, this factor is greater than 1, and vice versa. The formula for calculating the comprehensive risk score is as follows: ; Where R represents the comprehensive risk score, which ranges from [0,1] and is dimensionless; This represents the weighted baseline risk value for diseases, which is dimensionless. This is a case-specific risk indicator, dimensionless; Let be the weight coefficient of the i-th indicator, satisfying Dimensionless; The historical epidemic background factor for the region is dimensionless; The fusion index, with a value range of (0,1), is used to adjust the nonlinear coupling strength between the disease baseline and case characteristics, and is dimensionless. This ensures the final score does not exceed the upper limit of 1. The formula is designed so that even when either the disease baseline or any dimension of the case characteristics is significantly elevated, the overall score remains high, reflecting a conservative principle in risk assessment; simultaneously, through... The parameters avoid the problem of underscoring caused by simple multiplication, enabling reasonable differentiation even in medium-risk situations.
[0071] Step 3.4: Based on the comprehensive risk score R generated in Step 3.3, the cases are divided into four risk levels: low, medium, high, and emergency according to the preset threshold range. The level label, along with the risk score, candidate disease set, and case knowledge subgraph URI, are encapsulated into a risk decision package as the direct basis for resource scheduling.
[0072] The threshold ranges were set with reference to the grading standards of the animal epidemic emergency response plan, and were also localized and calibrated based on the local veterinary resource allocation. Low risk Medium risk. High risk, This is an emergency risk.
[0073] The risk level label is not only a classification result, but also carries corresponding response protocol guidelines. For example, low risk can automatically reply with handling suggestions, medium risk requires the intervention of county-level experts, high risk requires the response of prefecture-level experts, and emergency risk requires the immediate deployment of provincial expert groups.
[0074] Step 4: Hierarchical resource matching and dynamic triage routing based on the expert capability subgraph of the knowledge graph; After generating the risk decision package in step three, which includes a comprehensive risk score, risk level labels, a candidate disease set, and a case knowledge graph URI, this step utilizes a pre-built expert capability subgraph in the knowledge graph to perform hierarchical resource matching and triage routing strictly corresponding to the risk level. This process is not a simple personnel list query, but a complex constraint satisfaction and optimal path search problem on the graph. Its goal is to find the most suitable treatment entity for the current case risk characteristics while satisfying multiple hard constraints such as administrative division, professional field, language ability, and current workload, and to establish a traceable task assignment record, specifically including: Step 4.1: Based on the risk level labels in the risk decision package, activate the corresponding level of candidate disposal subject set in the expert capability subgraph of the knowledge graph to form an initial matching pool.
[0075] The expert capability subgraph is a dedicated projection of the global knowledge graph. Its nodes include the main entities responsible for handling matters, such as veterinary experts at all levels, rural veterinarians, laboratory testers, and administrative coordinators. The edges represent attribute relationships such as affiliation, professional expertise, service radius, language skills, and duty status.
[0076] The system directly indexes the corresponding level node based on the risk level label: low risk activates village and township-level disposal entities, medium risk activates county-level expert nodes, high risk activates prefecture-level expert nodes, and emergency risk activates provincial-level expert group nodes.
[0077] Within this level, the system further filters out obviously mismatched nodes based on basic attributes such as animal species and geographical coordinates in the case knowledge subgraph. For example, it removes experts who are only good at poultry diseases from the candidate pool of swine disease cases, or excludes rural veterinarians whose service radius does not cover the township where the case is located.
[0078] Step 4.2: For each candidate treatment subject in the initial matching pool, calculate its multidimensional matching score with the current case risk decision package. This score comprehensively reflects the professional fit, geographical proximity, language compatibility and load balancing.
[0079] The calculation of professional fit is based on the semantic overlap between the disease tags of the handling subject in the knowledge graph and the candidate disease set output in step two. If the candidate subject's area of expertise accurately covers the high-confidence candidate disease, a high score is obtained; if it only covers the low-confidence candidate or only covers the higher-level concept, the score is reduced proportionally.
[0080] Geographic proximity is calculated by normalizing the road network distance or straight-line distance between the respondent's permanent or real-time location and the case's geographical location; the closer the distance, the higher the score. Language compatibility is particularly important in areas with concentrated ethnic minorities. If the respondent's language skills match the language preferences of the case reporter, they receive full marks; otherwise, they receive basic marks.
[0081] Load balancing score is calculated by querying the knowledge graph and determining the ratio of the number of currently incomplete tasks to the rated capacity of the entity. A lower ratio indicates more available resources and a higher score. The raw scores for all four dimensions are normalized to the 0-1 range.
[0082] Step 4.3: Weight and fuse the multidimensional matching feature vectors to generate the final matching score for each candidate treatment subject, and select the first one in descending order as the target task subject. At the same time, create a task dispatch instance in the knowledge graph and establish an association edge with the case knowledge subgraph.
[0083] The weighted fusion uses a linear weighted summation method, with the weights of each dimension dynamically adjusted according to the risk level: for emergency and high-risk cases, the weight of professional fit is significantly increased, while the weight of geographical proximity is appropriately reduced to ensure priority is given to treatment quality; for low-risk cases, the weights of geographical proximity and load balancing are increased to promote efficient use of resources in nearby locations. The final matching score is calculated as follows: ; in, The final matching score of candidate subject j is dimensionless. This represents the normalized score of subject j on the k-th matching dimension (profession, geography, language, workload), which is dimensionless; This represents the weight coefficient of the k-th dimension, which is a function of the comprehensive risk score R, satisfying... and Dimensionless. This formula, by designing the weights as a function of risk scoring, achieves the effect of adaptive adjustment of the triage strategy according to the risk level, avoiding the applicability deviation of fixed weights in different risk scenarios. The system selects... The highest-ranking candidate entity is selected as the target task entity, and a task dispatch instance node is created in the knowledge graph. This node is connected to the root node of the case knowledge subgraph via a disposal object relationship edge, and simultaneously connected to the selected disposal entity node via an assignment relationship edge. This task instance records metadata such as dispatch time, expected response time, and risk level, serving as the anchor point for all subsequent operation tracking and status updates.
[0084] Step 4.4: Encapsulate the task dispatch instance and its associated case knowledge subgraph URI, candidate disease set, and suggested verification item list into a structured task instruction, push it to the terminal device of the target task subject, and mark the task status as pending response.
[0085] Task instructions are pushed via message queues or push services to ensure reliable delivery. The instructions not only include basic task information but also embed key identification points and suggested verification items generated in step two, providing decision support to the receiving entity upon receiving the task without requiring further research. The task status is initialized as pending response in the knowledge graph, and a timer is started to monitor response timeouts.
[0086] If no data is received within the specified time limit, the system will trigger a rematching process, selecting the second-best candidate subject based on the same matching logic or upgrading the risk level and reactivating a higher-level candidate pool.
[0087] Step 5: Incremental update and status reassessment of the case knowledge subgraph based on the data returned from on-site verification; After task assignment and structured task instructions are pushed to the handling entity's terminal in step four, this step receives multi-source feedback data generated by the handling entity during on-site verification, laboratory testing, or expert review. It incrementally updates the case knowledge subgraph constructed in step one and triggers a recalculation of the comprehensive risk score and a reassessment of the case status based on the updated data. This process breaks down the disconnect between online consultation and offline handling in traditional diagnostic and treatment systems. By continuously injecting real-time data into the knowledge graph, the digital representation of cases remains synchronized with the real-world state, thus supporting dynamic adjustments to handling strategies and warning levels. Specifically, this includes: Step 5.1: Receive the on-site verification data, laboratory test results, or expert review opinions transmitted back by the disposal entity through the terminal device, perform structured analysis and source credibility labeling on them, and generate incremental knowledge fragments to be integrated.
[0088] The returned data can take various forms, including supplementary photos taken on-site, completed standardized verification forms, PDF test reports issued by the laboratory, and handwritten diagnostic opinions from experts.
[0089] The system calls the corresponding parsing unit based on the data type: directly mapping form data to attribute value triples; generating image semantic fragments from the visual feature extraction logic in step one of supplementary photo reuse; extracting key fields such as test items, result values, and judgment conclusions from the test report through the document understanding component; and extracting semantic units such as diagnostic conclusions, treatment suggestions, and excluded diseases from expert opinions through the natural language understanding component.
[0090] Each incremental knowledge fragment is accompanied by a source identifier (such as the ID of the specific entity handling the matter, the laboratory qualification number), a collection timestamp, and a source credibility label (such as the credibility of official laboratory results being higher than that of on-site visual estimation).
[0091] Step 5.2: Align, detect conflicts, and selectively fuse the incremental knowledge fragment with the current version of the case knowledge subgraph constructed in Step 1 to generate a new version of the case knowledge subgraph and record the version change log. The alignment process reuses the ontology mapping logic in Step 1 to locate the entities and attributes in the incremental fragment to the corresponding nodes in the existing subgraph.
[0092] When incremental data and existing data give different values for the same attribute, the system arbitrates based on source credibility labels and data type priority: laboratory quantitative test results take precedence over on-site qualitative estimates, reports issued by official institutions take precedence over personal subjective judgments, and recent data takes precedence over long-term data. If the conflict cannot be automatically arbitrated, both values are retained and marked as pending manual review.
[0093] After fusion, the system assigns a new version number to the case knowledge subgraph and writes the set of newly added, modified, and deleted triples, along with conflict markers and arbitration criteria, into the version change log. This log not only supports historical retrospection and auditing but also provides a basis for selecting high-quality training samples for subsequent knowledge feedback.
[0094] Step 5.3: Based on the new version of the case knowledge subgraph, re-execute the candidate disease confidence calculation and comprehensive risk score quantification process in Steps 2 and 3 to generate an updated candidate disease set and comprehensive risk score, and compare it with the previous version to identify significant changes.
[0095] The recalculation process completely reuses the logic and formulas from the previous section, but the input data has been replaced with the new version of the subgraph content. If the incremental data includes a confirmed laboratory result, the confidence level of the corresponding disease will be forcibly increased to close to 1.0, while the confidence levels of other candidate diseases will be lowered accordingly; if on-site verification excludes a certain symptom, the confidence level of the disease relying on that symptom will be reduced.
[0096] The recalculation of the comprehensive risk score also reflects these changes, such as a decrease in the score after a diagnosis of a low-risk disease, or an increase in the score after discovering that the group mortality rate is much higher than initially reported. The system compares the confidence lists of the old and new versions with the risk scores item by item, identifies items whose changes exceed a preset threshold, and marks them as significant changes.
[0097] Step 5.4: Based on the significant changes in the reassessment results, update the status labels and treatment protocol guidelines of the case knowledge subgraph, and encapsulate the updated status labels, risk scores, confirmed case information (if any), and the new version URI of the subgraph into a status update event as an input signal for regional early warning.
[0098] The status label updates follow predefined status transition rules: if a disease is confirmed and the risk score remains above the threshold, the status changes from suspected to confirmed - under treatment; if the risk score drops below the threshold, the status changes to low risk - under observation; if the confidence level of all candidate diseases is below the threshold and no new evidence is available, the status changes to pending supplementary information. Treatment protocol guidelines are also adjusted accordingly, such as automatically associating the standard treatment plan and reporting process for the disease after confirmation.
[0099] The status update event, as a lightweight message object, carries only the key fields required in step six, avoiding the bandwidth overhead of transmitting the entire subgraph. This event is published to the internal message bus and subscribed to and consumed by the alerting component in step six.
[0100] Step Six: Generate spatial reasoning and regional epidemic risk early warning based on the knowledge graph of the spatiotemporal distribution of confirmed cases; After step five outputs a status update event containing confirmed case information, updated risk scores, and status labels, this step utilizes the spatial reasoning capabilities and temporal aggregation logic of knowledge graphs to analyze the spatiotemporal distribution patterns of multiple cases within the region, generating a risk warning signal reflecting the overall epidemic situation in the region. Specifically, this includes: Step 6.1: Listen to the status update event stream published in Step 5, filter out events with status tags of confirmed or high-risk - under treatment, extract their case geographical location, confirmed disease type, risk score and timestamp, and write them into the regional epidemic time series subgraph of the knowledge graph as new observation points.
[0101] The regional epidemic time series subgraph is a projection in the global knowledge graph specifically used for storing and analyzing regional epidemic dynamics. Its nodes include geographic grid units, time windows, disease types, warning levels, etc., while the edges express spatiotemporal relationships such as occurrence, proximity, and evolution.
[0102] Whenever an event that meets the criteria is received, the system creates an observation edge between the corresponding time window node and the geographic grid node, and attaches the disease type, risk score, etc., as edge attributes. If there is already an observation edge for the same disease in the same grid within the same time window, its aggregated statistics (such as the number of cases and the average risk score) are updated.
[0103] Step 6.2: Using the newly written observation point as the center, perform a spatiotemporal neighborhood search in the regional epidemic time series subgraph of the knowledge graph to obtain all similar disease observation points within the specified time window and spatial radius, and calculate the spatiotemporal clustering index.
[0104] Spatiotemporal neighborhood retrieval utilizes pre-built geospatial and time interval indexes in the map to efficiently locate neighboring nodes. The spatial radius is dynamically set based on the disease transmission characteristics; for example, the radius for respiratory infectious diseases is larger than that for contact infectious diseases. The time window is set based on the disease incubation period and reporting delay, typically 7 or 14 days.
[0105] The calculation of the spatiotemporal clustering index comprehensively considers three dimensions: the number of cases within the neighborhood, the spatial compactness, and the temporal concentration. Spatial compactness is calculated by the ratio of the average distance from all observation points within the neighborhood to the center point to the maximum possible distance; the smaller the ratio, the more concentrated the distribution. Temporal concentration is calculated by the ratio of the standard deviation of the observation point timestamps to the window length; the smaller the standard deviation, the more synchronous the outbreak.
[0106] Step 6.3: Integrate the spatiotemporal clustering index with auxiliary indicators such as the average risk score of cases in the neighborhood, the proportion of positive laboratory results, and the case growth rate to generate a regional early warning score, and classify the early warning level according to the preset threshold.
[0107] The laboratory positive rate is derived from the aggregation of laboratory test results from the same region during the same period in the regional epidemic time series sub-plot, reflecting the proportion of confirmed diagnoses; the case growth rate is calculated by comparing the number of cases in the current time window with the number of cases in the previous window, reflecting the evolution trend of the epidemic.
[0108] The regional early warning score is calculated using a weighted summation method, with the weights of each indicator dynamically configured based on the disease type and seasonal context. The early warning levels are classified according to the standards of the emergency response plan for major animal epidemics, into four levels: blue (general), yellow (relatively severe), orange (serious), and red (extremely severe).
[0109] The output of this step is a regional early warning score with clearly defined level labels. It condenses complex spatiotemporal patterns into an actionable signal, providing core parameters for subsequent report generation. Notably, this score is generated entirely based on structured data and inference rules within a knowledge graph, without relying on any external black-box models, ensuring the interpretability and auditability of the early warning logic.
[0110] Step 6.4: Based on the regional early warning score and level label, extract the response suggestion template that matches the level from the knowledge graph, and combine it with information such as the current dominant disease, scope of impact, and trend judgment to generate a complete regional early warning report. Store the report persistently and push it to the decision-making terminal of the competent department.
[0111] The response suggestion templates are pre-stored in the emergency management ontology of the knowledge graph. Each warning level corresponds to a set of standardized suggested measures. For example, a blue warning suggests strengthening monitoring and information reporting, a yellow warning suggests activating emergency reserve teams and allocating materials, an orange warning suggests implementing epidemic area lockdown and culling, and a red warning suggests requesting the higher-level government to activate the emergency response.
[0112] When filling in the template, the system also extracts background data such as the main circulating strains, susceptible animal populations, and vaccine coverage rates in the region from the regional epidemic time series sub-map, making the report more targeted and actionable. The generated report is stored in structured document format and pushed to the decision-making terminals of the provincial, municipal, and county-level competent authorities through a secure channel, while also sending copies to relevant technical support units.
[0113] The above description is merely an exemplary embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural transformations made using the contents of the present invention specification and drawings under the technical concept of the present invention, or direct / indirect applications in other related technical fields, are included within the patent protection scope of the present invention.
Claims
1. A method for the sub-diagnosis of animal diseases based on a multimodal large model architecture, characterized in that, include: Multimodal raw data consisting of symptom description text data and lesion site image data is mapped to the knowledge graph ontology layer to generate a case knowledge subgraph; Based on the aforementioned case knowledge subgraph, semantic retrieval and rule-based reasoning are performed to generate a candidate disease set and confidence level. By integrating the candidate disease set, confidence level, and case characteristics, a comprehensive risk score is quantified and risk levels are classified. Based on the risk level, the relevant entity is matched in the expert capability subgraph of the knowledge graph to generate a task dispatch instance; The system receives on-site verification data and incrementally updates the case knowledge subgraph, triggering a state reassessment and generating a state update event. The spatiotemporal distribution of confirmed cases in the aforementioned status update events is aggregated, and spatial reasoning is performed to generate regional early warning reports.
2. The animal disease triage and diagnosis method based on a multimodal large model architecture according to claim 1, characterized in that, Multimodal raw data, consisting of symptom description text data and lesion site image data, is mapped to the knowledge graph ontology layer to generate a case knowledge subgraph, including: Modal separation and metadata extraction are performed on the symptom description text data and lesion site image data to generate a modal data stream; Based on the modal data stream, text semantic segments and image semantic segments are extracted respectively to form a set of semantic segments for each modality; Align the sets of semantic fragments of each modality with the knowledge graph ontology mode layer to generate a set of normalized triples; Based on the normalized triple set, a case knowledge subgraph is generated and the root node Uniform Resource Identifier is returned.
3. The animal disease triage and diagnosis method based on a multimodal large model architecture according to claim 2, characterized in that, Based on the aforementioned case knowledge subgraph, semantic retrieval and rule-based reasoning are performed to generate a candidate disease set and confidence scores, including: Starting from the root node Uniform Resource Identifier, traverse the knowledge graph to obtain an initial candidate disease pool; For each disease in the initial candidate disease pool, a semantic similarity score with the case knowledge subgraph is calculated; The case data credibility correction factor and the knowledge timeliness decay factor are superimposed on the semantic similarity score to generate a comprehensive confidence score; Based on the comprehensive confidence score, a set of candidate diseases is extracted and accompanied by a list of key identification points and suggested verification items.
4. The animal disease triage and diagnosis method based on a multimodal large model architecture according to claim 3, characterized in that, By integrating the aforementioned candidate disease set, confidence levels, and case characteristics, a comprehensive risk score is quantified and risk levels are classified, including: The inherent risk attributes are extracted from the candidate disease set and combined with the confidence level to calculate the weighted disease risk baseline value; Based on the aforementioned case knowledge subgraph, real-time feature indicators are extracted to form a case-specific risk vector; The weighted disease risk baseline value is non-linearly fused with the case-specific risk vector and superimposed with regional historical epidemic background factors to generate a comprehensive risk score; The risk levels are classified based on the comprehensive risk score and packaged into a risk decision package.
5. The animal disease triage and diagnosis method based on a multimodal large model architecture according to claim 4, characterized in that, Based on the aforementioned risk level, the relevant entity is matched within the expert capability subgraph of the knowledge graph to generate a task assignment instance, including: Activate the corresponding set of candidate disposal subjects based on the risk level labels in the risk decision package; For each entity in the candidate disposal entity set, a multi-dimensional matching score is calculated, consisting of professional suitability, geographical proximity, language compatibility, and load balancing. The multi-dimensional matching scores are weighted and fused to generate the final matching score, and the first and second subjects are selected to create a task dispatch instance. The task dispatch instance and associated information are encapsulated into a structured task instruction and pushed to the target task's main terminal.
6. The animal disease triage and diagnosis method based on a multimodal large model architecture according to claim 5, characterized in that, The system receives on-site verification data and incrementally updates the case knowledge subgraph, triggering a state reassessment and generating a state update event, including: Receive on-site verification data and perform structured parsing and source credibility labeling to generate incremental knowledge fragments; Align the incremental knowledge fragments with the current version of the case knowledge subgraph, perform conflict arbitration based on source credibility tags and data type priority, merge them to generate a new version, and record the change log. Based on the new version of the case knowledge subgraph, the confidence scores and comprehensive risk scores of candidate diseases are recalculated and significant changes are identified. The status label and handling protocol guidelines are updated based on the significant changes and encapsulated as a status update event.
7. The animal disease triage and diagnosis method based on a multimodal large model architecture according to claim 6, characterized in that, The spatiotemporal distribution of confirmed cases in the aforementioned status update events is aggregated, and spatial reasoning is performed to generate regional early warning reports, including: Filter out confirmed or high-risk events from the status update events and write them into the regional epidemic time series subgraph as observation points; Perform spatiotemporal neighborhood retrieval and calculate the spatiotemporal clustering index centered on the observation point; The spatiotemporal clustering index is fused with auxiliary indicators such as the laboratory positive rate and the case growth rate to generate a regional early warning score and classify the early warning levels. Extract response suggestion templates based on the aforementioned warning level and fill them in to generate a regional warning report.
8. The animal disease triage and diagnosis method based on a multimodal large model architecture according to claim 3, characterized in that, The case data credibility correction factor and knowledge timeliness decay factor are superimposed on the semantic similarity score to generate a comprehensive confidence score, including: The confidence correction factor for case data was determined based on the confidence-weighted average of text semantic segments and image semantic segments; The knowledge timeliness decay factor is determined by normalizing the historical occurrence frequency of diseases in the current spatiotemporal context. The semantic similarity score, the case data credibility correction factor, and the knowledge timeliness decay factor are linearly combined according to preset weights to generate a comprehensive confidence score.
9. The animal disease triage and diagnosis method based on a multimodal large model architecture according to claim 5, characterized in that, The multidimensional matching scores are weighted and fused to generate the final matching score, including: The weighting coefficients of professional fit, geographical proximity, language compatibility, and load balancing are dynamically adjusted based on the comprehensive risk score. The final matching score is generated by multiplying the scores of each dimension in the multidimensional matching score with the corresponding weight coefficients and then summing the results.
10. An animal disease sub-diagnosis system based on a multimodal large model architecture for implementing the method as described in any one of claims 1-9, characterized in that, include: The data mapping unit is used to map multimodal raw data, consisting of symptom description text data and lesion site image data, to the knowledge graph ontology layer to generate a case knowledge subgraph. The semantic retrieval unit is used to perform semantic retrieval and rule reasoning based on the case knowledge subgraph to generate a candidate disease set and confidence level. The risk quantification unit is used to integrate the candidate disease set, confidence level, and case characteristics to quantify the comprehensive risk score and classify the risk level. The resource matching unit is used to match the handling subject in the knowledge graph expert capability subgraph according to the risk level and generate a task dispatch instance; The incremental update unit is used to receive on-site verification data to incrementally update the case knowledge subgraph, trigger state reassessment, and generate a state update event. The early warning generation unit is used to aggregate the spatiotemporal distribution of confirmed cases in the status update events and perform spatial reasoning to generate regional early warning reports.