A state-assisted interpretation method and device based on a multi-layer evidence chain

CN122390092BActive Publication Date: 2026-09-08ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610873956.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-17
Publication Date
2026-09-08
Estimated Expiration
2046-06-17

AI Technical Summary

Technical Problem

[0006]本申请的主要目的在于提供一种基于多层证据链的状态辅助解释方法及装置,用于解决现有大语言模型在多模态生理数据解读中存在的容易过度解读、缺乏跨模态权重仲裁、缺失处理采取静默跳过或忽略处理以及结论无法溯源的技术问题

Benefits of technology

[0023] This application constructs a structured semantic feature combination by performing multi-dimensional semantic feature transformation on the physiological characteristic values ​​of the target individual in advance, and uses the correlation between different modalities to perform forced cross-modal evidence weight integration and logical arbitration before inputting it into the model. This fundamentally curbs the "illusion" phenomenon in state evaluation of generative language processing models and the tendency to over-infer single modality indicators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122390092B_ABST
    Figure CN122390092B_ABST
Patent Text Reader

Abstract

The application discloses a state auxiliary explanation method and device based on a multi-layer evidence chain. The method comprises the following steps: acquiring multi-modal input data containing physiological characteristic values of a target individual and performing multi-dimensional semantic feature conversion to generate a structured semantic feature combination; based on the data correlation relationship between different modes, cross-modal evidence weight integration and logical arbitration are performed to determine the comprehensive confidence of each mode data and the arbitration state; the structured semantic feature combination is input into a pre-configured generative language processing model to generate an explanation conclusion carrying traceable source data and confidence identification; when the comprehensive confidence of the explanation conclusion is lower than a preset first confidence threshold, a state supplementary evaluation instruction is attached to the output end. The application can suppress large model hallucinations, establish a two-way binding traceability mechanism between the conclusion and the underlying data, and improve system security and reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of artificial intelligence and multimodal data processing, and more specifically, to a state-assisted interpretation method and apparatus based on a multi-layered chain of evidence. Background Technology

[0002] In modern medical and health, psychological assessment, and human-computer interaction systems, accurately assessing the internal state of a target individual (such as emotional and cognitive states, fatigue levels, and stress levels) has become a crucial research direction. To improve the accuracy of assessments, modern systems are gradually evolving from single-modality to multimodal data fusion. Among these, physiological characteristics, including biomarkers such as inflammatory markers, metabolite concentrations, and gene expression, serve as hard evidence objectively reflecting an individual's underlying physiological mechanisms and have become an indispensable core auxiliary dimension in multimodal assessments.

[0003] Current practices typically input unstructured raw numerical values ​​directly as cue words into the model. Large language models, lacking built-in knowledge of medical or physiological baseline standards, exhibit instability in their understanding of numerical reference ranges. This often leads to "over-interpretation," where a single minor numerical abnormality is easily inferred as strong causal evidence of serious organic disease, ignoring differences in individual baseline characteristics and levels of evidence.

[0004] Existing multimodal systems fail to establish a dynamic and rigorous evidence weighting mechanism when performing cross-modal integration. Under the current architecture, physiological biomarker data, subjective performance texts, and scale scores are often treated as equivalent evidence. When data from different modalities contradict each other (i.e., modal conflict), the system lacks reasonable arbitration logic to reduce the weight of weak evidence. Furthermore, when data from a specific modality is missing, existing systems typically handle it by silently skipping or ignoring it, failing to explicitly convey "insufficient evidence" to the system state and reducing the robustness of the evaluation results.

[0005] Currently, most explanatory conclusions generated by large models exhibit typical "black box" characteristics. These model-generated conclusions not only lack corresponding confidence metrics but also fail to establish a precise two-way traceability link with the original multimodal data sources. Professional systems, faced with evaluation reports output by the models, cannot verify the model's reasoning basis and the specific underlying data it references line by line, resulting in poor system auditability. This black box model limits the engineering implementation of AI-assisted explanation technology in serious scenarios with high requirements for security and interpretability. Therefore, how to construct a rigorous multi-layered evidence chain mechanism, semanticize physiological characteristic values, and achieve logical arbitration of cross-modal evidence and comprehensive traceability of conclusions has become a pressing technical challenge in this field. Summary of the Invention

[0006] The main purpose of this application is to provide a state-assisted interpretation method and device based on a multi-layered evidence chain to solve the technical problems of existing large language models in interpreting multimodal physiological data, such as easy over-interpretation, lack of cross-modal weight arbitration, silent skipping or ignoring of missing data, and inability to trace the source of conclusions.

[0007] To achieve the above objectives, the first aspect of this application provides a state-assisted interpretation method based on a multi-layered chain of evidence, comprising: acquiring multimodal input data containing physiological characteristic values ​​of a target individual, and performing multi-dimensional semantic feature transformation on the physiological characteristic values ​​to generate a structured semantic feature combination; based on the data association relationship between different modalities, performing cross-modal evidence weight integration and logical arbitration on the structured semantic feature combination to determine the comprehensive confidence level and arbitration state of each modality's data, specifically including: comparing the structured semantic feature combinations corresponding to multiple modalities; when it is detected that the structured semantic feature combinations of multiple modalities point to the same equivalent state, increasing the comprehensive confidence level and arbitration state of the data. The first numerical increment of the confidence level; when a directional conflict is detected in the combination of structured semantic features of multiple modalities, the weight of the modal data with a lower initial evidence strength rating is reduced, and conflict state description information is output, wherein the first evidence strength score of the first modal data and the second evidence strength score of the second modal data that have directional conflict are obtained respectively; when the first evidence strength score is less than the second evidence strength score, a first attenuation coefficient γ is assigned to the first modal data; the feature source node data that generated the directional conflict is extracted, and the conflict state description information is generated based on the feature source node data; the first attenuation coefficient γ is preset. A dynamic penalty function is generated by calculating the semantic distance scalar S between the first modality data and the second modality data through the internal feature space of the generative language processing model; reading the association strength matrix pre-existing in the feature map to obtain the first historical variance parameter corresponding to the first modality data and the second historical variance parameter corresponding to the second modality data; calculating the confidence penalty term λ based on the ratio of the first historical variance parameter to the second historical variance parameter; constructing a first exponential decay base α based on the logarithmic product of the semantic distance scalar S and the sum of the constant 1 and the confidence penalty term λ; and then comparing the second evidence strength score with the first evidence strength score. The difference is substituted into the smoothing activation function, and the output is superimposed on the first exponential decay base α to obtain the final first decay coefficient γ. The first decay coefficient γ is then used to update the global decision weight library. When data loss or isolated evidence of a single modality is detected, a corresponding limited confidence label or insufficient causal decision label is added. The structured semantic features combined by the weight integration and logical arbitration are input into a pre-configured generative language processing model to generate an explanation conclusion carrying traceable source data and confidence label. When the overall confidence of the explanation conclusion is lower than the preset first confidence threshold, a state supplementary evaluation instruction is added to the output.

[0008] Furthermore, a multi-dimensional semantic feature transformation is performed on the physiological feature values ​​to generate a structured semantic feature combination, including: extracting numerical deviation and direction features based on the distribution features of the physiological feature values; mapping the numerical deviation and direction features to equivalent description information of physiological state; determining the initial evidence strength level of the physiological feature values ​​based on a preset evidence grading standard; and extracting the data missing degree of the physiological feature values ​​and recording it as an uncertainty labeling parameter.

[0009] Further, based on the distribution characteristics of the physiological characteristic values, numerical deviation and directional features are extracted, including: calculating a first deviation distance D of the physiological characteristic values ​​from the healthy baseline distribution range; determining an abnormal increment marker or an abnormal decrease marker according to the direction of the first deviation distance D; and combining the first deviation distance D with the abnormal increment marker or the abnormal decrease marker into a feature deviation vector.

[0010] Further, determining the initial evidence strength grade of the physiological feature value based on a preset evidence grading standard includes: extracting a pre-constructed biomedical feature mapping matrix; inputting the first feature identifier corresponding to the physiological feature value into the biomedical feature mapping matrix to obtain the causal association strength coefficient and clinical confirmatory parameter corresponding to the first feature identifier; and outputting the initial evidence strength grade through a multi-level threshold judgment network based on the product of the causal association strength coefficient and the clinical confirmatory parameter, wherein a first-level high-intensity identifier is assigned when the product is greater than a first intensity judgment threshold, a second-level medium-intensity identifier is assigned when the product is between the first intensity judgment threshold and a second intensity judgment threshold, and a third-level low-intensity identifier is assigned when the product is less than the second intensity judgment threshold.

[0011] Furthermore, based on the data association relationships between different modalities, cross-modal evidence weight integration and logical arbitration are performed on the structured semantic feature combinations to determine the comprehensive confidence level and arbitration status of each modality's data. This includes: comparing structured semantic feature combinations corresponding to multiple modalities; when it is detected that structured semantic feature combinations of multiple modalities point to the same equivalent state, increasing the first numerical increment of the comprehensive confidence level; when it is detected that there is a directional conflict in the structured semantic feature combinations of multiple modalities, reducing the weight of the modality data with a lower initial evidence strength rating and outputting conflict state description information; when it is detected that data is missing in a specific modality or that single-modality isolated evidence is presented, adding a corresponding limited confidence level identifier or insufficient causal determination identifier.

[0012] Furthermore, when a directional conflict is detected in the combination of structured semantic features of multiple modalities, the weight of the modal data with a lower initial evidence strength rating is reduced, and conflict state description information is output, including: obtaining the first evidence strength score of the first modal data and the second evidence strength score of the second modal data that have a directional conflict; when the first evidence strength score is less than the second evidence strength score, a first attenuation coefficient γ is assigned to the first modal data; the feature source node data that generated the directional conflict is extracted, and the conflict state description information is generated based on the feature source node data.

[0013] Furthermore, the first attenuation coefficient γ is generated by a preset dynamic penalty function, and the method further includes: calculating the semantic distance scalar S between the first modality data and the second modality data through the internal feature space of the generative language processing model; reading the association strength matrix in the pre-existing feature map to obtain the first historical variance parameter corresponding to the first modality data. The second historical variance parameter corresponding to the second modal data According to the first historical variance parameter Divide by the second historical variance parameter The confidence penalty term λ is calculated based on the ratio; a first exponential decay base α is constructed based on the logarithmic product of the semantic distance scalar S and the constant 1 plus the confidence penalty term λ; the difference between the second evidence strength score and the first evidence strength score is substituted into the smoothing activation function, and the output result is superimposed on the first exponential decay base α to obtain the final first decay coefficient γ, and the global decision weight library is updated using the first decay coefficient γ.

[0014] Furthermore, when data loss in a specific modality is detected or isolated evidence of a single modality is presented, a corresponding limited confidence level identifier or insufficient causal determination identifier is added, including: identifying the missing data dimension in the multimodal input data; when the missing data dimension belongs to the core determination modality set, limiting the overall confidence level to a preset second upper limit value and binding the limited confidence level identifier; when only a single modality supports the preset determination conclusion in the combination of structured semantic features, writing the insufficient causal determination identifier into the data structure of the output result.

[0015] Furthermore, the structured semantic feature combination after weight integration and logical arbitration is input into a pre-configured generative language processing model to generate an explanation conclusion carrying traceable source data and confidence level identifiers, including: concatenating the structured semantic feature combination and the arbitration state into a first prompt word template; parsing the first prompt word template using the generative language processing model; extracting and combining metadata of multiple dimensions based on the parsing results; and encapsulating the metadata of multiple dimensions into a standard output file as the explanation conclusion.

[0016] Furthermore, the metadata of the multiple dimensions includes at least the conclusion text, confidence score, evidence source tracing list, conflict analysis record, and uncertainty assessment description; the evidence source tracing list contains the original collection identifier and timestamp information of the corresponding physiological feature values ​​in the structured semantic feature combination.

[0017] Furthermore, when the overall confidence level of the interpretation conclusion is lower than a preset first confidence level threshold, a supplementary evaluation instruction is added to the output, including: extracting the defect modality identifier that causes the low value in the overall confidence level; matching the pre-stored evaluation completion database to retrieve at least one supplementary detection item scheme corresponding to the defect modality identifier; converting the supplementary detection item scheme into a standardized scheduling instruction; sending the standardized scheduling instruction to the corresponding external acquisition device interface through the message bus, and rendering a highlighted prompt message in a designated area of ​​the interactive interface, requesting secondary verification data that is strongly correlated with the defect modality identifier.

[0018] To achieve the above objectives, a second aspect of this application provides a state-assisted interpretation device based on a multi-layered chain of evidence, comprising: a feature transformation module, used to acquire multimodal input data containing physiological feature values ​​of a target individual, and to perform multi-dimensional semantic feature transformation on the physiological feature values ​​to generate a structured semantic feature combination; and an evidence arbitration module, connected to the feature transformation module, used to perform cross-modal evidence weight integration and logical arbitration on the structured semantic feature combination based on the data association relationship between different modalities, to determine the comprehensive confidence level and arbitration status of each modality's data; the evidence arbitration module is specifically used to: compare the structured semantic feature combinations corresponding to multiple modalities; and when multiple modalities' structured semantic feature combinations are detected... When semantic feature combinations point to the same equivalent state, the first numerical increment of the comprehensive confidence is increased; when directional conflicts are detected in the structured semantic feature combinations of multiple modalities, the weight of the modal data with a lower initial evidence strength rating is reduced, and conflict state description information is output, wherein the first evidence strength score of the first modal data and the second evidence strength score of the second modal data that have directional conflicts are obtained respectively; when the first evidence strength score is less than the second evidence strength score, a first attenuation coefficient γ is assigned to the first modal data; the feature source node data that generated the directional conflict is extracted, and the conflict state description information is generated based on the feature source node data; the first attenuation The coefficient γ is generated by a preset dynamic penalty function. The semantic distance scalar S between the first modality data and the second modality data is calculated using the internal feature space of the generative language processing model. The association strength matrix in the pre-existing feature map is read to obtain the first historical variance parameter corresponding to the first modality data and the second historical variance parameter corresponding to the second modality data. The confidence penalty term λ is calculated based on the ratio of the first historical variance parameter to the second historical variance parameter. A first exponential decay base α is constructed based on the logarithmic product of the semantic distance scalar S, a constant 1, and the confidence penalty term λ. The difference between the second evidence strength score and the first evidence strength score is substituted into the smoothing activation. The function is used to superimpose the output result onto the first exponential decay base α to obtain the final first decay coefficient γ, and the first decay coefficient γ is used to update the global judgment weight library; when data loss of a specific modality is detected or isolated evidence of a single modality is presented, a corresponding limited confidence level label or insufficient causal judgment label is added; a feedback generation module is connected to the evidence arbitration module, which is used to input the structured semantic feature combination after the weight integration and logical arbitration into the pre-configured generative language processing model to generate an explanation conclusion carrying traceable source data and confidence level label, and when the comprehensive confidence level of the explanation conclusion is lower than the preset first confidence level threshold, a state supplementary evaluation instruction is added to the output.

[0019] To achieve the above objectives, a third aspect of this application provides an electronic device, comprising: a memory for storing a computer program; and a processor communicatively connected to the memory, wherein the processor, when executing the computer program, implements the method described in any of the preceding claims.

[0020] To achieve the above objectives, a fourth aspect of this application provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the method described in any of the preceding claims.

[0021] To achieve the above objectives, a fifth aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the method described in any of the preceding claims.

[0022] Compared with existing technologies, the state-assisted interpretation method and apparatus based on a multi-layered chain of evidence proposed in this application have the following advantages:

[0023] This application constructs a structured semantic feature combination by performing multi-dimensional semantic feature transformation on the physiological characteristic values ​​of the target individual in advance, and uses the correlation between different modalities to perform forced cross-modal evidence weight integration and logical arbitration before inputting it into the model. This fundamentally curbs the "illusion" phenomenon in state evaluation of generative language processing models and the tendency to over-infer single modality indicators.

[0024] Furthermore, this application mandates that the model's output interpretation conclusions must include traceable source data and confidence level indicators, and triggers targeted supplementary evaluation instructions when the overall confidence level is insufficient. This design establishes a two-way binding traceability mechanism between each interpretation conclusion and the underlying data source, providing business stakeholders with transparent operational space for item-by-item verification and professional auditing, thereby enhancing the security, credibility, and engineering implementation value of the AI-assisted evaluation system in serious medical or psychological application environments. Attached Figure Description

[0025] Figure 1 This is a flowchart illustrating a state-assisted interpretation method based on a multi-layered chain of evidence provided in an embodiment of this application.

[0026] Figure 2 This is a schematic diagram of the logical flow of the multi-dimensional semantic feature transformation process provided in the embodiments of this application.

[0027] Figure 3 This is a flowchart illustrating the cross-modal evidence weighting integration and logical arbitration process provided in the embodiments of this application.

[0028] Figure 4 This is a logical structure block diagram of a state-assisted interpretation device based on a multi-layered chain of evidence provided in an embodiment of this application.

[0029] Figure 5 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application.

[0030] Explanation of reference numerals in the attached figures:

[0031] In the diagram: 401 Feature conversion module, 402 Evidence arbitration module, 403 Generation feedback module, 501 Processor, 502 Memory, 503 Communication interface, 504 System bus. Detailed Implementation

[0032] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0033] With the evolution of artificial intelligence technology, large language models have acquired the ability to process multimodal data. However, in strong logic applications such as emotion and cognitive state assisted assessment, when the input includes specific biomarkers (such as inflammatory indicators, metabolite expression, etc.) and cross-modal data such as behavioral text and scale scores, general models often exhibit uncontrollable interpretability. To address these challenges, this application provides a novel state-assisted interpretation framework based on a multi-layered evidence chain. This framework can construct an evidence chain with logical gradients and traceability attributes before data is input into a large model, thereby ensuring the reliability and auditability of the output conclusions.

[0034] See Figure 1 , Figure 1 This is a flowchart illustrating a state-assisted interpretation method based on a multi-layered chain of evidence provided in an embodiment of this application. Figure 1 As shown, the method provided in this application mainly includes the following steps:

[0035] In the process of acquiring multimodal input data and performing semantic feature transformation S101, multimodal input data containing the physiological feature values ​​of the target individual is acquired, and multidimensional semantic feature transformation is performed on the physiological feature values ​​to generate a structured semantic feature combination.

[0036] In this embodiment, the multimodal input data mainly constitutes the state semantic package of the target individual, which includes data in at least two dimensions: the mechanism axis and the modal evidence axis. The target individual can be a user who needs to assess their health or psychological state. The multimodal input data not only includes traditional subjective survey scale ratings, textual self-reports, or voice emotional features, but more importantly, it includes objective physiological characteristic values. The physiological characteristic values ​​refer to objective physical or biochemical indicators that can reflect the underlying physiological state of the body, collected through medical testing or wearable devices. It should be noted that although this embodiment uses biomarkers such as inflammatory markers, metabolite concentrations, or gene expression levels as typical representatives of physiological characteristic values, the physiological characteristic values ​​in this application are not limited to these. In other embodiments, the physiological characteristic values ​​can also be heart rate variability, skin conductivity, or brainwave frequency band energy values, as long as they can objectively reflect the functional state of the target individual.

[0037] In practice, the system calls an external multimodal data acquisition interface to receive raw physiological feature values. Since the raw values ​​themselves lack medical semantic boundaries for the language model, multi-dimensional semantic feature transformation must be performed. This transformation process is strictly divided into four mandatory dimensions, L1 to L4, which are to extract the abnormality direction and degree, state equivalence description, evidence strength grading, and uncertainty source in sequence.

[0038] Combination Figure 2 , Figure 2 This is a schematic diagram of the logical flow of the multi-dimensional semantic feature transformation process provided in the embodiments of this application. For example... Figure 2 As shown, the multi-dimensional semantic feature transformation includes the following specific sub-steps:

[0039] In extracting numerical deviation and directional features S201, numerical deviation and directional features are extracted based on the distribution characteristics of the physiological feature values. To ensure the objectivity of numerical comparison, the system extracts the distribution baseline of this physiological feature in a healthy population. The system calculates the first deviation distance D of the physiological feature value from the healthy baseline distribution interval using the following formula:

[0040]

[0041] in, V actual This represents the physiological characteristic values ​​actually collected from the target individual; V baseline This represents the mean of the baseline distribution of the healthy population corresponding to this feature; This represents the standard deviation of the distribution of healthy individuals. In this embodiment, the first deviation distance D typically ranges from 0 to 5, and is used to characterize the severity of the abnormality. Following this, the system... determine an abnormal increment marker or an abnormal decrement marker based on the positive and negative directions. Finally, the calculated first deviation distance D and the above abnormal direction marker are combined into a feature deviation vector, which can accurately anchor the relative spatial position of the original data in the overall distribution.

[0042] In the physiological state equivalent description information mapping step S202, the numerical deviation degree and direction features are mapped into physiological state equivalent description information. The system is pre-configured with a natural language feature dictionary, which discretizes the feature deviation vector according to interval division and matches corresponding texts. For example, when the first deviation distance D≤1.0 and it is an abnormal increment, it is mapped to "at the upper limit of normal"; when 1.0<D≤3.0, it is mapped to "mildly elevated"; when D>3.0, it is mapped to "abnormally high expression". Through this step, pure digital signals can be equivalently translated into natural language qualitative descriptions that are easy for the model to understand.

[0043] In the initial evidence intensity grading step S203, the initial evidence intensity grading of the physiological characteristic value is determined based on a preset evidence grading standard. In this step, a biomedical feature mapping matrix pre-constructed by a medical expert system is extracted. The biomedical feature mapping matrix is a data structure pre-constructed by integrating clinical guidelines and medical literature, and its dimensional features at least include feature identification nodes, target state nodes, as well as causal association intensity weight values and clinical confirmation degree reference values connecting each node. The system first generates a first feature identifier corresponding to the current physiological characteristic value according to the physiological state equivalent description information generated in the foregoing steps, then inputs the first feature identifier corresponding to the current physiological characteristic value into the matrix, and queries to obtain the corresponding causal association intensity coefficient C and clinical confirmation degree parameter R. Based on the product result P of the causal association intensity coefficient and the clinical confirmation degree parameter, where P=C×R, the initial evidence intensity grading is output through a multi-level threshold judgment network. The specific logic is: when the product result P is greater than a preset first intensity judgment threshold, the network judges that the indicator has a strong causal relationship with the target state and assigns a first-level high-intensity identifier; when the product result P is between the first intensity judgment threshold and a second intensity judgment threshold, a second-level medium-intensity identifier is assigned; when the product result P is less than the second intensity judgment threshold, it is regarded as a marginally correlated signal and a third-level low-intensity identifier is assigned.

[0044] In the uncertainty labeling parameter S204, the degree of data missing for the physiological feature values ​​is extracted and recorded as uncertainty labeling parameters. The system identifies quality issues in the data acquisition process, such as signal-to-noise ratio and whether some historical samples are missing, and formats these issues into uncertainty labels. After completing the above four-layer transformation, the system packages the information from these four dimensions to generate a structured semantic feature combination that covers the underlying numerical details and the top-level semantic logic. By performing semantic transformation from L1 to L4, the raw biological data can be given clear clinical gradients and causal weights, thereby effectively preventing downstream interpretive models from over-interpreting and exaggerating single weak indicators due to a lack of reference.

[0045] In the cross-modal evidence weight integration and logical arbitration S102, based on the data association relationship between different modalities, cross-modal evidence weight integration and logical arbitration are performed on the structured semantic feature combination to determine the comprehensive confidence level and arbitration status of each modality's data.

[0046] At this point, the data from each modality has been semantically encoded. However, data from different modalities may corroborate each other, contradict each other, or be missing. Therefore, strict arbitration rules must be established for integration. Figure 3 , Figure 3 This is a flowchart illustrating the cross-modal evidence weighting integration and logical arbitration process provided in this application embodiment. The system primarily compares the structured semantic feature combinations corresponding to multiple modalities in parallel through four main logical branches.

[0047] In the equivalent state consistency determination S301, when multiple modalities' structured semantic features are detected to point to the same equivalent state, the system determines that the evidence chain forms a closed loop. At this time, the first numerical increment of the comprehensive confidence score is increased. For example, when physiological characteristics and scale characteristics jointly indicate a "high stress state," the system adds the first numerical increment to the default comprehensive confidence score, improving the reliability of the final conclusion.

[0048] In the direction conflict reduction judgment S302, when a direction conflict is detected in the combination of structured semantic features of multiple modalities, the system forcibly performs weight reduction processing to prevent the large model from outputting self-contradictory analysis results. Specifically, the first evidence strength score of the first modality data and the second evidence strength score of the second modality data that have the direction conflict are obtained respectively; when the first evidence strength score is less than the second evidence strength score, the first modality data is assigned a first attenuation coefficient γ for weight penalty. At the same time, the system background extracts the feature source node data that generated the direction conflict and generates structured conflict state description information based on the feature source node data, which is then displayed with the report. By suppressing low-strength evidence when a conflict occurs, the uniqueness and clear distinction between primary and secondary aspects of the final interpretation logic can be ensured.

[0049] To make the generation of the first attenuation coefficient γ smoother and more statistically consistent, in one optional implementation, this coefficient is generated by a preset dynamic penalty function. The system calculates the semantic distance scalar S between the first modality data and the second modality data using the internal feature space of the generative language processing model. This semantic distance scalar S is used to quantify the degree of semantic divergence between the two modal contents. The system reads the association strength matrix from the pre-existing feature map to obtain the first historical variance parameter corresponding to the first modality data. The second historical variance parameter corresponding to the second modal data The variance parameter here characterizes the volatility and unreliability of this type of data during historical collection. Based on the first historical variance parameter... Divide by the second historical variance parameter The confidence penalty term λ is calculated based on the ratio. The system constructs a first exponential decay base α based on the logarithmic product of the semantic distance scalar S, the constant 1, and the confidence penalty term λ. The system performs the semantic distance and penalty logarithmic product operation:

[0050]

[0051] The constant 1 is used to prevent negative values ​​in the logarithmic result. After obtaining the decay base, the system substitutes the difference between the second evidence strength score (set as E2) and the first evidence strength score (set as E1) into the smoothing activation function, and superimposes the output result onto the first exponential decay base α to obtain the final first decay coefficient γ. As a specific implementation, the smoothing activation function specifically adopts the Sigmoid function form, and its calculation logic is as follows:

[0052]

[0053] Wherein, γ represents the dynamic decay coefficient that is finally applied to the weights of the low-scoring modes, and its value is preferably constrained to between 0.1 and 0.9; E 1 and E 2 This represents the intensity score. After calculation, the system updates the corresponding weight allocation matrix in the global decision weight library using the first attenuation coefficient γ. By introducing this composite exponential attenuation smoothing action that incorporates historical variance and semantic distance, not only is the abrupt weight changes caused by hard-coded rules avoided, but the weight reduction process also ensures that it accurately reflects the true physical intensity and uncertainty of modal conflicts.

[0054] In the modal data missing constraint S303, when data missing for a specific modality is detected or isolated evidence of a single modality is presented, explicit annotation and compensation must be performed. The system identifies the missing data dimension in the multimodal input data. If the missing data dimension belongs to the core decision modality set, the system forcibly limits the overall confidence level to a preset second upper limit value and simultaneously binds the limited confidence level identifier in the internal bus. This is equivalent to marking a limited confidence level at the system's underlying layer.

[0055] In the single-modal isolated evidence warning S304, during reasoning about a certain state, when only a single modality in the structured semantic feature combination supports the preset judgment conclusion, the system writes the insufficient causal judgment flag into the data structure of the output result, explicitly expressing that it is insufficient to support a strong causal conclusion. Through the explicit arbitration mechanism of missing and isolated evidence, the system can effectively prevent the model from forcibly fabricating evidence when information is incomplete.

[0056] In the generation of the explanation conclusion S103, the structured semantic feature combination after weight integration and logical arbitration is input into a pre-configured generative language processing model to generate an explanation conclusion carrying traceable source data and confidence level indicators. The generative language processing model in this application refers to an artificial intelligence engine with long text generation capabilities, such as a large-scale pre-trained general language model or an expert model fine-tuned for the medical field. In a preferred embodiment, the generative language processing model can also be a language model based on the Transformer architecture, a long short-term memory network, or a dedicated multimodal reasoning model, as long as it can achieve cross-modal parsing and text generation; this application does not impose specific limitations on this.

[0057] The core of this step is to ensure the transparency and traceability of the final report. To constrain the output format of the large model, the system constructs a specific call request through prompt word engineering. The structured semantic features are combined with the arbitration status generated in the previous step to form a first prompt word template. The generative language processing model is used to parse the first prompt word template, strictly following the instructions, and extracting and combining metadata of multiple dimensions based on the parsing results. The metadata of multiple dimensions is encapsulated into a standard output file as the interpretation conclusion.

[0058] According to the system's strong constraint protocol, the metadata of the multiple dimensions includes at least five key elements: conclusion text, confidence level value, evidence source tracing list, conflict analysis record, and uncertainty assessment description.

[0059] The conclusion text is a detailed textual explanation synthesized by the large model based on various weights. The confidence score reflects the overall reliability of this explanation. The evidence source tracing list must include the original collection identifier and timestamp information of the corresponding physiological feature values ​​in the structured semantic feature combination. Based on this, the conflict analysis record includes the first conflict node identifier, the first conflict direction information, and the decision tree execution path when processing the conflict state description information. The uncertainty assessment description includes the first noise estimation feature value and the first probability density distribution information predicted based on the current data coverage. By forcing the large model to output this five-dimensional metadata carrying traceability markers, it ensures that each conclusion can be traced back to the original sensor or laboratory report at the time of collection through the evidence source list, achieving deep two-way traceability. If any doubts arise, reviewers only need to extract the metadata to complete a white-box audit.

[0060] In the state supplementary evaluation instruction S104, when the overall confidence level of the interpretation conclusion is lower than the preset first confidence level threshold, a state supplementary evaluation instruction is added to the output.

[0061] This step is crucial for forming a closed loop of intelligent interaction. The system continuously monitors the overall confidence level of each output conclusion. When the confidence level of a conclusion falls below a preset first confidence threshold, the system enters a low-confidence response subroutine. It extracts the defect modality identifier that causes the low value from the overall confidence level. It matches the pre-stored evaluation completion database and retrieves at least one supplementary detection item scheme corresponding to the defect modality identifier. The supplementary detection item scheme is converted into a standardized scheduling instruction. This standardized scheduling instruction is sent to the corresponding external acquisition device interface via the system's underlying message bus, ready to initiate a new round of acquisition at any time; simultaneously, a highlighted prompt is rendered in a designated area of ​​the interactive interface, requesting secondary verification data strongly correlated with the defect modality identifier. Upon receiving the interface interaction instruction, the system renders and highlights the gap in the evidence chain leading to the low confidence level. By automatically triggering supplementary instructions and UI interaction rendering at low confidence levels, proactive guided completion is achieved, closing the loop of the multimodal assisted evaluation business flow.

[0062] In summary, the method embodiments provided in this application, by employing four layers of mandatory semantic rules, transform the black-box raw numerical values ​​into features with clear causal strength; combined with four cross-modal arbitration logics, it dynamically and reasonably suppresses contradictory evidence and identifies isolated evidence; finally, by leveraging five-dimensional source metadata and a low-confidence feedback loop, it completely eradicates the "illusion" phenomenon in state analysis of large models. This method improves the rigor, security, and interpretability of AI-assisted state interpretation systems in practical applications.

[0063] Based on the same subject matter as the methods described above, this application also provides an apparatus system. See [link to relevant documentation]. Figure 4 , Figure 4 This is a logical structure block diagram of a state-assisted interpretation device based on a multi-layered chain of evidence provided in an embodiment of this application. For example... Figure 4 As shown, the device includes a feature conversion module 401, an evidence arbitration module 402, and a generation feedback module 403 that are interconnected.

[0064] The feature transformation module 401 is mainly configured to acquire multimodal input data containing the physiological feature values ​​of the target individual, and perform multi-dimensional semantic feature transformation on the physiological feature values ​​to generate a structured semantic feature combination. This module integrates a data parser and a multi-layer semantic mapping dictionary, and is responsible for executing the feature reshaping process at layers L1 to L4.

[0065] The evidence arbitration module 402, connected to the feature transformation module 401, is configured to perform cross-modal evidence weight integration and logical arbitration on the structured semantic feature combination based on the data association relationships between different modalities, and determine the comprehensive confidence level and arbitration status of each modality's data. This module embeds a conflict weight suppression algorithm and missing data handling compensation logic.

[0066] A feedback generation module 403, connected to the evidence arbitration module 402, is configured to input the structured semantic feature combination after weight integration and logical arbitration into a pre-configured generative language processing model to generate an interpretation conclusion carrying traceable source data and confidence level indicators. Furthermore, this module includes an instruction scheduling unit, used to append supplementary evaluation instructions to the output and distribute them to the interaction bus when the overall confidence level of the interpretation conclusion is lower than a preset first confidence level threshold.

[0067] The device described in this embodiment, through its modular physical structure, decouples the evidence processing and generation tasks at each stage. The modules work collaboratively to achieve efficient, high-throughput, secure, and reliable multimodal state attribution interpretation.

[0068] Based on the above implementation principles, this application also provides a hardware implementation scheme for an electronic device. See also... Figure 5 , Figure 5 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application. For example... Figure 5 As shown, the electronic device mainly includes a processor 501 and a memory 502. In addition, in order to meet the basic data communication and interaction requirements, it may also include a communication interface 503 and a system bus 504 connecting the above-mentioned components.

[0069] The memory 502, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the state-assisted interpretation method described above in this application embodiment. The processor 501 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 502, thereby implementing the methods provided in the aforementioned embodiments.

[0070] The processor 501 can be a central processing unit, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices. The memory 502 can include high-speed random access memory (RAM) and non-volatile memory, such as at least one disk storage device, flash memory, or other non-volatile solid-state storage device. The communication interface 503 is used to realize data communication transmission between the electronic device and the outside world. It should be noted that the electronic device structure shown in this embodiment is merely exemplary; those skilled in the art can achieve the purposes of this application by using server or terminal devices containing more or fewer components.

[0071] Furthermore, this application also provides a computer-readable storage medium storing computer instructions. When executed by a processor, these computer instructions implement the steps of the state-assisted interpretation method based on a multi-layered chain of evidence provided in the foregoing embodiments. This storage medium can take any suitable form, including but not limited to optical discs, USB flash drives, portable hard drives, and solid-state drives.

[0072] This application also provides a computer program product containing instructions that, when run on a computer or related computing node, causes the computer to execute a state-assisted interpretation method based on a multi-layered chain of evidence as detailed in the foregoing embodiments, thereby covering the core technical solution of this application at the software distribution level.

[0073] The above description is only a preferred embodiment of this application and does not limit the scope of patent protection of this application. Any equivalent structural or procedural changes made based on the content of this application’s specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of this application.

Claims

1. A state-assisted interpretation method based on a multi-layered chain of evidence, comprising: Acquire multimodal input data containing the physiological characteristic values ​​of the target individual, and perform multidimensional semantic feature transformation on the physiological characteristic values ​​to generate a structured semantic feature combination; Based on the data association relationships between different modalities, cross-modal evidence weight integration and logical arbitration are performed on the structured semantic feature combination to determine the comprehensive confidence level and arbitration status of each modal data; Specifically, it includes: Compare the structured semantic feature combinations corresponding to multiple modalities; When a combination of structured semantic features from multiple modalities is detected to point to the same equivalent state, the first numerical increment of the comprehensive confidence is increased. When a directional conflict is detected in the combination of structured semantic features of multiple modalities, the weight of the modal data with a lower initial evidence strength rating is reduced, and conflict state description information is output. Specifically, the first evidence strength score of the first modal data that has a directional conflict and the second evidence strength score of the second modal data are obtained respectively. When the first evidence strength score is less than the second evidence strength score, the first modal data is assigned a first attenuation coefficient γ; Extract the feature source node data that generates the directional conflict, and generate the conflict state description information based on the feature source node data; the first attenuation coefficient γ is generated by a preset dynamic penalty function, and the semantic distance scalar S between the first modality data and the second modality data is calculated through the internal feature space of the generative language processing model; Read the correlation strength matrix from the pre-existing feature map and obtain the first historical variance parameter corresponding to the first modality data. The second historical variance parameter corresponding to the second modal data ; Based on the first historical variance parameter Divide by the second historical variance parameter The confidence penalty term λ is calculated using the ratio. The first exponential decay base α is constructed based on the logarithmic product of the semantic distance scalar S, the constant 1, and the confidence penalty term λ. The difference between the second evidence strength score and the first evidence strength score is substituted into the smoothing activation function, and the output is superimposed on the first exponential decay base α to obtain the final first decay coefficient γ, and the global decision weight library is updated using the first decay coefficient γ. When data loss for a specific modality is detected or isolated evidence for a single modality is presented, a corresponding limited confidence level indicator or insufficient causal determination indicator is added. The structured semantic features, after weight integration and logical arbitration, are input into a pre-configured generative language processing model to generate explanatory conclusions carrying traceable source data and confidence level indicators. When the overall confidence level of the interpretation conclusion is lower than the preset first confidence level threshold, a supplementary evaluation instruction is added to the output.

2. The method as described in claim 1, characterized in that, Perform multi-dimensional semantic feature transformation on the physiological feature values ​​to generate a structured semantic feature combination, including: Numerical deviation and directional features are extracted based on the distribution characteristics of the aforementioned physiological characteristic values; The numerical deviation and directional features are mapped to physiological state equivalent description information; The initial evidence strength grading of the physiological characteristic values ​​is determined based on a preset evidence grading standard. The degree of data missing for the physiological feature values ​​is extracted and recorded as an uncertainty labeling parameter.

3. The method as described in claim 2, characterized in that, Based on the distribution characteristics of the aforementioned physiological characteristics, numerical deviation and directional features are extracted, including: Calculate the first deviation distance D of the physiological characteristic value from the healthy baseline distribution range; Determine the abnormal increment marker or abnormal decrease marker based on the direction of the first deviation distance D; The first deviation distance D is combined with the abnormal increment marker or the abnormal decrement marker to form a feature deviation vector.

4. The method as described in claim 2, characterized in that, The initial evidence strength grading of the physiological characteristic values ​​is determined based on a preset evidence grading standard, including: Extract the pre-constructed biomedical feature mapping matrix; A first feature identifier corresponding to the physiological feature value is generated based on the equivalent description information of the physiological state, and the first feature identifier corresponding to the physiological feature value is input into the biomedical feature mapping matrix to obtain the causal association strength coefficient and clinical confirmatory parameter corresponding to the first feature identifier. Based on the product of the causal association strength coefficient and the clinical confirmatory parameter, the initial evidence strength grade is output through a multi-level threshold determination network. When the product is greater than the first strength determination threshold, a first-level high-intensity label is assigned; when the product is between the first and second strength determination thresholds, a second-level medium-intensity label is assigned; and when the product is less than the second strength determination threshold, a third-level low-intensity label is assigned.

5. The method as described in claim 1, characterized in that, When data loss for a specific modality is detected or isolated evidence of a single modality is presented, a corresponding limited confidence level indicator or insufficient causal determination indicator is added, including: Identify the missing dimensions in the multimodal input data; When the missing data dimension belongs to the core judgment modality set, the comprehensive confidence level is limited to a preset second upper limit value, and the limited confidence level identifier is bound. When only a single modality supports the preset judgment conclusion in the combination of structured semantic features, the insufficient causal judgment identifier is written into the data structure of the output result.

6. The method as described in claim 1, characterized in that, The structured semantic features, after weight integration and logical arbitration, are input into a pre-configured generative language processing model to generate explanatory conclusions carrying traceable source data and confidence level indicators, including: The structured semantic features and the arbitration status are combined to form a first prompt word template; The first prompt word template is parsed using the generative language processing model described above; Metadata with multiple dimensions is extracted and combined based on the parsing results; The metadata from multiple dimensions is encapsulated into a standard output file as the interpretation conclusion.

7. The method as described in claim 6, characterized in that, The metadata, which spans multiple dimensions, includes at least the conclusion text, confidence level values, a list of evidence sources, conflict analysis records, and uncertainty assessment descriptions. The evidence source tracing list includes the original collection identifier and timestamp information of the corresponding physiological feature values ​​in the structured semantic feature combination.

8. The method as described in claim 1, characterized in that, When the overall confidence level of the interpretation conclusion is lower than a preset first confidence threshold, a supplementary evaluation instruction is added to the output, including: Extract the defect mode identifiers that lead to low values ​​from the overall confidence score; Match the pre-stored evaluation and completion database to retrieve at least one supplementary detection item scheme corresponding to the defect modality identifier; The supplementary testing plan is converted into standardized scheduling instructions; The standardized scheduling instructions are sent to the corresponding external acquisition device interface via the message bus, and highlighted prompts are rendered in a designated area of ​​the interactive interface, requiring secondary verification data that is strongly correlated with the defect modality identifier.

9. A state-assisted interpretation device based on a multi-layered chain of evidence, comprising: The feature transformation module is used to acquire multimodal input data containing the physiological feature values ​​of the target individual, and to perform multi-dimensional semantic feature transformation on the physiological feature values ​​to generate a structured semantic feature combination. The evidence arbitration module, connected to the feature transformation module, is used to perform cross-modal evidence weight integration and logical arbitration on the structured semantic feature combination based on the data association relationship between different modalities, and to determine the comprehensive confidence level and arbitration status of each modality's data; the evidence arbitration module is specifically used to: compare the structured semantic feature combinations corresponding to multiple modalities; when it is detected that the structured semantic feature combinations of multiple modalities point to the same equivalent state, increase the first numerical increment of the comprehensive confidence level; When a directional conflict is detected in the combination of structured semantic features of multiple modalities, the weight of the modal data with a lower initial evidence strength rating is reduced, and conflict state description information is output. Specifically, the first evidence strength score of the first modal data and the second evidence strength score of the second modal data in which the directional conflict occurs are obtained respectively. When the first evidence strength score is less than the second evidence strength score, a first attenuation coefficient γ is assigned to the first modal data. The feature source node data that generated the directional conflict is extracted, and the conflict state description information is generated based on the feature source node data. The first attenuation coefficient γ is generated by a preset dynamic penalty function, and the semantic distance scalar S between the first modal data and the second modal data is calculated through the internal feature space of the generative language processing model. Pre-existing features are read. The association strength matrix in the graph is used to obtain the first historical variance parameter corresponding to the first modality data and the second historical variance parameter corresponding to the second modality data; the confidence penalty term λ is calculated based on the ratio of the first historical variance parameter to the second historical variance parameter; a first exponential decay base α is constructed based on the logarithmic product of the semantic distance scalar S and the constant 1 plus the confidence penalty term λ; the difference between the second evidence strength score and the first evidence strength score is substituted into the smoothing activation function, and the output result is superimposed on the first exponential decay base α to obtain the final first decay coefficient γ, and the global judgment weight library is updated using the first decay coefficient γ; when data loss or isolated evidence of a single modality is detected in a specific modality, a corresponding limited confidence level indicator or insufficient causal judgment indicator is added; The feedback generation module, connected to the evidence arbitration module, is used to input the structured semantic feature combination after weight integration and logical arbitration into a pre-configured generative language processing model to generate an explanation conclusion carrying traceable source data and confidence level identifiers. When the overall confidence level of the explanation conclusion is lower than a preset first confidence level threshold, a supplementary evaluation instruction is added to the output.

10. An electronic device, comprising: Memory, used to store computer programs; A processor, communicatively connected to the memory, implements the method as described in any one of claims 1 to 8 when executing the computer program.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the method as described in any one of claims 1 to 8.

12. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Multi-modal data driven general report generation method and system based on large model

    CN121031543A

  • Multi-modal medical data fusion prediction and interpretable report generation system, method and device, processor and storage medium thereof

    CN121641324A