Anti-illusion training method and device for multi-modal model, equipment and storage medium

By adjusting the modal weights through delayed mutual information causal discovery and dynamic gating mechanism, adversarial samples are constructed for causal regularization training, which solves the data fusion distortion problem caused by unbalanced modal weights in multimodal models and improves the accuracy of financial risk prediction and medical diagnosis.

CN120670841APending Publication Date: 2025-09-19PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510724318.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Multimodal models have data fusion distortion problems in financial risk prediction and medical diagnosis, especially data neglect and misjudgment caused by uneven distribution of modal weights.

Method used

A time-delayed mutual information causal discovery algorithm is used to extract cross-modal causal relationships, a dynamic gating mechanism is used to adjust modal weights, and adversarial multimodal samples are constructed for causal regularization constrained adversarial training to generate an anti-hallucination model.

Benefits of technology

It improves the reasoning accuracy of multimodal models, solves the data fusion distortion problem caused by improper configuration of modal weights, and improves the accuracy and consistency of decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670841A_ABST
    Figure CN120670841A_ABST
Patent Text Reader

Abstract

The invention discloses an anti-illusion training method and device for a multi-modal model, equipment and a medium, and the method comprises the steps: firstly introducing a time delay mutual information causal discovery algorithm in a data preprocessing stage, and building a cross-modal causal atlas between modal data; and dynamically adjusting the fusion weight of each modal data according to the dynamic gating mechanism of the second stage and the causal relationship between the modal data corresponding to the cross-modal causal atlas so as to solve the cross-modal conflict resolution capability and solve the problem of data fusion distortion caused by traditional static weight distribution. And on the basis of the causal relationship between the modal data corresponding to the cross-modal causal atlas, constructing an adversarial multi-modal sample on the basis of the original multi-modal data, and carrying out constrained adversarial training on the multi-modal model on the basis of causal regularization to obtain an anti-illusion multi-modal model, so that the reasoning precision of the multi-modal model is improved. The method can be applied to the financial risk prediction field and the medical diagnosis field so as to improve the risk prediction accuracy and the diagnosis accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent decision-making technology, which is applied to the fields of financial risk prediction and medical diagnosis, and in particular to an anti-hallucination training method, device, computer equipment and computer-readable storage medium for a multimodal model. Background Art

[0002] Currently, multimodal models used in financial risk prediction suffer from widespread data fusion distortion. For example, while mainstream platforms in this field, such as Bloomberg's AIM system, integrate multimodal data, including news text, financial report charts, and market data, the AIM system overweights the text modality in its decision-making process. This leads to the systematic neglect of implicit risk features in the charts (such as unusual cross-references in the notes to the financial statements). This problem is particularly pronounced during unexpected market events. For example, when a company's official statement contradicts supply chain sensor data, existing systems, due to the heavy weighting of rigid text modality data in their decision-making process, often mistakenly accept the text information. This led to the stock price misjudgment incident during the 2022 Meta earnings season, caused by a misinterpretation of the ARPU chart. This same challenge also exists in medical diagnostic systems. Imaging diagnostic platforms, such as IBM Watson Health, also suffer from an imbalance in attention allocation when processing multimodal medical data. For example, when CT images show ground-glass nodules in the lungs and the electronic medical record does not record related symptoms, the system's attention to image features drops to less than 15%, resulting in 23% missed diagnoses of early lung cancer.

[0003] Therefore, how to solve the data fusion distortion in the current multimodal model has become an urgent problem to be solved. Summary of the Invention

[0004] The present invention provides a method, apparatus, computer device and computer-readable storage medium for anti-hallucination training of a multimodal model to solve the technical problem of data fusion distortion in current multimodal models.

[0005] In a first aspect, a method for anti-hallucination training of a multimodal model is provided, comprising:

[0006] Based on the time-delay mutual information causal discovery algorithm, the original multimodal data is preprocessed, and the causal relationship between each modal data is extracted from the original multimodal data to obtain a cross-modal causal graph;

[0007] Based on the dynamic gating mechanism and the cross-modal causal graph, the weight of each modal data in the multimodal model to be optimized is dynamically adjusted to obtain the target weight of each modal data;

[0008] Based on the cross-modal causal graph, modal contradictions are constructed on the original multimodal data to generate adversarial multimodal samples, and based on the target weights and the cross-modal causal graph, causal regularization constrained adversarial training is performed on the multimodal model.

[0009] In a second aspect, a multimodal model anti-hallucination training device is provided, comprising:

[0010] The causal graph generation module is used to preprocess the original multimodal data based on the time-delay mutual information causal discovery algorithm, and extract the causal relationship between each modal data from the original multimodal data to obtain a cross-modal causal graph;

[0011] A modal weight gating module is used to dynamically adjust the weight of each modal data in the multimodal model to be optimized based on the dynamic gating mechanism and the cross-modal causal graph to obtain the target weight of each modal data;

[0012] The model anti-hallucination training module is used to construct modal contradictions on the original multimodal data based on the cross-modal causal graph, generate adversarial multimodal samples, and perform causal regularization constrained adversarial training on the multimodal model based on the target weights and the cross-modal causal graph.

[0013] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the anti-hallucination training method for the multimodal model are implemented.

[0014] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the anti-hallucination training method of the multimodal model are implemented.

[0015] In the scheme implemented by the above-mentioned anti-hallucination training method, device, computer equipment and computer-readable storage medium for multimodal models, based on the time-delayed mutual information causal discovery algorithm, the original multimodal data is preprocessed, and the causal relationship between each modal data is extracted from the original multimodal data to obtain a cross-modal causal graph; based on the dynamic gating mechanism and the cross-modal causal graph, the weight of each modal data in the multimodal model to be optimized is dynamically adjusted to obtain the target weight of each modal data; based on the cross-modal causal graph, the original multimodal data is modally contradictory constructed to generate adversarial multimodal samples, and based on the target weight and the cross-modal causal graph, the multimodal model is subjected to causal regularization constrained adversarial training. Through the above method, the time-delayed mutual information causal discovery algorithm is first introduced in the data preprocessing stage to establish a cross-modal causal graph between each modal data. Then, through the dynamic gating mechanism of the second stage and the causal relationship between the modal data corresponding to the cross-modal causal graph, the fusion weights of each modal data are dynamically adjusted, enabling the multimodal model to master the cross-modal conflict resolution capability and solve the data fusion distortion problem caused by traditional static weight allocation. Based on the causal relationship between the modal data corresponding to the cross-modal causal graph, adversarial multimodal samples are constructed based on the original multimodal data. Based on causal regularization, constrained adversarial training of the multimodal model is performed to obtain a multimodal model that is resistant to hallucinations and improves the inference accuracy of the multimodal model. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0017] Figure 1 1 is a flow chart of a method for anti-hallucination training of a multimodal model according to an embodiment of the present invention;

[0018] Figure 2 yes Figure 1 A schematic flow chart of a specific implementation of step S10;

[0019] Figure 3 yes Figure 1 Another specific implementation flow diagram of step S20;

[0020] Figure 4 2 is a schematic diagram of an anti-hallucination training architecture of a multimodal model in one embodiment of the present invention;

[0021] Figure 5 1 is a schematic structural diagram of an anti-hallucination training device of a multimodal model according to an embodiment of the present invention;

[0022] Figure 6 It is a structural diagram of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION

[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0024] The anti-hallucination training method of the multimodal model involved in the embodiments of the present application is mainly applied to computer equipment. The anti-hallucination training generation device of the multimodal model can be a PC, a portable computer, a mobile terminal, or other device with display and processing functions.

[0025] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features therein may be combined with each other.

[0026] Reference Figure 1 , Figure 1 A flowchart of an anti-hallucination training method for a multimodal model provided in this application.

[0027] like Figure 1 As shown, an embodiment of the present application provides an anti-hallucination training method for a multimodal model, and the anti-hallucination training method for a multimodal model includes steps S10 to S30.

[0028] In this embodiment, the anti-hallucination training method of the multimodal model includes the following steps:

[0029] Step S10: pre-processing the original multimodal data based on the time-delay mutual information causal discovery algorithm, and extracting the causal relationship between each modal data from the original multimodal data to obtain a cross-modal causal graph;

[0030] Currently, multimodal systems in the financial and healthcare sectors commonly suffer from data fusion distortion. In the field of financial risk forecasting, while mainstream platforms like Bloomberg's AIM system can integrate news text, financial report charts, and market data, the text modality accounts for over 68% of the weight allocated in actual decision-making, leading to a systematic disregard for implicit risk features in charts (such as unusual cross-references in financial statement notes). In other words, in financial scenarios, multimodal systems rely too heavily on text and ignore more structured multimodal signals like charts, creating a "modal dominance illusion." This phenomenon is particularly pronounced during sudden market events. For example, when a company's official statement contradicts supply chain sensor data, existing systems, due to rigid modal prioritization, often mistakenly prioritize text information. This led to a stock price misjudgment during the 2022 Meta earnings season due to a misinterpretation of an ARPU chart.

[0031] Medical diagnostic systems also face the challenge of cross-modal conflict. Imaging diagnostic platforms, such as IBM Watson Health, suffer from an imbalance in attention allocation when processing multimodal medical data. Clinical testing has shown that when CT images reveal ground-glass nodules in the lungs but the electronic medical record lacks relevant symptoms, the system's attention to the image features drops to less than 15%, resulting in 23% missed diagnoses of early-stage lung cancer.

[0032] What's more serious is that the modal prior bias accumulated by traditional models during the training process will continue to amplify over time as they are deployed. For example, in skin cancer diagnosis, systems trained based on Caucasian samples generally pay less attention to the lesion areas of dark-skinned patients, and the misdiagnosis rate is 19 percentage points higher than the baseline model. In other words, due to the bias in the distribution of training samples and the imbalance in the modal attention mechanism, the system has accumulated an attention bias towards "Caucasian skin color" samples during training, resulting in a decrease in the attention area in the visual modality when dealing with dark-skinned patients. Therefore, this problem (caused by training data bias and attention mechanism bias) and the modal weight configuration problem (multimodal systems rely too much on text and ignore more structural multimodal signals such as charts) are both "modal attention allocation imbalance" and belong to the same category of "modal dominance bias."

[0033] To address the problem of data fusion distortion in multimodal models, this embodiment addresses the aforementioned issues of systematic neglect, attention deviation, and training bias, respectively, and reconstructs the relationship between attention weights and modalities at the causal level. Therefore, the two are structurally related issues at the low level. Specifically, this is achieved through a three-level anti-hallucination processing structure consisting of data preprocessing, dynamic gating mechanism, and adversarial training:

[0034] First, during the data preprocessing phase, a time-delayed mutual information causal discovery algorithm and a dynamic Bayesian network were used to model cross-modal causal relationships. In the financial scenario, the time-lagged correlation (τ = 5 minutes) between social media sentiment and stock index futures fluctuations was captured, and a causal relationship graph was reconstructed every 15 minutes. In the medical emergency scenario, the causal links between vital signs and imaging features were updated at 0.3-second intervals.

[0035] The second stage is a dynamic gating mechanism that uses an adaptive weight allocation strategy driven by KL divergence. When a cross-modal conflict is detected, a triple verification mechanism is triggered: freezing gradients, historical case matching, and attention recalibration.

[0036] The third stage is the anti-hallucination training framework, which includes adversarial sample generation and causal regularization. It constructs modality-contradictory sample pairs and uses the pseudo-correlation attention penalty term in the loss function.

[0037] Through the above three stages, the original multimodal data is obtained, the causal graph is generated through the causal discovery module, and then the weights are adjusted by the dynamic gating mechanism. Finally, adversarial training and regularization are performed in the training framework.

[0038] Specifically, during the data preprocessing process in the first stage, the causal relationship between cross-modal data is extracted from the original multimodal data through the time-delay mutual information causal discovery algorithm, and a cross-modal causal graph is constructed to guide the modal weight allocation of subsequent modules.

[0039] Multimodal data includes textual data (such as social media sentiment and news), image data (such as financial report PDFs), and numerical data (such as stock index futures prices) in financial scenarios, as well as received images (CT / MRI), text (electronic medical records), and time series signals (vital signs) in medical scenarios. A causal graph represents the causal relationships between modalities using a directed graph (e.g., the causal edge weight from image to diagnosis is 0.7).

[0040] Among them, such as Figure 2 As shown, the step S10 specifically includes:

[0041] Step S11, aligning each modal data in the original multimodal data according to a time unit, and constructing a time series modal data pair based on the sliding window and the aligned modal data;

[0042] Step S12, based on each time series modal data pair, calculating the mutual information of each modal data at different time delays to obtain the time delay mutual information;

[0043] Step S13: determining the causal time lag corresponding to the maximum mutual information based on the time-delay mutual information, and modeling the causal relationship between the modal data based on the dynamic Bayesian network and the time lag corresponding to the maximum mutual information to obtain the cross-modal causal graph.

[0044] In this embodiment, the impact of cross-modal data (such as financial text sentiment and stock prices, medical images and vital signs) often has a time delay (for example, it takes 5 minutes for sentiment to be reflected in stock prices). Therefore, the time series data is first decomposed into discrete time slices (for example, one slice per minute in financial scenarios and one slice per 0.3 seconds in medical scenarios);

[0045] Multimodal data is then aligned to the time granularity required by the scenario. For example, for modal data in financial scenarios, it is aligned to a 1-minute granularity, with a 15-minute sliding window to construct sample pairs. For modal data in medical scenarios, it is aligned to the millisecond level (0.3-second intervals), and data slices are updated in real time. This sliding window mechanism retains the latest time slice data, reduces redundant data, and avoids historical noise interference.

[0046] Then, the mutual information of cross-modal variables (such as text sentiment X and stock price fluctuation Y) at different time delays τ is calculated, and the cross-time slice dependencies between variables are explicitly modeled based on the dynamic Bayesian network (DBN), for example: social media sentiment (t-τ) → stock index fluctuation (t) (τ = 5 minutes), vital signs (t-0.3s) → image features (t) (medical scenario).

[0047] Among them, the DBN structure definition includes:

[0048] The node definition includes time slice t-1:X t-1 (emotion), Y t-1 (stock price) and time slice t: Y t (Current stock price).

[0049] Edge definition: Cross-time edge: X t-1 →Y t ,Y t-1 →Y t (Stock price autocorrelation).

[0050] Specifically, calculating the mutual information under different time delays includes the following steps:

[0051] 1. Prepare time series data: Given two variable time series (such as text sentiment X and stock price fluctuation Y)

[0052] X={x1,x2,...,xT}, Y={y1,y2,...,yT}.

[0053] 2. Select the time delay range: Set the delay range τin[1,20] (minutes)

[0054] 3. Construct delayed variable pairs: For each τ, construct a delayed sequence pair

[0055] X (τ) ={x1,x2,...,x T-τ},Y (τ) ={y 1+τ ,y 2+τ ,...,y T}

[0056] 4. Calculate mutual information:

[0057]

[0058] Here, p(x, y) is the joint probability of text sentiment x and stock price volatility y, p(x) is the marginal distribution of text sentiment x, and p(y) is the marginal distribution of stock price volatility y. These probabilities are estimated using a binning method. The range of values ​​for x and y is divided into a fixed number of bins. The probability of data occurring in each bin is counted to estimate p(x), p(y), and p(x, y).

[0059] For example, when calculating the mutual information between social media sentiment and stock index futures in a financial scenario, the mutual information is maximum when τ = 5 minutes, indicating that sentiment has a 5-minute lag effect on stock prices. This is illustrated using the text modality (TextSentiment), image modality (ImageFeature), and numerical modality (PriceVolatility) in a financial scenario as examples:

[0060] The causal relationship between multiple time steps is:

[0061] TextSentiment(t-1)→PriceVolatility(t);

[0062] ImageFeature(t-1)→PriceVolatility(t);

[0063] PriceVolatility(t-1)→PriceVolatility(t).

[0064] Through dynamic Bayesian network modeling, a cross-modal causal relationship graph is constructed:

[0065] The nodes in the causal relationship graph are used to represent various modal variables (such as text sentiment, image features, vital signs, etc.), and the edges in the causal relationship graph are used to represent the causal direction and time lag determined based on mutual information (such as text sentiment (t-τ) → stock price fluctuation (t)).

[0066] In the first stage, the time sensitivity and causality of multimodal data are uniformly modeled through dynamic Bayesian networks, thus providing a dynamic and interpretable causal reasoning basis for subsequent modules.

[0067] Step S20: Based on the dynamic gating mechanism and the cross-modal causal graph, dynamically adjust the weight of each modal data in the multimodal model to be optimized to obtain the target weight of each modal data;

[0068] In this embodiment, each modal feature vector corresponding to each modal data extracted in the first stage (such as the text encoding feature f text , image encoding features f image , numerical encoding feature f num ) and cross-modal causal graphs (such as the causal strength and time lag of text sentiment → stock price fluctuations), input the dynamic gating mechanism, and pass through the second stage of dynamic gating, that is, adaptive weight allocation driven by KL divergence, to quantify the cross-modal feature distribution differences, detect potential conflicts, and dynamically adjust the modal fusion weights according to the real-time feature distribution differences to resolve potential conflicts.

[0069] Among them, such as Figure 3 As shown, step S20 includes:

[0070] Step S21, obtaining the data feature distribution corresponding to each modal data, and calculating the distribution difference value between each modal data;

[0071] Step S22: performing cross-modal conflict detection on each modal data based on the KL divergence and the distribution difference value between each modal data;

[0072] Step S23: When cross-modal conflicts are detected in each modal data, the weights of each modal data are dynamically adjusted based on historical cases and the cross-modal causal graph to obtain target weights for each modal data.

[0073] Specifically, the probability distribution of the feature vector of each modality is calculated by kernel density estimation or binning method. For example, the text modality distribution P text : Based on the probability density function of the sentiment score, the image modal distribution P image : Histogram distribution of activation values ​​based on the lesion area.

[0074] Based on the KL divergence calculation formula, the distribution difference between the two modes is calculated:

[0075]

[0076] P(i) and P(j) are the probability distributions corresponding to the two modal data respectively, and i is the sequence number of each data in each modal data.

[0077] When the financial report chart distribution P chart P with conference call voice distribution audio D KL >θ (i.e. the threshold is set to 1.2), then the financial report chart distribution P is determined chart P with conference call voice distribution audio The data is contradictory.

[0078] If the CT image distribution P CT and medical record text distribution P text D KL >θ, then determine the CT image distribution P CT and medical record text distribution P text The data is contradictory.

[0079] It is understandable that an adaptive threshold may be set according to the historical data distribution (eg, θ=μ+2σ, μ is the average KL divergence, σ is the standard deviation).

[0080] Wherein, the step S23 includes:

[0081] When cross-modal conflict is detected in each modal data, the gradient propagation data corresponding to the modal data with cross-modal conflict is frozen;

[0082] Based on the fusion vectors corresponding to each modality data, search and match historical conflict cases to identify historical similar cases;

[0083] Based on the weight correction strategy or attention recalibration strategy of the historical similar cases, the weights of the modal data are dynamically adjusted to obtain the target weight.

[0084] After step S22, the method further includes:

[0085] When there is no cross-modal conflict in each modal data, the weight of each modal data is modified based on the cross-modal causal graph to obtain the target weight of each modal data.

[0086] In this embodiment, if there is a contradiction between the two modal data, a triple verification mechanism is triggered: first, the gradient propagation of the disputed modality is frozen, then similar historical cases are called for pattern matching, and finally, the decision is corrected through cross-modal attention recalibration:

[0087] When calculating the weights of each modality, the dynamic gating module uses the causal strength value extracted from the causal graph as the initial weight benchmark. For example:

[0088] In a financial scenario, if the causal graph shows that the causal contribution of a financial report chart (image mode) to stock price prediction is 0.4 (higher than 0.3 of the text mode), then in the absence of conflict, the chart will be given a higher weight.

[0089] In a medical scenario, if the causal weight of CT images for lung cancer diagnosis in the causal graph is 0.7 (much higher than the 0.1 in the electronic medical record text), the gating mechanism will default to strengthening the decision-making dominance of the imaging modality.

[0090] That is, the causal strength value α in the causal graph is used as the initial value of the gating weight, and then dynamically adjusted through the KL divergence:

[0091] Weight final =α+γ·D KL (P||Q)

[0092] Among them, γ is the KL divergence adjustment coefficient, which is used to balance the causal prior and real-time feature differences.

[0093] When a difference in feature distribution between modalities is detected (KL divergence exceeds the threshold), the system calls the causal graph to verify the causal credibility of the conflicting modalities.

[0094] In a financial example, if the cash flow statement (image mode) in the financial report PDF shows a decrease in profits, while the conference call audio (text mode) claims an increase, then:

[0095] 1) Query the causal graph to confirm that the causal weight of financial report charts on stock price prediction is higher than that of voice text.

[0096] 2) Determine whether the graph modality is more credible and trigger gradient freezing (suppressing the propagation of speech-text gradients).

[0097] In medical cases, if a CT scan shows a lung nodule but the medical record does not describe the symptoms, then:

[0098] 1) Based on the strong causal link of "imaging → diagnosis" in the causal graph (weight 0.8 vs. text 0.05), the imaging modality is given priority.

[0099] 2) Call similar historical cases (nodules are found in the image but not recorded in the text) to correct the current diagnosis.

[0100] When calling historical similar cases, give priority to historical samples that match the current causal graph structure:

[0101] In financial scenarios, when dealing with the conflict between "financial report charts and voice", only historical cases in the causal graph that also use charts as the dominant mode (such as the Meta financial report misjudgment incident) are retrieved to avoid introducing interference samples with mismatched causal structures.

[0102] In a medical example, when diagnosing patients with dark skin, if the causal graph shows that the causal weight of the imaging modality needs to be improved, cases of "imaging features correcting text misjudgments" are screened from the historical library for transfer learning.

[0103] Therefore, the causal graph makes the dynamic gating mechanism explainable (i.e., the decision conforms to the causal logic) and robust (i.e., resistant to statistical noise interference).

[0104] The process of cross-modal attention recalibration based on historical similar cases is as follows:

[0105] 1. By freezing the gradient of the disputed modality, the back propagation of the disputed modality (such as speech) is suspended to prevent the wrong signal from affecting the model update.

[0106] 2. Retrieve cases similar to the current conflict scenario from the historical database:

[0107] 1) Construct a cross-modal unified vector space: First, use a pre-trained language model (such as BEAT) to extract text modal data features f text , use convolutional networks (such as ResNet) to extract image modality data features f image , use a temporal encoder (such as LSTM) to extract numerical sequence features f num , then concatenate the features of each modality to obtain multimodal features, and use the fusion layer to map the multimodal features to a shared semantic space to eliminate modal differences and make different modal data (text / image / value) comparable in a unified semantic space:

[0108] f fused =FusionLayer(f text , f image , f num ).

[0109] 2) Based on the cosine similarity calculation formula, calculate the current fusion feature f fused Similarity with all cases in the historical database.

[0110] Among them, a historical case index library is pre-built, and the fusion vectors of each historical case are stored in the vector database. The causal weight distribution corresponding to each modal data when the conflict occurs (such as 0.2 for text and 0.7 for image), conflict type label (predefined 78 types of error patterns, such as text-image contradiction) and correction result (that is, the final correct decision of the annotation, such as ignoring text and accepting image) are also recorded.

[0111] 3. Filter similar cases based on case confidence: retain cases with high confidence labels in the historical database (such as cases of correctly predicted stock price fluctuations in financial scenarios and imaging cases of pathological diagnosis in medical scenarios).

[0112] Specifically, similar cases can be retrieved from the historical case library based on the fusion features of the current conflict modal sample + the current causal graph.

[0113] Specifically, we can first filter out historical cases with high semantic similarity, and then filter out historical cases whose causal graph structure is the same as the causal graph structure of the current modal sample as historical similar cases.

[0114] Among them, the causal graph of historically similar cases contains the conflicting modalities of the current modal sample (such as text-diagnosis), and the difference between the causal weight of the dominant modality in the historically similar cases and the causal weight of the current modal sample is less than a preset threshold (such as 10%, when the current image weight is 0.7, the historical case needs to be between 0.63-0.77).

[0115] 4. Use reliable features from historical cases to dynamically reconstruct the current model’s attention to correct the current model’s attention distribution:

[0116] 1) Take the fusion feature of the current modality sample as the query vector, and the fusion feature of the matched historical similar cases as the key-value pair, that is, construct the attention query-key-value pair: the current fusion feature f fused and the integration characteristics of historical cases

[0117] 2) Based on the attention calculation company, calculate the cross-modal attention weight:

[0118] 3) Fusion correction based on the gating mechanism, that is, dynamically adjusting the historical experience weight according to the KL divergence.

[0119] Step S30: Based on the cross-modal causal graph, modal contradictions are constructed for the original multimodal data to generate adversarial multimodal samples, and causal regularization constrained adversarial training is performed on the multimodal model based on the target weight and the cross-modal causal graph.

[0120] In this embodiment, adversarial samples of modally contradictory data (such as a text about a 20% profit growth accompanied by a -5% financial report table) are constructed based on real samples (i.e., original multimodal data, text + image + numerical value, and their annotated labels).

[0121] Among them, real samples can be financial report texts, charts, stock price series + real rise and fall labels in financial scenarios, or CT images, electronic medical record texts, vital signs time series data + real diagnostic labels in medical scenarios.

[0122] Furthermore, in order to ensure that the tampered modal data is aligned with the original data in terms of timestamp and spatial dimensions, the modal alignment of real samples and adversarial samples is verified.

[0123] In financial scenarios, ensure that the time window of the tampered text is consistent with the time window of the financial report chart;

[0124] In medical scenarios, ensure that disturbed images are synchronized with the corresponding vital signs;

[0125] Thus, adversarial sample pairs (tampered text / image / value + original label) are generated.

[0126] The method of constructing modal contradictions on the original multimodal data based on the cross-modal causal graph to generate adversarial multimodal samples includes:

[0127] Based on the cross-modal causal graph, the original multimodal data is subjected to text tampering, image perturbation or numerical distortion, and corresponding conflict labels are added to the modified contradictory samples to generate the adversarial multimodal samples.

[0128] In this embodiment, text tampering involves modifying the text description to make it contradictory with the real data (such as reversing the emotional polarity), and a predefined contradiction rule library can be used (such as pairing "profit growth" with "net profit decline" in finance); image perturbation involves adding misleading visual features to the chart (such as abnormal K-line patterns), and misleading visual features can be generated through GAN (such as adding false lesions in medical images); numerical distortion involves adjusting the trend of time series data (such as randomly offsetting the heart rate data in vital signs by ±20%).

[0129] It is understandable that causal graphs (e.g., the causal strength of "text sentiment → stock price fluctuations") and dynamic fusion weights (real-time weight distribution ratios of each modality) can be used to guide adversarial sample generation, and regularization terms can be added to the loss function. For example, information in the causal graph can be used to adjust attention weights or constrain modal parameters to avoid spurious correlations.

[0130] Specifically, real data and adversarial data can be input into the model in a certain proportion for training, or trained alternately.

[0131] The loss function includes cross entropy and pseudo-correlation attention penalty terms, which are used to suppress 78 types of error models.

[0132] Then, based on the adversarial sample pairs, the model is trained for multimodal consistency:

[0133] 1. Mix real samples and adversarial samples in proportion (default 1:1) to form a training batch;

[0134] 2. Text / image / numerical modalities are respectively encoded and extracted through encoders (such as BERT, ResNet, LSTM);

[0135] 3. Based on the real-time weights received in the second stage (such as text weight and image weight), the weighted fusion features are obtained by fusion.

[0136] 4. Based on the above training batches and their corresponding weighted fusion features, perform model training.

[0137] The step of performing causal regularization constrained adversarial training on the multimodal model based on the target weight and the cross-modal causal graph includes:

[0138] Calculating a regularization term based on the cross-modal causal graph and the target weight;

[0139] The regularization term gradient is calculated and back-propagated, and the weight of the pseudo-correlation penalty term in the multimodal model is adjusted so that the target weight converges to the weight in the cross-modal causal graph.

[0140] In this embodiment, false associations are blocked through a cross-modal causal graph, that is, undefined modal associations in the causal graph (such as the pseudo-correlation edge of "value→image" in historical cases) are excluded to avoid erroneous corrections caused by data bias.

[0141] The regularization term is used to ensure that the corrected weight distribution (such as attention weight) output by the model during the retrieval and matching process is consistent with the modal causal relationship of the current causal graph, while adapting to the business logic after real-time adjustment.

[0142]

[0143] in, is the causal relationship constraint in the cross-modal causal graph (i.e., the causal graph structure constraint), Align the target weights (i.e., the modified weights obtained by matching historical death cases).

[0144] Then, backpropagation is performed based on the regularization term, the regularization term is added to the total model loss, and the parameters of the attention weight generation module are updated through gradient descent, so that the output weight gradually approaches the causal graph and the dynamically adjusted target distribution.

[0145] Causal regularization uses orthogonality constraints and gradient boosting during the training phase. Among them, the regularization term includes modal independence constraints (orthogonality) and causal path reinforcement (gradient boosting). After retrieving historical death cases and applying the cross-modal causal graph for filtering, the regularization term is dynamically adjusted according to the fusion weights of each modality in the cross-modal causal graph to further optimize the fusion features or attention weights to ensure consistency with the current causal structure. During dynamic attention recalibration, the regularization term is calculated to make the attention distribution of the current model consistent with the weights in the causal graph, and it can also penalize attention weights that are inconsistent with the causal graph.

[0146] In a specific embodiment, in the loss function design, in addition to the standard cross entropy loss, a pseudo-correlation attention penalty term is specially added.

[0147] Cross entropy loss function formula: For a multi-classification task (three-class prediction of "up, down, flat"), the standard cross entropy loss is:

[0148]

[0149] Among them, C is the number of prediction categories (such as rise, fall, and flat in financial transactions), y i is the true label (one-hot) of the i-th category (such as the true trend of the sample in financial transactions, if the true trend is "rising", then y i =[1,0,0]), The probability of this category predicted by the model (such as the confidence level of the model's prediction of the i-th category trend in financial transactions (such as the probability of an increase of 0.8)).

[0150] Calculate the pseudo-correlation attention penalty based on cross entropy:

[0151] L total =L CE +β·L atten

[0152] L atten is the pseudo-correlation attention penalty term, where α i is the model’s attention distribution on the i-th modal token, is the prior attention distribution, KL is the KL divergence calculation formula, the same as D mentioned above KL (P||Q). β is a weight adjustment term, a manually defined parameter, and its default value is 1.

[0153] It can be seen that in the above scheme, the regularization term is used to force the model to reduce the attention weight of the text modality on the risk level, and the modality weight is dynamically fine-tuned according to similar historical cases (i.e., historical experience) to avoid weight mutations.

[0154] like Figure 4 As shown, this embodiment constructs a three-level linkage anti-hallucination processing architecture to address the modality dominance hallucination problem caused by the fusion of multi-source data in financial decision support and medical diagnosis scenarios.

[0155] First, a time-delayed mutual information causal discovery algorithm is introduced during data preprocessing, modeling cross-modal causal relationships using a dynamic Bayesian network. For financial trading scenarios, this module reconstructs a causal relationship graph every 15 minutes, capturing, for example, the time-lagged correlation (τ = 5 minutes) between social media sentiment and stock index futures fluctuations. For medical emergency scenarios, the causal link between vital sign data and imaging features is continuously updated at 0.3-second intervals.

[0156] The dynamic gating mechanism in the second stage adopts an adaptive weight allocation strategy driven by KL divergence. This module dynamically adjusts the fusion weights by calculating the differences in the feature distribution of each modality in real time. When a cross-modal conflict is detected (such as the contradiction between the cash flow statement data in the financial report PDF and the voice recording of the conference call in the financial scenario), the system automatically triggers a triple verification mechanism: first, the gradient propagation of the disputed modality is frozen, then similar historical cases are called for pattern matching, and finally the decision correction is completed through cross-modal attention recalibration. In medical imaging diagnosis, this mechanism can increase the attention coverage of key lesion areas from 5% of the traditional model to 63%.

[0157] The third-level anti-hallucination training framework comprises two modules: adversarial example generation and causal regularization. By constructing modal conflicting pairs (e.g., generating a text description of "net profit growth of 20%" paired with a numerical table of -5% in the financial statements for the same period), the model is forced to learn cross-modal consistency verification capabilities. In the loss function design, in addition to the standard cross-entropy loss, a pseudo-correlation attention penalty is specifically incorporated to specifically suppress 78 types of erroneous patterns, such as the false correlation between weather data and stock price fluctuations, which are common in financial scenarios. The medical training dataset MedConFLICT contains 120,000 manually annotated image-text conflict cases. Through transfer learning, the model acquires cross-modal conflict resolution capabilities.

[0158] Through the above approach, this embodiment provides a dynamic anti-interference framework based on attentional causal interpretation. By establishing a cross-modal causal graph, developing an adaptive gating mechanism, and constructing an adversarial training system, it significantly improves the accuracy of multimodal joint reasoning in complex scenarios.

[0159] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0160] In one embodiment, a multimodal model anti-hallucination training device is provided, and the multimodal model anti-hallucination training device corresponds to the multimodal model anti-hallucination training processing method in the above embodiment. Figure 5 As shown, the anti-hallucination training device for the multimodal model includes a causal graph generation module 101, a modal weight gating module 102, and a model anti-hallucination training module 103. The functional modules are described in detail as follows:

[0161] The causal graph generation module 101 is used to pre-process the original multimodal data based on the time-delay mutual information causal discovery algorithm, and extract the causal relationship between each modal data from the original multimodal data to obtain a cross-modal causal graph;

[0162] A modal weight gating module 102 is configured to dynamically adjust the weights of each modal data in the multimodal model to be optimized based on a dynamic gating mechanism and the cross-modal causal graph to obtain a target weight for each modal data;

[0163] The model anti-hallucination training module 103 is used to construct modal contradictions on the original multimodal data based on the cross-modal causal graph, generate adversarial multimodal samples, and perform causal regularization constrained adversarial training on the multimodal model based on the target weight and the cross-modal causal graph.

[0164] Furthermore, the causal graph generation module 101 is further configured to:

[0165] Aligning each modal data in the original multimodal data according to a time unit, and constructing a time series modal data pair based on the sliding window and the aligned modal data;

[0166] Based on each time series modal data pair, the mutual information of each modal data at different time delays is calculated to obtain the time delay mutual information;

[0167] The causal time lag corresponding to the maximum mutual information is determined based on the time-delay mutual information, and the causal relationship between the modal data is modeled based on a dynamic Bayesian network and the time lag corresponding to the maximum mutual information to obtain the cross-modal causal graph.

[0168] Furthermore, the modal weight gating module 102 is further configured to:

[0169] Obtain the data feature distribution corresponding to each modal data, and calculate the distribution difference value between each modal data;

[0170] Based on the KL divergence and the distribution difference between the modal data, cross-modal conflict detection is performed on the modal data;

[0171] When cross-modal conflicts are detected in the modal data, the weights of the modal data are dynamically adjusted based on historical cases and the cross-modal causal graph to obtain the target weights of the modal data.

[0172] Furthermore, the modal weight gating module 102 is further configured to:

[0173] When cross-modal conflict is detected in each modal data, the gradient propagation data corresponding to the modal data with cross-modal conflict is frozen;

[0174] Based on the fusion vectors corresponding to each modality data, search and match historical conflict cases to identify historical similar cases;

[0175] Based on the weight correction strategy or attention recalibration strategy of the historical similar cases, the weights of the modal data are dynamically adjusted to obtain the target weight.

[0176] Furthermore, the modal weight gating module 102 is further configured to:

[0177] When there is no cross-modal conflict in each modal data, the weight of each modal data is modified based on the cross-modal causal graph to obtain the target weight of each modal data.

[0178] Furthermore, the model anti-hallucination training module 103 is further used to:

[0179] Based on the cross-modal causal graph, the original multimodal data is subjected to text tampering, image perturbation or numerical distortion, and corresponding conflict labels are added to the modified contradictory samples to generate the adversarial multimodal samples.

[0180] Furthermore, the model anti-hallucination training module 103 is further used to:

[0181] Calculating a regularization term based on the cross-modal causal graph and the target weight;

[0182] The regularization term gradient is calculated and back-propagated, and the weight of the pseudo-correlation penalty term in the multimodal model is adjusted so that the target weight converges to the weight in the cross-modal causal graph.

[0183] For specific definitions of the multimodal model anti-hallucination training device, please refer to the definitions of the multimodal model anti-hallucination training method above and will not be repeated here. The various modules in the above-mentioned multimodal model anti-hallucination training device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.

[0184] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 6As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps of an anti-hallucination training method for a multimodal model.

[0185] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0186] Based on the time-delay mutual information causal discovery algorithm, the original multimodal data is preprocessed, and the causal relationship between each modal data is extracted from the original multimodal data to obtain a cross-modal causal graph;

[0187] Based on the dynamic gating mechanism and the cross-modal causal graph, the weight of each modal data in the multimodal model to be optimized is dynamically adjusted to obtain the target weight of each modal data;

[0188] Based on the cross-modal causal graph, modal contradictions are constructed on the original multimodal data to generate adversarial multimodal samples, and based on the target weights and the cross-modal causal graph, causal regularization constrained adversarial training is performed on the multimodal model.

[0189] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0190] Aligning each modal data in the original multimodal data according to a time unit, and constructing a time series modal data pair based on the sliding window and the aligned modal data;

[0191] Based on each time series modal data pair, the mutual information of each modal data at different time delays is calculated to obtain the time delay mutual information;

[0192] The causal time lag corresponding to the maximum mutual information is determined based on the time-delay mutual information, and the causal relationship between the modal data is modeled based on a dynamic Bayesian network and the time lag corresponding to the maximum mutual information to obtain the cross-modal causal graph.

[0193] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0194] Obtain the data feature distribution corresponding to each modal data, and calculate the distribution difference value between each modal data;

[0195] Based on the KL divergence and the distribution difference between the modal data, cross-modal conflict detection is performed on the modal data;

[0196] When cross-modal conflicts are detected in the modal data, the weights of the modal data are dynamically adjusted based on historical cases and the cross-modal causal graph to obtain the target weights of the modal data.

[0197] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0198] When cross-modal conflict is detected in each modal data, the gradient propagation data corresponding to the modal data with cross-modal conflict is frozen;

[0199] Based on the fusion vectors corresponding to each modality data, search and match historical conflict cases to identify historical similar cases;

[0200] Based on the weight correction strategy or attention recalibration strategy of the historical similar cases, the weights of the modal data are dynamically adjusted to obtain the target weight.

[0201] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0202] When there is no cross-modal conflict in each modal data, the weight of each modal data is modified based on the cross-modal causal graph to obtain the target weight of each modal data.

[0203] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0204] Based on the cross-modal causal graph, the original multimodal data is subjected to text tampering, image perturbation or numerical distortion, and corresponding conflict labels are added to the modified contradictory samples to generate the adversarial multimodal samples.

[0205] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0206] Calculating a regularization term based on the cross-modal causal graph and the target weight;

[0207] The regularization term gradient is calculated and back-propagated, and the weight of the pseudo-correlation penalty term in the multimodal model is adjusted so that the target weight converges to the weight in the cross-modal causal graph.

[0208] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can refer to the relevant description in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0209] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0210] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0211] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A method for anti-hallucination training of a multimodal model, characterized in that: include: Based on the time-delay mutual information causal discovery algorithm, the original multimodal data is preprocessed, and the causal relationship between each modal data is extracted from the original multimodal data to obtain a cross-modal causal graph; Based on the dynamic gating mechanism and the cross-modal causal graph, the weight of each modal data in the multimodal model to be optimized is dynamically adjusted to obtain the target weight of each modal data; Based on the cross-modal causal graph, modal contradictions are constructed on the original multimodal data to generate adversarial multimodal samples, and based on the target weights and the cross-modal causal graph, causal regularization constrained adversarial training is performed on the multimodal model.

2. The anti-hallucination training method of the multimodal model according to claim 1, characterized in that: The time-delay mutual information-based causal discovery algorithm preprocesses the original multimodal data and extracts the causal relationship between each modal data from the original multimodal data to obtain a cross-modal causal graph, including: Aligning each modal data in the original multimodal data according to a time unit, and constructing a time series modal data pair based on the sliding window and the aligned modal data; Based on each time series modal data pair, the mutual information of each modal data at different time delays is calculated to obtain the time delay mutual information; The causal time lag corresponding to the maximum mutual information is determined based on the time-delay mutual information, and the causal relationship between the modal data is modeled based on a dynamic Bayesian network and the time lag corresponding to the maximum mutual information to obtain the cross-modal causal graph.

3. The anti-hallucination training method of the multimodal model according to claim 1, characterized in that: Based on the dynamic gating mechanism and the cross-modal causal graph, the weight of each modal data in the multimodal model to be optimized is dynamically adjusted to obtain the target weight of each modal data, including: Obtain the data feature distribution corresponding to each modal data, and calculate the distribution difference value between each modal data; Based on the KL divergence and the distribution difference between the modal data, cross-modal conflict detection is performed on the modal data; When cross-modal conflicts are detected in the modal data, the weights of the modal data are dynamically adjusted based on historical cases and the cross-modal causal graph to obtain the target weights of the modal data.

4. The anti-hallucination training method of the multimodal model according to claim 3, characterized in that: When cross-modal conflicts are detected in the modal data, the weights of the modal data are dynamically adjusted based on historical cases and the cross-modal causal graph to obtain target weights for the modal data, including: When cross-modal conflict is detected in each modal data, the gradient propagation data corresponding to the modal data with cross-modal conflict is frozen; Based on the fusion vectors corresponding to each modality data, search and match historical conflict cases to identify historical similar cases; Based on the weight correction strategy or attention recalibration strategy of the historical similar cases, the weights of the modal data are dynamically adjusted to obtain the target weight.

5. The anti-hallucination training method of the multimodal model according to claim 3, characterized in that: After performing cross-modal conflict detection on each modal data based on the KL divergence and the distribution difference value between each modal data, the method further includes: When there is no cross-modal conflict in each modal data, the weight of each modal data is modified based on the cross-modal causal graph to obtain the target weight of each modal data.

6. The anti-hallucination training method of the multimodal model according to claim 1, wherein: The method of constructing modal contradictions on the original multimodal data based on the cross-modal causal graph to generate adversarial multimodal samples includes: Based on the cross-modal causal graph, the original multimodal data is subjected to text tampering, image perturbation or numerical distortion, and corresponding conflict labels are added to the modified contradictory samples to generate the adversarial multimodal samples.

7. The anti-hallucination training method of a multimodal model according to any one of claims 1 to 6, characterized in that: The performing causal regularization constrained adversarial training on the multimodal model based on the target weight and the cross-modal causal graph includes: Calculating a regularization term based on the cross-modal causal graph and the target weight; The regularization term gradient is calculated and back-propagated, and the weight of the pseudo-correlation penalty term in the multimodal model is adjusted so that the target weight converges to the weight in the cross-modal causal graph.

8. A multimodal model anti-hallucination training device, characterized in that: include: The causal graph generation module is used to preprocess the original multimodal data based on the time-delay mutual information causal discovery algorithm, and extract the causal relationship between each modal data from the original multimodal data to obtain a cross-modal causal graph; A modal weight gating module is used to dynamically adjust the weight of each modal data in the multimodal model to be optimized based on the dynamic gating mechanism and the cross-modal causal graph to obtain the target weight of each modal data; The model anti-hallucination training module is used to construct modal contradictions on the original multimodal data based on the cross-modal causal graph, generate adversarial multimodal samples, and perform causal regularization constrained adversarial training on the multimodal model based on the target weights and the cross-modal causal graph.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the anti-hallucination training method of the multimodal model according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the anti-hallucination training method of the multimodal model are implemented.

Citation Information

Cited By

  • Multi-modal causal reasoning method and device based on large language model

    CN121052389A

  • Well engineering-oriented multi-modal abstract generation method and device and storage medium

    CN121256057A

  • A method, apparatus and storage medium for generating multimodal summaries for well engineering

    CN121256057B

  • Physical constraint-based multi-mode indoor carbon dioxide concentration prediction method

    CN121580357A