A method, system, terminal, and storage medium for missing modality inference based on manifold-aware diffusion completion and evidence gating fusion.

By fusing manifold-aware diffusion completion with evidence gating, the problem of error propagation in missing modality completion in multimodal understanding tasks is solved. This approach enables explicit modeling and adaptive fusion of modality reliability in the time dimension, thereby improving the accuracy and stability of missing modality completion.

CN122047515BActive Publication Date: 2026-07-17SHENZHEN MSU-BIT UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN MSU-BIT UNIVERSITY
Filing Date
2026-04-15
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In existing technologies for multimodal understanding tasks, the missing modality completion process is prone to error propagation and time-varying noise instability, resulting in inaccurate completion results. Furthermore, there is a lack of explicit modeling of completion credibility and modality reliability, making it difficult to adaptively reduce the impact of unreliable modalities in the time dimension.

Method used

A missing modality inference method based on manifold-aware diffusion completion and evidence gating fusion is adopted. By mapping multimodal time-series samples to the manifold latent space, noise is predicted using a denoising network and a denoising loss function is constructed. The loss function is constructed by combining Dirichlet parameters and confidence scores, and residual regularization and geometric regularization terms are introduced to train the missing modality inference model and output the predicted class and final class probability of the missing sample.

Benefits of technology

In the latent space, semantic offset of the completion trajectory is reduced, the reliability and uncertainty of each modality are explicitly modeled, adaptive weighted fusion is used to suppress the interference of unreliable modalities, and the accuracy and stability of missing modality completion are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122047515B_ABST
    Figure CN122047515B_ABST
Patent Text Reader

Abstract

This invention relates to the field of data processing technology and discloses a method, system, terminal, and storage medium for missing modality inference based on manifold-aware diffusion completion and evidence gating fusion. The method includes: performing manifold-aware diffusion completion in the latent space to obtain missing latent variables consistent with the latent space geometry and reduce semantic shifts in the completion trajectory; using Dirichlet evidence modeling to output evidence parameters for each modality-time step; mapping confidence levels to gating weights during the fusion stage and adaptively weighting and fusing multimodal evidence; finally, introducing residual regularization and geometric regularization terms to constrain the temporal smoothness of the completion trajectory and performing pooling supervision on the sample-level effective time step set. This invention can characterize the reliability and uncertainty of each modality at different time steps, automatically suppressing the interference of unreliable modalities on decision-making and abrupt drift in temporal completion when evidence is insufficient, conflicting, or noise-enhanced, thereby improving the accuracy of missing modality completion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method, system, terminal, and computer-readable storage medium for missing modality inference based on manifold-aware diffusion completion and evidence gating fusion. Background Technology

[0002] In recent years, multimodal understanding tasks (such as dialogue emotion recognition) have been widely used in scenarios such as intelligent customer service, remote conferencing and human-computer interaction. They typically improve the accuracy and robustness of emotion recognition by fusing signals from multiple sources such as text, voice, and vision.

[0003] However, in real-world environments, multimodal data often suffers from modal missingness and time-varying reliability (fluctuations in modal quality at different times) due to sensor failures, environmental noise, network packet loss, and privacy restrictions. This leads to a significant decrease in the performance of traditional fusion models that rely on complete inputs and unstable output.

[0004] Existing methods often employ a two-stage approach of "completing first, then fusing," but this paradigm tends to treat the completion result as deterministic information, which can easily cause completion errors to propagate downstream. At the same time, it lacks explicit modeling of completion credibility and modal reliability, making it difficult to adaptively reduce the impact of unreliable modes in the time dimension.

[0005] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0006] The main objective of this invention is to provide a missing modality inference method, system, terminal, and computer-readable storage medium based on manifold-sensing diffusion completion and evidence gating fusion, aiming to solve the problem of inaccurate completion results caused by unstable error propagation and time-varying noise during missing modality completion in the prior art.

[0007] To achieve the above objectives, this invention provides a missing modality inference method based on manifold-sensing diffusion completion and evidence gating fusion. The missing modality inference method based on manifold-sensing diffusion completion and evidence gating fusion includes the following steps:

[0008] Map all observable samples from the multimodal time series input by the user to the manifold latent space to obtain the corresponding observable latent variables;

[0009] Noise is added to all the observable latent variables to obtain the corresponding diffused state variables. A denoising network is used to predict each of the diffused state variables to obtain the corresponding prediction noise. A denoising loss function is constructed using all the prediction noise.

[0010] After training the denoising network using the denoising loss function, reverse inference is performed on all missing samples in the multimodal time series samples to output the completion latent variable for each missing sample;

[0011] For each of the completed latent variables, a corresponding Dirichlet parameter and a corresponding confidence level are constructed. All the Dirichlet parameters are fused according to all the confidence levels to obtain the time step probability. The time step probability is used to perform a pooling operation on the effective time step set to obtain the class probability. A supervised loss function is constructed according to the class probability.

[0012] For each of the completed latent variables, a model residual is constructed. Based on the model residual, a residual regularization term is constructed. Constraints are applied to different completed latent variables and observable latent variables to obtain geodesic velocity amplitude and geodesic acceleration amplitude. Based on the geodesic velocity amplitude and geodesic acceleration amplitude, a geometric regularization term is constructed.

[0013] A missing modality inference model is trained based on the denoising loss function, the supervised loss function, the residual regularization term, and the geometric regularization term. Using the missing modality inference model, all missing samples are predicted based on all class probabilities, and the predicted class and final class probability of each missing sample are output.

[0014] Furthermore, to achieve the above objectives, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a missing modality inference program based on manifold-aware diffusion completion and evidence gating fusion stored in the memory and executable on the processor. When the missing modality inference program based on manifold-aware diffusion completion and evidence gating fusion is executed by the processor, it implements the steps of the missing modality inference method based on manifold-aware diffusion completion and evidence gating fusion as described above.

[0015] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a missing modality inference program based on manifold-aware diffusion completion and evidence gating fusion, wherein when the missing modality inference program based on manifold-aware diffusion completion and evidence gating fusion is executed by a processor, it implements the steps of the missing modality inference method based on manifold-aware diffusion completion and evidence gating fusion as described above.

[0016] In this invention, manifold-aware diffusion completion is performed in the latent space to obtain missing latent variables consistent with the geometry of the latent space and reduce semantic shifts in the completion trajectory. Dirichlet evidence modeling is used to output evidence parameters for each modality-time step. In the fusion stage, confidence scores are mapped to gating weights for adaptive weighted fusion of multimodal evidence. Finally, residual regularization and geometric regularization terms are introduced to constrain the temporal smoothness of the completion trajectory, and pooling supervision is performed on the sample-level effective time step set. This invention can characterize the reliability and uncertainty of each modality at different time steps, automatically suppress the interference of unreliable modalities on decision-making and the abrupt drift of temporal completion when evidence is insufficient, conflicting, or noise is amplified, thereby improving the accuracy of missing modal completion. Attached Figure Description

[0017] Figure 1 This is a flowchart of a preferred embodiment of the missing modality inference method based on manifold sensing diffusion completion and evidence gating fusion of the present invention;

[0018] Figure 2 This is a flowchart of the first stage of a preferred embodiment of the missing modality inference method based on manifold sensing diffusion completion and evidence gating fusion of the present invention;

[0019] Figure 3 This is a flowchart of the second stage of a preferred embodiment of the missing modality inference method based on manifold sensing diffusion completion and evidence gating fusion of the present invention;

[0020] Figure 4 This is a schematic diagram illustrating the constraints of model training in a preferred embodiment of the missing modality inference method based on manifold-sensing diffusion completion and evidence gating fusion of the present invention.

[0021] Figure 5 This is a comparison chart of the visual distribution of a preferred embodiment of the missing modality inference method based on manifold-sensing diffusion completion and evidence gating fusion of the present invention;

[0022] Figure 6 This is a comparison diagram verifying the effectiveness of a preferred embodiment of the missing modality inference method based on manifold sensing diffusion completion and evidence gating fusion of the present invention;

[0023] Figure 7 This is a structural diagram of a preferred embodiment of the missing modality inference system based on manifold sensing diffusion completion and evidence gating fusion of the present invention;

[0024] Figure 8 This is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0026] The goal of multimodal temporal understanding tasks (such as dialogue emotion recognition, behavior recognition, and multi-sensor recognition) is to model the semantics and states that change over time and output reliable classification results by fusing text, speech, visual, or multi-source sensor signals. However, in real-world scenarios, factors such as sensor failure, environmental noise, network packet loss, and privacy restrictions often lead to modality loss and time-varying modality quality. This makes the traditional two-stage "complete first, then fuse" approach prone to problems such as completion error propagation, temporal drift, and overconfidence, resulting in a significant decrease in inference performance and stability.

[0027] Therefore, this invention discloses a missing inference framework based on manifold-aware diffusion completion and evidence gating fusion (MD2E-MI, Multi-Dimensional Dynamic Embedding with Mutual Information), which aims to alleviate the robust inference challenge under incomplete multimodal conditions.

[0028] The missing modality inference method based on manifold-sensing diffusion completion and evidence gating fusion described in the preferred embodiment of the present invention, such as... Figure 1 As shown, the missing modality inference method based on manifold-aware diffusion completion and evidence gating fusion includes the following steps:

[0029] Step S10: Map all observable samples in the multimodal time series samples input by the user to the manifold latent space to obtain the corresponding observable latent variables.

[0030] Among them, such as Figure 2 As shown, the first step is to identify the acquired multimodal time series samples, determine which samples are observable and which are missing (i.e. missing modes), and construct an effective time step set. Then, the final supervision and average pooling operation is performed on the set to avoid the pollution of decision-making by all missing time steps.

[0031] For example, in multimodal emotion computing and human-computer interaction data, including the joint analysis of text, speech and video, in real-world scenarios, a certain modality may be temporarily missing, its quality may be degraded, or it may be affected by noise. In such cases, the model needs to fill in the missing information while judging the credibility of each modality at each time point, so as to ensure that the final classification result is more stable.

[0032] Secondly, in multi-sensor behavior recognition and state perception data, including the analysis of timing signals such as acceleration and gyroscopes in wearable devices and mobile phone sensors, common problems with this type of data include sensor disconnection, long-term missing measurements, or continuous noise pollution.

[0033] For more common data applications, including multi-source time-series monitoring and decision data, such as smart device operation monitoring, environmental perception, and health monitoring, their common characteristics are "multi-source data + time series + partial missing or faulty data + need for stable decision-making". The significance of the method disclosed in this invention in dealing with such problems is that it not only completes the data, but also explicitly models the reliability of different modalities and different time steps, reducing the risk of error completion propagating to downstream tasks.

[0034] Specifically, multimodal time-series samples input by the user are obtained, and an availability mask is constructed for each sample in the multimodal time-series samples to distinguish between all observable samples and all missing samples:

[0035] ;

[0036] ;

[0037] Where X represents a multimodal time series sample, Let M represent the m-th observable sample at time step t, where M represents the total number of modes and T represents the total number of time steps. express Availability mask, when When equal to 1, it represents the availability mask of observable samples; when... A value of 0 indicates an availability mask for missing samples;

[0038] Each observable sample is input into its corresponding encoder, and all observable samples are mapped to the manifold latent space using different encoders to obtain the corresponding observable latent variables:

[0039] ;

[0040] ;

[0041] ;

[0042] in, Let m represent the observable latent variable of the m-th observable sample at time step t. Indicates encoder, Represents the manifold latent space. This represents the set of effective time steps for mode m. Represents the set of time steps for all observable samples. Let represent the union of the first to the Mth sets.

[0043] Among them, such as Figure 2 As shown, after all observable samples are determined, in the A1 stage, all observable samples are input into the encoders dedicated to their respective modalities (for example, observable samples of the text modality are input into the text encoder) for manifold latent space encoding, so as to map all observable samples into the manifold latent space. In the embodiments disclosed in this invention, the latent space is modeled as a Riemannian manifold and equipped with basic operators such as exponential mapping, logarithmic mapping, parallel translation and geodesic distance for subsequent calculations.

[0044] It should be noted that the latent space can be selected as a spherical manifold, hyperbolic manifold, Lie group manifold, or other Riemannian manifold with defined geodesic distance and basic operators, depending on the task characteristics. As long as basic geometric operators such as exponential mapping, logarithmic mapping, parallel translation, and geodesic distance can be provided, the manifold uniform completion and trajectory regularization of the present invention can be realized.

[0045] In this approach, the location of missing samples in the multimodal time series samples is not encoded. Instead, corresponding latent variables are generated through subsequent steps such as completion and inverse diffusion to ensure that the entire time series is complete and usable in the latent space, thereby improving the consistency of samples in the latent space.

[0046] Step S20: Add noise to all the observable latent variables to obtain the corresponding diffused state variables, use a denoising network to predict each of the diffused state variables to obtain the corresponding prediction noise, and use all the prediction noise to construct a denoising loss function.

[0047] Among them, such as Figure 2 As shown, in the A2 stage, diffusion occurs on the Riemannian manifold. Geometric operators such as exponential mapping, logarithmic mapping, and parallel translation are used to achieve consistent alignment of the tangent space during the diffusion noise addition and denoising process, thereby obtaining missing latent variables consistent with the latent space geometry and reducing the semantic offset of the completed trajectory.

[0048] The latent space geometry can be approximated as Euclidean space, or, to reduce engineering complexity, the manifold back end can be degenerated into an Euclidean version. In implementation, explicit geometric operators can be omitted, and conventional vector space operations can be used instead. This approach can significantly reduce the implementation threshold and operator dependency, but its robustness may decrease in scenarios with high missing rates or strong trajectory structure constraints.

[0049] Specifically, noise is added to all the observable latent variables to obtain the corresponding diffused state variables:

[0050] ;

[0051] in, express The diffusion state variable, Indicated by The exponential mapping of the tangent point, Represents the noise amplitude function. This represents tangent space noise. Indicates diffusion time;

[0052] Each of the diffusion state variables, the diffusion time, and the full-modal context are input into the constructed denoising network for prediction, and the noise estimate of each of the observable samples in the tangent space is output:

[0053] ;

[0054] in, express Noise estimation, This represents a denoising network. Represents a mask in the full modal context. Represents the manifold points after diffusion;

[0055] The true noise of each observable sample is shifted to the corresponding manifold point of the observable sample to obtain the corresponding predicted noise:

[0056] ;

[0057] in, express Predictive noise, Indicates will translate to Popular spots express Real noise;

[0058] Construct a denoising loss function using all the predicted noise and the corresponding estimated noise:

[0059] ;

[0060] ;

[0061] in, Represents the denoising loss function. Indicates the denoising weights. Represents the gate function. This represents the activation function. Indicates hyperparameters, Indicates the difference parameter.

[0062] First, forward diffusion is performed to train the denoising network. During the training phase, the diffusion time is sampled, and noise is injected into the tangent space of the Riemannian manifold. An exponential mapping based on the observable latent variables is used to obtain the diffused state variables. The diffused state refers to the "noisy latent variable" obtained by injecting noise into the original latent variable at a certain diffusion time point during training. It resides in the same manifold space but is more "fuzzy" than the original latent variable. This noisy latent variable serves as the input to the denoising network, training it to recover or predict the corresponding noise components given noise intensity and context.

[0063] In the embodiments disclosed in this invention, the denoising network is a model used to predict or remove noise. During the training phase, the parameters are updated by continuously comparing the "predicted noise" and the "actual noise" to gradually learn to predict noise components under different noise intensities. During the inference phase, the denoising network (with fixed parameters that are no longer updated) is used to gradually restore random noise into reasonable latent variables during the reverse generation process, thereby completing the missing modal or time step information.

[0064] Based on this, such as Figure 2 As shown, in stage A3, the noise estimate in the tangent space can be predicted based on the diffused state variables, diffusion time, and full modal context, and then compared with the true noise. Since the true noise is near the manifold point where the currently observed latent variable is located, a linear space is used to approximate the local direction and displacement. This linear space contains "vectors" (such as the true noise); the manifold is a curved surface or space, and the tangent space is a local coordinate system after "flattening" the vector at a certain point, used to perform linear operations such as addition, subtraction, difference, and length calculation.

[0065] Furthermore, a translation operator is used to move the actual noise to the estimated noise location, i.e., near the "manifold point obtained after diffusion." A linear space is also used to approximate the local direction and displacement, containing "vectors" (e.g., noise estimates). Since the output of the denoising network is a noise estimate given for the "current diffusion state," it is defined in the local coordinate system of the diffusion state point. Because the actual noise and predicted noise reside in two different local coordinate systems, a consistent comparison requires first transforming the actual noise to the local coordinate system of the diffusion state, then calculating the difference between the two, thereby constructing the expression for the denoising loss function.

[0066] Furthermore, in order to adjust the learning intensity at different diffusion times and different time steps, a denoising weight parameter is added to the expression of the denoising loss function. The denoising weight parameter includes a gating function and an activation function that depend only on the time step, thereby ensuring the reproducibility of the model's predictions.

[0067] Step S30: After training the denoising network using the denoising loss function, perform reverse inference on all missing samples in the multimodal time series samples and output the completion latent variable for each missing sample.

[0068] In this context, to meet the requirement of "outputting usable latent variables," generative completion in stage A can be replaced by alternative diffusion models such as reversible flow models, variational autoencoders, and conditional generative networks. The alternative models must be able to generate a complete representation consistent with the latent space under missing conditions and support controllable sampling or uncertainty metrics. For time-latency sensitive deployments, fewer inverse diffusion steps, distilled fast generators, or single-step approximate updates can be used to significantly reduce inference time within acceptable accuracy loss. When the missing data is structural (e.g., consecutive block missing data), completion can be triggered for key time periods or key modes, while other time periods can be directly imputed using observations or lightweight interpolation, thereby reducing unnecessary computation.

[0069] Specifically, a corresponding discrete time is constructed for each missing sample based on the constructed inverse diffusion time step, and the denoising network is trained using the denoising loss function:

[0070] ,

[0071] ;

[0072] in, This represents the number of the i-th reverse diffusion step. Indicates the number of backdiffusion steps;

[0073] The back-update process for each missing sample is induced using the trained denoising network to obtain the imputation latent variables for each missing sample:

[0074] ;

[0075] ;

[0076] in, This represents the completed latent variable for the m-th missing sample after the second iteration at time step t. This indicates the reverse update process. This represents the completed latent variable for the m-th missing sample after the first iteration at time step t. Let X represent the (i-1)th reverse diffusion step number, and let X represent the multimodal time series sample. Represents a mask in the full modal context. This represents the imputation latent variable for the m-th missing sample at time step t. Let represent the completed latent variable after the i-th iteration update at time step t for the m-th missing sample.

[0077] Specifically, the reverse update process for each missing sample is induced based on the trained denoising network. This process can adopt a defined discretization rule and perform a geometrically consistent reverse update on the manifold latent space to obtain the completion latent variables for each missing sample.

[0078] Specifically, by performing manifold-aware diffusion completion in the latent space (which can be a Riemannian manifold), geometric operators such as exponential mapping, logarithmic mapping, and parallel translation are used to achieve consistent alignment of the tangent space during the diffusion noise addition and denoising process, thereby obtaining missing latent variables consistent with the latent space geometry and reducing the semantic offset of the completion trajectory.

[0079] Step S40: For each of the completed latent variables, construct the corresponding Dirichlet parameters and the corresponding confidence level. Based on all the confidence levels, fuse all the Dirichlet parameters to obtain the time step probability. Use the time step probability to perform a pooling operation on the effective time step set to obtain the class probability. Construct a supervised loss function based on the class probability.

[0080] Among them, such as Figure 3 As shown, in order to characterize the reliability and uncertainty of each modality at different time steps, this invention uses Dirichlet evidence modeling in stage B1 to output the evidence parameters of each modality-time step and calculates the confidence level based on the evidence strength; in stage B2, the confidence level is mapped to gating weights, and multimodal evidence is adaptively weighted and fused to automatically suppress the interference of unreliable modalities on decision-making when there is insufficient evidence, conflict, or increased noise; in stage B3, pooling is performed only on the set of effective time steps at the sample level to obtain the class probability; finally, in stage B4, the missing modality is predicted based on the class probability to obtain the predicted modality type and the corresponding final class probability.

[0081] In this invention, the confidence level is constructed using the total Dirichlet evidence strength. Alternatively, the "reliability or uncertainty" index can be constructed using methods such as predicted distribution entropy, ensemble model divergence, or multiple sampling consistency, but the interpretable link of "confidence level → gating weights → fusion decision" must be maintained. Besides softmax (a mathematical function that transforms any real vector into a probability distribution) gating, normalized linear gating, temperature scaling gating, threshold truncation gating, or gating with upper or lower bound constraints can also be used to meet different business preferences for "conservative or aggressive fusion." In addition to weighted summation of Dirichlet parameters, logarithmic evidence fusion, weighted product fusion, or hierarchical fusion (fusion of similar modalities first, then cross-modal fusion) can also be used, as long as adaptive weighting of unreliable modalities can be achieved and interpretable fusion results can be output.

[0082] Specifically, for each of the completed latent variables, an evidence representation and Dirichlet parameters are constructed, and a class probability estimate for each of the completed latent variables is constructed based on all the Dirichlet parameters:

[0083] ;

[0084] ;

[0085] ;

[0086] in, The evidence representation of the imputation latent variable for the m-th missing sample at the current time step. This represents the activation function. This represents a K-dimensional vector of the imputation latent variables for the m-th missing sample at the current time step, where k represents the dimension index. The Dirichlet parameter represents the imputation latent variable for the m-th missing sample at the current time step. This represents the sum of Dirichlet parameters. The mean of the Dirichlet parameters is used as an estimate of the class probability of the m-th missing sample at the current time step.

[0087] Construct a confidence level for the class probability estimate of each missing sample:

[0088] ;

[0089] in, This represents the confidence level of the m-th missing sample at the current time step.

[0090] For each time step, the confidence level for each of the missing modes is mapped to a gating weight:

[0091] ;

[0092] in, This represents the gating weight of the m-th missing sample at the current time step. Indicates the gating hyperparameters. It is an exponential function. Indicates the first Confidence level of missing samples other than the number of missing samples;

[0093] The gate weights are used to perform a weighted fusion of all the Dirichlet parameters to obtain fusion parameters, and the time step probability of the current time step is calculated based on the fusion parameters.

[0094] ;

[0095] ;

[0096] in, This represents the fusion parameters at the current time step. This represents the gating weights after fusion at the current time step. This represents the probability of the current time step.

[0097] By performing a pooling operation on the effective set of time steps using the probability of each time step, the class probabilities are obtained:

[0098] ;

[0099] in, Represents class probability, Represents the set of effective time steps;

[0100] Calculate the class probability corresponding to the missing sample with the preset real label to construct the supervised loss function:

[0101] ;

[0102] in, Represents the supervised loss function. This represents the probability of the true label being y.

[0103] In the embodiments disclosed in this invention, an evidence representation is constructed on the completed latent variables. The evidence prediction head outputs a K-dimensional vector, which is then used to obtain the evidence representation through the Softplus activation function. Dirichlet parameters (which determine the probability distribution shape of each component in the Dirichlet distribution and control the concentration and sparsity of the distribution) are then defined, and the mean of the Dirichlet parameters is used as the class probability estimate for the corresponding missing mode-time step. To measure the reliability of this class probability estimate, this invention further constructs a corresponding confidence level based on the evidence strength. The greater the evidence strength, the higher the confidence level, and the more accurate the class probability estimate.

[0104] Furthermore, for each time step, the confidence scores of each modality are mapped to the gating weights of the extended modality. The Dirichlet parameters are then weighted and fused using these gating weights, and the fused time step probabilities are provided by the mean of the Dirichlet parameters. Finally, average pooling is performed on the set of effective time steps to obtain the final class probabilities.

[0105] The model first outputs a classification result at each "effective time step" (i.e., a time point where at least one modality is available and not contaminated by complete missing values), obtaining the predicted probability for each class at each time step. Then, the predicted probabilities from these effective time steps are averaged to obtain a "sample-level" final predicted probability distribution. Next, the probability value corresponding to the true label's class is extracted from this final class probability distribution as a measure of "how much the model trusts the true class." The supervised loss is defined as the negative logarithm of this probability: the higher the true class probability, the smaller the loss; the lower the true class probability, the larger the loss. Minimizing this loss during training is equivalent to continuously pushing the model to increase the final probability of the true class, thereby making the classification more accurate and stable, while avoiding interference from invalid (completely missing) time steps in the decision-making process.

[0106] Step S50: Construct model residuals for each of the completed latent variables, construct residual regularization terms based on the model residuals, constrain different completed latent variables and observable latent variables to obtain geodesic velocity amplitude and geodesic acceleration amplitude, and construct geometric regularization terms based on the geodesic velocity amplitude and the geodesic acceleration amplitude.

[0107] Among them, such as Figure 4 As shown, this invention introduces residual regularization terms and geometric velocity and acceleration regularization constraints to complete the temporal smoothness of the trajectory, and performs pooling supervision on the set of effective time steps at the sample level to improve stability and auditability in missing scenarios.

[0108] In addition to residual regularization terms and geometric velocity and acceleration constraints, more general time smoothing regularizations (such as first-order and second-order difference smoothing or geodesic distance-based smoothing terms) can be introduced to adapt to different sampling frequencies and sequence lengths. The strength of trajectory regularization can be dynamically adjusted by confidence level or gating weights to make "reliable segments more constrained and low-reliability segments generated more conservatively", thereby further suppressing abrupt drift of noisy segments. For sequence boundary positions, one-sided difference, mirror extension, or fixed boundary strategies can be used to ensure that the regularization term calculation is stable and does not introduce additional bias.

[0109] Specifically, across the entire multimodal time series sample dimension, for each current time step, the residual for the current time step is constructed based on the completed latent variables or observable latent variables of the previous time step and the completed latent variables or observable latent variables of the next time step, and a residual regularization term is constructed based on the residual:

[0110] ;

[0111] ;

[0112] in, denotes the residual of the m-th missing sample at the t-th time step, denotes the logarithmic mapping with as the base point, denotes the completed latent variable of the m-th missing sample at the t-th time step, denotes the completed latent variable of the m-th missing sample at the (t + 1)-th time step, denotes the completed latent variable of the m-th missing sample at the (t - 1)-th time step, denotes the residual regularization term, denotes the set of valid time steps, and T denotes the total number of time steps;

[0113] Calculate the geodesic velocity magnitude and the geodesic acceleration magnitude to constrain the completed latent variable and the observable latent variable:

[0114] ;

[0115] ;

[0116] where, denotes the geodesic velocity magnitude of the m-th missing sample at the t-th time step, denotes the constraint process, denotes the geodesic acceleration magnitude of the m-th missing sample at the t-th time step;

[0117] Construct a geometric regularization term according to the geodesic velocity magnitude and the geodesic acceleration magnitude:

[0118] ;

[0119] where, denotes the geometric regularization term.

[0120] Among them, in order to suppress the sudden drift of the completed latent trajectory in the time dimension, a residual regularization term is constructed on the completed manifold trajectory of the present invention. For any sample and time step satisfying 2 < t < T - 1, the corresponding residual is defined, and then the residual regularization term is constructed by using the logarithmic mapping operator.

[0121] Among them, the residual means "taking the current time step as the base point, mapping the manifold displacements of the next step and the previous step into the same tangent space, and then taking the difference", which can measure whether the completed manifold trajectory is smooth in the time dimension and whether there is sudden drift. If the trajectory is smooth, the displacements before and after should be close and the difference is small; if there is a jump or inconsistency, the difference is large. This residual is squared and accumulated over the valid time steps to form a residual regularization term, which is used to constrain the temporal consistency of the completed trajectory.

[0122] Furthermore, in the embodiments disclosed in this invention, constraints are simultaneously applied to geodesic velocity and geodesic acceleration, which can construct geodesic velocity amplitude and geodesic acceleration amplitude respectively, thereby defining a geometric regularization term to suppress time drift.

[0123] Step S60: Train a missing modality inference model based on the denoising loss function, the supervised loss function, the residual regularization term, and the geometric regularization term. Using the missing modality inference model, predict all the missing samples based on all the class probabilities, and output the predicted class and final class probability of each missing sample.

[0124] This invention achieves robust inference for incomplete multimodal inputs through a unified scheme of latent space consistent diffusion completion, evidence and confidence-driven gated fusion, and temporal trajectory regularization constraints. It can effectively reduce the risk of error propagation and improve prediction accuracy, stability, and reliability under conditions such as high missing rate, continuous block missing, and strong noise.

[0125] During the training phase, a hybrid strategy of random missing data and continuous block missing data can be adopted to cover the missing data patterns in real-world scenarios and improve deployment robustness. Alternatively, training can be phased, first training the generative completion module to obtain stable completion capabilities, and then jointly training the evidence gating fusion module for end-to-end optimization. This approach is suitable for scenarios with large data scales, unstable training, or clear division of labor in engineering iterations. During deployment, the system can output gating weights, confidence curves over time, and evidence strength for key time periods for online monitoring, alarms, and manual review, thereby improving system controllability and interpretability.

[0126] Specifically, corresponding weight hyperparameters are constructed for the denoising loss function, the residual regularization term, and the geometric regularization term, respectively, and a target loss function is constructed based on the denoising loss function, the supervised loss function, the residual regularization term, and the geometric regularization term:

[0127] ;

[0128] in, Represents the target loss function. Represents the supervised loss function. express The weight hyperparameters, Represents the denoising loss function. express The weight hyperparameters, express The weight hyperparameters;

[0129] The constructed missing modality inference model is trained end-to-end using the target loss function, and all the missing samples are re-predicted according to the probabilities of all the categories, outputting the predicted category and the corresponding final category probability of each missing sample.

[0130] In the embodiments disclosed in this invention, a target loss function is constructed from a denoising loss function, a supervised loss function, a residual regularization term, and a geometric regularization term. Corresponding weight hyperparameters are added to the denoising loss function, residual regularization term, and geometric regularization term. An end-to-end training method is adopted. During training, the denoising loss function is calculated for the observation location to learn the denoising network. During inference, inverse diffusion is performed on the missing location to obtain the completion latent variables of the missing sample at the current time step. Then, evidence, gating, fusion, and final prediction are calculated for all stages B. For any input sample, the model outputs the predicted category and its corresponding final category probability, and simultaneously outputs the gating weights and confidence scores for each time step, explaining which modal evidence the model relies on for decision-making at different time periods, thereby improving system auditability and engineering controllability.

[0131] Furthermore, to verify the effectiveness and robustness of the proposed method (MD2E-MI), this invention conducted extensive experiments on two typical tasks:

[0132] (1) Experiments were conducted on multimodal dialogue and sentiment benchmark datasets (CMU-MOSEI, Multimodal Viewpoint Sentiment and Emotion Intensity Dataset; CMU-MOSI, Emotion Intensity Corpus) to evaluate the degradation of classification performance under different missing rates (MR);

[0133] (2) Experiments were conducted on physical sensor time series datasets (UCI-HAR, Human Activity Recognition Dataset; PAMAP2, Physical Activity Monitoring Dataset, Second Edition) to evaluate the stability and generalization ability under long-term continuous missing data (R2 Block Failure, Miss (missing rate) = 0.8).

[0134] The sentiment dataset was evaluated at different missing rates (0, 0.3, 0.5, 0.7, and 0.9), with the metrics being Acc-2 (binary classification accuracy), F1-Score (F1 score), and Acc-7 (seven-class classification accuracy). The sensor dataset reported Acc (accuracy) and Macro-F1 (macro-average F1 score), and provided the mean ± standard deviation of three random seeds.

[0135] Based on the above experimental conditions, in the experiments disclosed in this invention, this method was compared with several other methods: (1) For multimodal representation learning experiments, the comparison methods included MMIN (Missing Modality Imagination Network) and GCNet (Graph Complete Network); (2) For explicit generation and imputation enhancement experiments, the comparison method included IMDer (Implicit Multimodal Dynamics Estimator); (3) For graph structure reasoning experiments, the comparison method included SDR-GNN (Spectral Domain Reconstruction Graph Neural Network); and the performance was compared under the same missing protocol. The comparison results are shown in Table 1.

[0136] Table 1: Performance Comparison Table

[0137]

[0138] In Table 1, the data obtained by each method under different MR values ​​represent Acc-2, F1-Score, and Acc-7, respectively. Drop@0.9 represents the relative performance degradation from MR=0.0 to MR=0.9. As shown in Table 1, the performance degradation of this method is slower as the missing rate increases, especially maintaining a significant advantage under the extreme missing rate scenario of MR=0.9: Acc-2 reaches 77.1% on CMU-MOSEI and 71.8% on CMU-MOSI. Furthermore, Extreme Drop@0.9 (Extreme emphasizes the most severe performance drop, not the average fluctuation) is significantly smaller than the comparative methods, indicating that this invention is more robust to extreme missing rates.

[0139] Furthermore, to simulate long-term sensor failures that more closely resemble real-world scenarios, this invention employs the R2 BlockFailure protocol (Miss=0.8, continuous missing values ​​covering 80% of time steps) for evaluation on UCI-HAR and PAMAP2. The results are shown in Table 2.

[0140] Table 2: Comparison of Results

[0141]

[0142] In Table 2, the data obtained by each method in different datasets and with different loss rates represent Acc and Macro-F1, respectively. As can be seen from Table 2, the present invention also maintains the best results under long block missing (Miss=0.8) and has a smaller standard deviation, indicating that the present invention has more stable recovery and discrimination capabilities under the realistic conditions of severe continuous missing.

[0143] Furthermore, to verify the contribution of each core module to the overall performance, this invention conducted a step-by-step ablation experiment on CMU-MOSEI (missing rate MR=0.8), removing the imputation mechanism (using zero-filling), the manifold diffusion back end (degenerating to Euclidean diffusion), evidence gating (uniform fusion), and residual geometric regularization, and comparing the results with the complete model. The results are shown in Table 3.

[0144] Table 3: Ablation Experiment Results

[0145]

[0146] Among them, removing the imputation mechanism will lead to a significant performance drop, indicating that explicit recovery is required under high missing values; removing the manifold back end, residuals and geometric regularization will significantly weaken the performance, indicating that maintaining manifold consistency and temporal physical smoothness is a necessary condition during the recovery process; removing evidence gating will reduce the performance, indicating that evidence-based reliability weighted fusion can effectively suppress the negative impact of low-quality recovery modes on the final prediction.

[0147] To intuitively analyze the quality of latent representations learned by different methods under missing conditions, this invention further visualizes and compares the feature spaces; for example... Figure 5 As shown, Figure 5 (a) describes the embedding distribution results of SDR-GNN. Figure 5 (b) in the figure describes the embedding distribution results of this method; Figure 5 In (a), the category point clouds show obvious overlap and blurred boundaries; Figure 5 In (b), the three sample clusters are more compact and have larger intervals, significantly improving inter-class separation. Therefore, the baseline method shows significant cluster overlap, while the present invention has more compact intra-class clusters and larger inter-class intervals, indicating that "manifold consistent completion and evidence gating fusion" can learn more stable and discriminative representations in missing scenarios.

[0148] Furthermore, in addition to overall accuracy, this invention also evaluates three key engineering properties: robust degradation patterns with increasing missing rates, whether gating weights can adaptively adjust over time, and whether confidence levels can be used for reliable decision-making; such as Figure 6 As shown in (a) of the figure, the performance curves at different missing rates are presented. With increasing missing rate, the accuracy of MD2E-MI decreases the slowest, exhibiting the best robustness. Figure 6As shown in (b) of the diagram, the gating changes dynamically over time and reduces the contribution of damaged modes in the missing interval. The gating weights are dynamically adjusted over time, automatically reducing the contribution of damaged modes in the missing interval. Figure 6 As shown in (c), the correlation between completion error and confidence level was verified. Confidence level and completion error showed a significant negative correlation, indicating that uncertainty is calibrable and can be used for reliable decision-making.

[0149] This invention can characterize the reliability and uncertainty of each modality at different time steps, and automatically suppress the interference of unreliable modalities on decision-making and the abrupt drift of time sequence completion when there is insufficient evidence, conflict or increased noise, thereby improving the accuracy of missing modal completion.

[0150] Furthermore, such as Figure 7 As shown, based on the above-mentioned missing modality inference method based on manifold-aware diffusion completion and evidence gating fusion, the present invention also provides a missing modality inference system based on manifold-aware diffusion completion and evidence gating fusion, wherein the missing modality inference system based on manifold-aware diffusion completion and evidence gating fusion includes:

[0151] Modality definition module 51 is used to map all observable samples in the multimodal time series samples input by the user to the manifold latent space to obtain the corresponding observable latent variables;

[0152] The perception diffusion module 52 is used to add noise to all the observable latent variables to obtain the corresponding diffused state variables, use a denoising network to predict each of the diffused state variables to obtain the corresponding prediction noise, and use all the prediction noise to construct a denoising loss function.

[0153] The completion module 53 is used to train the denoising network using the denoising loss function, and then perform reverse inference on all missing samples in the multimodal time series samples to output the completion latent variable for each missing sample.

[0154] The evidence prediction module 54 is used to construct a corresponding Dirichlet parameter and a corresponding confidence level for each of the completed latent variables, fuse all the Dirichlet parameters according to all the confidence levels to obtain the time step probability, perform a pooling operation on the effective time step set using the time step probability to obtain the class probability, and construct a supervised loss function based on the class probability.

[0155] The time regularization module 55 is used to construct model residuals for each of the completed latent variables, construct residual regularization terms based on the model residuals, constrain different completed latent variables and observable latent variables to obtain geodesic velocity amplitude and geodesic acceleration amplitude, and construct geometric regularization terms based on the geodesic velocity amplitude and geodesic acceleration amplitude.

[0156] The prediction module 56 is used to train a missing modality inference model based on the denoising loss function, the supervised loss function, the residual regularization term and the geometric regularization term, and use the missing modality inference model to predict all the missing samples based on all the class probabilities, and output the predicted class and final class probability of each missing sample.

[0157] Furthermore, such as Figure 8 As shown, based on the above-mentioned missing modality inference method and system based on manifold-aware diffusion completion and evidence gating fusion, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 8 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0158] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code installed on the terminal. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores a missing modality inference program 40 based on manifold-aware diffusion completion and evidence gating fusion. This missing modality inference program 40 can be executed by the processor 10, thereby implementing the missing modality inference method based on manifold-aware diffusion completion and evidence gating fusion in this application.

[0159] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the missing modality inference method based on manifold-aware diffusion completion and evidence gating fusion.

[0160] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The components of the terminal communicate with each other via a system bus.

[0161] In one embodiment, when the processor 10 executes the missing modality inference program 40 based on manifold-aware diffusion completion and evidence gating fusion in the memory 20, it implements the steps of the missing modality inference method based on manifold-aware diffusion completion and evidence gating fusion as described above.

[0162] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a missing modality inference program based on manifold-aware diffusion completion and evidence gating fusion, wherein when the missing modality inference program based on manifold-aware diffusion completion and evidence gating fusion is executed by a processor, it implements the steps of the missing modality inference method based on manifold-aware diffusion completion and evidence gating fusion as described above.

[0163] In summary, this invention provides a missing modality inference method and related equipment based on manifold-aware diffusion completion and evidence gating fusion. The method includes: performing manifold-aware diffusion completion in the latent space to obtain missing latent variables consistent with the geometry of the latent space and reduce semantic shifts in the completion trajectory; using Dirichlet evidence modeling to output evidence parameters for each modality-time step; mapping confidence levels to gating weights during the fusion stage and performing adaptive weighted fusion of multimodal evidence; finally, introducing residual regularization and geometric regularization terms to constrain the temporal smoothness of the completion trajectory and performing pooling supervision on the sample-level effective time step set. This invention can characterize the reliability and uncertainty of each modality at different time steps, automatically suppress the interference of unreliable modalities on decision-making and the abrupt drift of temporal completion when evidence is insufficient, conflicting, or noise-enhanced, thereby improving the accuracy of missing modality completion.

[0164] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes that element.

[0165] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The computer-readable storage medium can be a memory, magnetic disk, optical disk, etc.

[0166] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A missing modality inference method based on manifold-sensing diffusion completion and evidence gating fusion, characterized in that, The missing modality inference method based on manifold-aware diffusion completion and evidence gating fusion includes: Map all observable samples from the multimodal sentiment time series samples input by the user to the manifold latent space to obtain the corresponding observable latent variables; Noise is added to all the observable latent variables to obtain the corresponding diffused state variables. A denoising network is used to predict each of the diffused state variables to obtain the corresponding prediction noise. A denoising loss function is constructed using all the prediction noise. After training the denoising network using the denoising loss function, reverse inference is performed on all missing sentiment samples in the multimodal sentiment time series samples, and the completion latent variable for each missing sentiment sample is output. For each of the completed latent variables, a corresponding Dirichlet parameter and a corresponding confidence level are constructed. All the Dirichlet parameters are fused according to all the confidence levels to obtain the time step probability. The time step probability is used to perform a pooling operation on the effective time step set to obtain the class probability. A supervised loss function is constructed according to the class probability. For each of the completed latent variables, a model residual is constructed. Based on the model residual, a residual regularization term is constructed. Constraints are applied to different completed latent variables and observable latent variables to obtain geodesic velocity amplitude and geodesic acceleration amplitude. Based on the geodesic velocity amplitude and geodesic acceleration amplitude, a geometric regularization term is constructed. A missing modality inference model is trained based on the denoising loss function, the supervised loss function, the residual regularization term, and the geometric regularization term. Using the missing modality inference model, all missing sentiment samples are predicted based on all the category probabilities, and the predicted category and final category probability of each missing sentiment sample are output.

2. The missing modality inference method based on manifold-sensing diffusion completion and evidence gating fusion as described in claim 1, characterized in that, The process of mapping all observable samples from the multimodal sentiment time-series samples input by the user to the manifold latent space to obtain the corresponding observable latent variables specifically includes: Obtain multimodal sentiment time-series samples input by the user, and construct an availability mask for each sample in the multimodal sentiment time-series samples to distinguish all observable samples from all missing sentiment samples: ; ; Where X represents a multimodal sentiment time-series sample, Let M represent the m-th observable sample at time step t, where M represents the total number of modes and T represents the total number of time steps. express Availability mask, when When equal to 1, it represents the availability mask of observable samples; when... A value of 0 indicates an availability mask for missing sentiment samples; Each observable sample is input into its corresponding encoder, and all observable samples are mapped to the manifold latent space using different encoders to obtain the corresponding observable latent variables: ; ; ; in, Let m represent the observable latent variable of the m-th observable sample at time step t. Indicates encoder, Represents the manifold latent space. This represents the set of effective time steps for mode m. Represents the set of time steps for all observable samples. Let represent the union of the first to the Mth sets.

3. The missing modality inference method based on manifold-sensing diffusion completion and evidence gating fusion as described in claim 2, characterized in that, The process involves adding noise to all the observable latent variables to obtain corresponding diffused state variables, using a denoising network to predict each diffused state variable to obtain corresponding prediction noise, and constructing a denoising loss function using all the prediction noise. Specifically, this includes: Noise is added to all the observable latent variables to obtain the corresponding diffused state variables: ; in, express The diffusion state variable, Indicates The exponential mapping of the tangent point, Represents the noise amplitude function. This represents tangent space noise. Indicates diffusion time; Each of the diffusion state variables, the diffusion time, and the full-modal context are input into the constructed denoising network for prediction, and the noise estimate of each of the observable samples in the tangent space is output: ; in, express Noise estimation, This represents a denoising network. Represents a mask in the full modal context. Represents the manifold points after diffusion; The true noise of each observable sample is shifted to the corresponding manifold point of the observable sample to obtain the corresponding predicted noise: ; in, express Predictive noise, Indicates will Translate to Popular spots express Real noise; Construct a denoising loss function using all the predicted noise and the corresponding estimated noise: ; ; in, Represents the denoising loss function. Indicates the denoising weights. Represents the gate function. This represents the activation function. Indicates hyperparameters, Indicates the difference parameter.

4. The missing modality inference method based on manifold-sensing diffusion completion and evidence gating fusion as described in claim 1, characterized in that, After training the denoising network using the denoising loss function, back-inference is performed on all missing sentiment samples in the multimodal sentiment time series samples to output the completion latent variable for each missing sentiment sample, specifically including: Based on the constructed inverse diffusion time step, a corresponding discrete time is constructed for each of the missing sentiment samples, and the denoising network is trained using the denoising loss function: , ; in, This represents the number of the i-th reverse diffusion step. Indicates the number of backdiffusion steps; The back-update process for each missing sentiment sample is induced using the trained denoising network to obtain the completion latent variables for each missing sentiment sample: ; ; in, This represents the completed latent variable after the second iteration update at time step t for the m-th missing sentiment sample. This indicates the reverse update process. This represents the completed latent variable after the first iteration update at time step t for the m-th missing sentiment sample. Let X represent the (i-1)th reverse diffusion step, and let X represent the multimodal sentiment time series sample. Represents a mask in the full modal context. This represents the latent variable for completing the m-th missing sentiment sample at time step t. Let represent the complete latent variable of the m-th missing sentiment sample after the i-th iteration update at time step t.

5. The missing modality inference method based on manifold-sensing diffusion completion and evidence gating fusion according to claim 2, characterized in that, For each completed latent variable, a corresponding Dirichlet parameter and a corresponding confidence level are constructed. All Dirichlet parameters are fused based on all confidence levels to obtain a time-step probability. This time-step probability is then used to perform pooling operations on the effective time-step set to obtain a class probability. Finally, a supervised loss function is constructed based on the class probabilities, specifically including: For each of the completed latent variables, construct an evidence representation and Dirichlet parameters, and based on all the Dirichlet parameters, construct a class probability estimate for each of the completed latent variables: ; ; ; in, Let f(m) represent the evidence for the completion of the latent variables for the m-th missing sentiment sample at the current time step. This represents the activation function. Let k represent the K-dimensional vector of the imputation latent variables for the m-th missing sentiment sample at the current time step, where k represents the dimension index. Let represent the Dirichlet parameter of the imputation latent variable for the m-th missing sentiment sample at the current time step. This represents the sum of Dirichlet parameters. The mean of the Dirichlet parameters is used as an estimate of the class probability of the m-th missing sentiment sample at the current time step. For each missing sentiment sample, construct a corresponding confidence level based on the category probability estimate: ; in, This represents the confidence level of the m-th missing sentiment sample at the current time step. For each time step, the confidence level for each of the missing modes is mapped to a gating weight: ; in, This represents the gating weight of the m-th missing sentiment sample at the current time step. Indicates the gating hyperparameters. It is an exponential function. Indicates the first Confidence of missing sentiment samples other than the missing sentiment samples; The gate weights are used to perform a weighted fusion of all the Dirichlet parameters to obtain fusion parameters, and the time step probability of the current time step is calculated based on the fusion parameters. ; ; in, This represents the fusion parameters at the current time step. This represents the gating weights after fusion at the current time step. This represents the probability of the current time step. By performing a pooling operation on the effective set of time steps using the probability of each time step, the class probabilities are obtained: ; in, Represents class probability, Represents the set of effective time steps; Calculate the category probability corresponding to the missing sentiment sample with the preset real label to construct the supervised loss function: ; in, Represents the supervised loss function. This represents the probability of the true label being y.

6. The missing modality inference method based on manifold-sensing diffusion completion and evidence gating fusion according to claim 2, characterized in that, The process of constructing model residuals for each completed latent variable, constructing residual regularization terms based on the model residuals, constraining different completed latent variables and observable latent variables to obtain geodesic velocity amplitudes and geodesic acceleration amplitudes, and constructing geometric regularization terms based on the geodesic velocity amplitudes and geodesic acceleration amplitudes specifically includes: Across the entire multimodal sentiment time-series sample, for each current time step, the residual for the current time step is constructed based on the completed latent variables or observable latent variables of the previous and subsequent time steps, and a residual regularization term is constructed based on the residual: ; ; in, This represents the residual of the m-th missing sentiment sample at time step t. Indicated by Logarithmic mapping with base point, This represents the latent variable for completing the m-th missing sentiment sample at time step t. This represents the latent variable for completing the m-th missing sentiment sample at time step t+1. This represents the latent variable for completing the m-th missing sentiment sample at time step t-1. Represents the residual regularization term. This represents the set of valid time steps, where T represents the total number of time steps. Calculate the geodesic velocity amplitude and geodesic acceleration amplitude to constrain the completed latent variables and the observable latent variables: ; ; in, This represents the geodesic velocity amplitude of the m-th missing sentiment sample at time step t. Represents the constraint process. This represents the geodesic acceleration amplitude of the m-th missing sentiment sample at time step t; Construct a geometric regularization term based on the geodesic velocity amplitude and the geodesic acceleration amplitude: ; in, This represents a geometric regularity term.

7. The missing modality inference method based on manifold-sensing diffusion completion and evidence gating fusion as described in claim 6, characterized in that, The missing modality inference model is trained based on the denoising loss function, the supervised loss function, the residual regularization term, and the geometric regularization term. Using this model, all missing sentiment samples are predicted based on all class probabilities, and the predicted class and final class probability of each missing sentiment sample are output. Specifically, this includes: Construct corresponding weight hyperparameters for the denoising loss function, the residual regularization term, and the geometric regularization term, respectively, and construct a target loss function based on the denoising loss function, the supervised loss function, the residual regularization term, and the geometric regularization term: ; in, Describes the target loss function. Represents the supervised loss function. express The weight hyperparameters, Represents the denoising loss function. express The weight hyperparameters, express The weight hyperparameters; The constructed missing modality inference model is trained end-to-end using the target loss function, and all the missing sentiment samples are re-predicted according to the probabilities of all the categories, outputting the predicted category and the corresponding final category probability of each missing sentiment sample.

8. A missing modality inference system based on manifold-aware diffusion completion and evidence gating fusion, characterized in that, The missing modality inference system based on manifold-aware diffusion completion and evidence gating fusion is used to implement the missing modality inference method based on manifold-aware diffusion completion and evidence gating fusion as described in any one of claims 1-7. The missing modality inference system based on manifold-aware diffusion completion and evidence gating fusion includes: The modality definition module is used to map all observable samples from the multimodal sentiment time series samples input by the user to the manifold latent space to obtain the corresponding observable latent variables; The perceptual diffusion module is used to add noise to all the observable latent variables to obtain the corresponding diffused state variables, use a denoising network to predict each of the diffused state variables to obtain the corresponding prediction noise, and use all the prediction noise to construct a denoising loss function. The completion module is used to train the denoising network using the denoising loss function, and then perform reverse inference on all missing sentiment samples in the multimodal sentiment time series samples, outputting the completion latent variable for each missing sentiment sample; The evidence prediction module is used to construct corresponding Dirichlet parameters and corresponding confidence levels for each of the completed latent variables, fuse all the Dirichlet parameters according to all the confidence levels to obtain time step probabilities, perform pooling operations on the effective time step set using the time step probabilities to obtain class probabilities, and construct a supervised loss function based on the class probabilities. The time regularization module is used to construct model residuals for each of the completed latent variables, construct residual regularization terms based on the model residuals, constrain different completed latent variables and observable latent variables to obtain geodesic velocity amplitude and geodesic acceleration amplitude, and construct geometric regularization terms based on the geodesic velocity amplitude and geodesic acceleration amplitude. The prediction module is used to train a missing modality inference model based on the denoising loss function, the supervised loss function, the residual regularization term, and the geometric regularization term. Using the missing modality inference model, it predicts all the missing sentiment samples based on all the category probabilities and outputs the predicted category and final category probability of each missing sentiment sample.

9. A terminal, characterized in that, The terminal includes: a memory, a processor, and a missing modality inference program based on manifold-aware diffusion completion and evidence gating fusion stored in the memory and executable on the processor. When the missing modality inference program based on manifold-aware diffusion completion and evidence gating fusion is executed by the processor, it implements the steps of the missing modality inference method based on manifold-aware diffusion completion and evidence gating fusion as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a missing modality inference program based on manifold-aware diffusion completion and evidence gating fusion. When the missing modality inference program based on manifold-aware diffusion completion and evidence gating fusion is executed by a processor, it implements the steps of the missing modality inference method based on manifold-aware diffusion completion and evidence gating fusion as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Multimodal emotion recognition method and system based on hypergraph diffusion and evidence fusion, terminal and storage medium

    CN121009512A

  • Multi-modal fusion method based on dual uncertainty of evidence

    CN121327756A