Multi-view tracking method and system based on diffusion clustering, terminal and storage medium

Through a diffusion clustering-based method, the incomplete multi-view data is processed using the Dirichlet diffusion module and the adaptive Kalman filter, which solves the problems of missing and noise observation, and realizes efficient clustering performance and information integration in missing views.

CN120356137AActive Publication Date: 2025-07-22SHENZHEN MSU-BIT UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510847211.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

When the prior art deals with incomplete multi-view clustering, it is difficult to effectively process missing and noise observation data, resulting in inaccurate processing results or modal missing.

Method used

Using a diffusion clustering-based method, iterative forward diffusion is performed through the Dirichlet diffusion module, combining the adaptive Kalman filter and the dynamic evidence fusion module, the uncertainty is directly modeled and processed, avoiding the explicit reconstruction of missing views, and diffusion training is performed using Dirichlet distribution and Sharma-Mittal divergence.

Benefits of technology

It has achieved improvements in clustering performance under missing and noise observations, and can effectively integrate multi-source information, maintain the accuracy and robustness of clustering results, especially in the absence of up to 70% of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356137A_ABST
    Figure CN120356137A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and discloses a multi-view tracking method and system based on diffusion clustering, a terminal and a storage medium, and the method comprises the steps: converting a feature into a specific feature representation of a view through a feature coding module, inputting the feature representation into a Dirichlet diffusion module, carrying out the multiple times of reverse diffusion, obtaining refined Dirichlet parameters, and carrying out the multi-view tracking through the Dirichlet diffusion module. And converting the information into belief quality, inputting the belief quality into a dynamic evidence fusion module for synchronous fusion, and finally converting the fused information into a stable clustering distribution result. The invention provides a robust and interpolation-free multi-view clustering and tracking solution capable of directly modeling and processing uncertainty, an end-to-end deep clustering framework is constructed, the problems of data missing, noise propagation and insufficient uncertainty quantization in the prior art are solved, the propagation of reconstruction noise is avoided, and the reconstruction efficiency is improved. The clustering performance under missing and noise observation is enhanced, and multi-source information can be effectively integrated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to a multi-view tracking method, system, terminal and computer-readable storage medium based on diffusion clustering. Background Art

[0002] The goal of multi-view clustering (MVC) is to partition this multi-view data into semantically consistent groups without external supervision. In this data, each object is usually described by multiple complementary modalities such as images, point clouds, text, or audio, which has indispensable application value in recommendation systems, multimedia retrieval, and bioinformatics systems.

[0003] However, in actual deployment, the assumption that each object can be completely observed in each view rarely holds. Problems such as sensor failures, privacy filtering, and network delays often lead to the absence of certain parts of the multi-view matrix, making it difficult to directly apply traditional MVC methods. This practical but challenging setting is formalized as the incomplete multi-view clustering (IMVC) problem, resulting in inaccurate processing results or missing modalities when dealing with missing and noisy observation data in multi-views.

[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0005] The main purpose of the present invention is to provide a multi-view tracking method, system, terminal and computer-readable storage medium based on diffusion clustering, aiming to solve the problem that in the existing technology, IMVC has inaccurate or missing data when dealing with missing and noisy observation data.

[0006] To achieve the above purpose, the present invention provides a multi-view tracking method based on diffusion clustering. The multi-view tracking method based on diffusion clustering includes the following steps: Obtain the original data corresponding to multiple views, map all the original data to the latent space to obtain multiple feature representations, and project the multiple feature representations into multiple initial view evidences; Construct a Dirichlet diffusion module according to all the initial view evidences, perform iterative forward diffusion through the Dirichlet diffusion module to obtain the cumulative diffusion rate, and perform backward diffusion according to the cumulative diffusion rate to output the Dirichlet prediction parameters of the backward transformation; Obtain the binary indication of the real label set, construct the UPCE loss function according to the Dirichlet prediction parameter and the binary indication, and optimize the constructed dynamic evidence fusion module through the UPCE loss function; Map multiple pieces of the initial view evidence into multiple belief masses, input all the belief masses into the dynamic evidence fusion module, and output a belief vector; Input the belief vector into an adaptive Kalman filter, and output the final clustering assignment results of all the views.

[0007] Optionally, for the multi-view tracking method based on diffusion clustering, wherein, obtaining the original data corresponding to multiple views, mapping all the original data into a latent space to obtain multiple feature representations, and projecting the multiple feature representations into multiple pieces of initial view evidence, specifically includes: Obtain the original data corresponding to multiple views, and respectively input all the original data into multiple feature encoders to obtain the hidden vector corresponding to each piece of the original data: ; Wherein, represents the index of the view, represents the -th hidden vector of the view, represents the -th feature encoder of the view, represents the -th original data of the view, represents the set of real numbers, represents the -th dimension of the view, represents the -th -dimensional set of real numbers of the view; Define the affine transformation of each view, and perform batch normalization processing on all the hidden vectors according to each affine transformation to obtain the feature representation corresponding to each piece of the original data in the latent space: ; Wherein, represents the -th feature representation of the view, represents the -th affine transformation of the view, represents batch normalization processing, represents the -th shared projection of the view, represents the -th translation vector of the view, represents -dimensional set of real numbers, Denote the dimension of; Map all the said feature representations to initial view evidences corresponding to each said view, and output: ; wherein, denotes the initial view evidence of the -th view, denotes the smooth non-linear activation function, denotes the shared projection, denotes the translation vector, denotes the constant.

[0008] Optionally, for the multi-view tracking method based on diffusion clustering, wherein, according to all the said initial view evidences, construct a Dirichlet diffusion module, and perform iterative forward diffusion through the said Dirichlet diffusion module to obtain the cumulative diffusion rate, and perform backward diffusion according to the said cumulative diffusion rate to output the Dirichlet prediction parameters of the backward transformation, specifically including: Perform normalization processing on each said initial view evidence to obtain a clustering assignment vector corresponding to each said view, and model each said clustering assignment vector as a multi-dimensional probability simplex, and use the said multi-dimensional probability simplex as a Dirichlet random variable to construct a Dirichlet diffusion module: ; ; ; wherein, denotes the initial view evidence of the -th view, denotes the concentration parameter of the first dimension, denotes the concentration parameter of the -th dimension, denotes the transpose, denotes the -th clustering probability, denotes the total evidence of the -th dimension, denotes the probability simplex of the denotes the normalized clustering probability, denotes the real number set of the denotes the -th dimension of the clustering probability; Construct a diffusion schedule, perform iterative forward diffusion on the Dirichlet diffusion module according to the diffusion schedule to obtain a cumulative diffusion rate, construct a single-step posterior distribution based on the cumulative diffusion rate, and predict Dirichlet prediction parameters through a reverse denoising network according to the single-step posterior distribution.

[0009] Optionally, in the multi-view tracking method based on diffusion clustering, the construction of the diffusion schedule, performing iterative forward diffusion on the Dirichlet diffusion module according to the diffusion schedule to obtain a cumulative diffusion rate, constructing a single-step posterior distribution based on the cumulative diffusion rate, and predicting Dirichlet prediction parameters through a reverse denoising network specifically includes: Construct a diffusion schedule and perform multiple forward diffusions according to the diffusion schedule: ; Among them, the diffusion schedule is , represents the total discrete time, represents the index of discrete time, represents the forward diffusion kernel of the th iteration step, represents the index of the iteration step number, represents the th iteration step's single-step size, represents the th iteration step's final input data, represents the th iteration step's input data, represents the Dirichlet diffusion module, represents the th forward process, represents the pure noise center; Construct a cumulative diffusion rate according to the single-step size in each forward diffusion, construct a posterior scaling factor, and construct a single-step posterior distribution according to the posterior scaling factor: ; ; ; Among them, represents the th step's cumulative diffusion rate, represents the th step's cumulative diffusion rate, represents 's index, represents the th single-step size, Represents the posterior scaling factor, Represents the single-step posterior distribution, Represents the initial input data, Represents the original view evidence; Determine the final input data according to the single-step posterior distribution, input the final input data into the time conditional encoder in the Dirichlet diffusion module, output the Dirichlet prediction parameters, and input the prediction parameters into the reverse denoising network to output the reverse denoised Dirichlet distribution: ; ; Among them, Represents the final input data of the step iteration, Represents the Dirichlet distribution of the step iteration, Represents the Dirichlet prediction parameter, Represents the smoothing non-linear activation function, Represents the subscript of the Dirichlet diffusion module, Represents a constant.

[0010] Optionally, in the multi-view tracking method based on diffusion clustering, where obtaining the binary indication of the true label set, constructing a UPCE loss function according to the Dirichlet prediction parameter and the binary indication, and optimizing the constructed dynamic evidence fusion module through the UPCE loss function specifically includes: Set the true label set and obtain the binary indication of the true label set: ; When Then ; Among them, Represents the binary indication, Represents the index of the true label set, Represents the th binary indication, Represents the dimension of the multi-dimensional probability simplex, Represents the true label set; Input the Dirichlet prediction parameter and the binary indication into the constructed UPCE loss module to output the UPCE loss function: ; Among them, Represents the UPCE loss function, Denotes the UPCE loss module, Denotes the Dirichlet prediction parameter, Denotes the subscript of the Dirichlet diffusion module, Denotes the Dirichlet diffusion module, Denotes the random variable of the Dirichlet prediction parameter distribution, Denotes the -dimensional random variable, Denotes the distribution of the Dirichlet prediction parameter; Simplify the UPCE loss function through a Beta random variable, and optimize the constructed dynamic evidence fusion module through the simplified UPCE loss function: ; ; Among them, Denotes the Beta distribution, Denotes the Beta random variable, Denotes the total evidence of the correct label, Denotes the original view evidence, Denotes the -dimensional total evidence, Denotes in the form of the negative logarithm matrix of Denotes in the form of the negative logarithm matrix of

[0011] Optionally, for the multi-view tracking method based on diffusion clustering, where mapping the initial view evidence to multiple belief masses and inputting all the belief masses into the dynamic evidence fusion module to output a belief vector specifically includes: For each view, normalize the corresponding initial view evidence to obtain the belief mass and the corresponding ignorance degree corresponding to each view: ; ; ; Among them, Denotes the belief mass of the -th class in the -th view, Denotes the ignorance degree of the -th class in the -th view, Denotes the proposition that "the first class is true", Denotes "the The proposition of "class is true" Indicating "the class is true" proposition Indicating the index of Indicating the dimension of the multi - dimensional probability simplex Indicating the th view - corresponding initial view evidence of the class Indicating the th view - corresponding original view evidence Indicating the identification framework Input all the belief masses into the optimized dynamic evidence fusion module, which fuses two of the belief masses multiple times until all the belief masses are fused into a belief vector and outputs: ; ; Wherein, Indicates the first belief mass Indicates the second belief mass Indicates the combination using Dempster's combination rule of and , Indicates any subset in that is assigned a non - zero belief mass any subset in that is assigned a non - zero belief mass and the intersection of Indicates the conflict coefficient Indicates the empty set ; Wherein, Indicates the belief vector at the th discrete time Indicates the belief mass for the first class Indicates the belief mass for the th class

[0012] Optionally, in the multi - view tracking method based on diffusion clustering, wherein, inputting the belief vector into an adaptive Kalman filter to output the final clustering assignment results of all the views, specifically including: Inputting the belief vector into an adaptive Kalman filter for preliminary prediction to obtain a predicted clustering result: ; ; Among them, represents the predicted clustering result at the -th discrete time, represents the state transition matrix, represents the clustering assignment result at the -th discrete time, represents the covariance at the -th discrete time, represents the predicted covariance predicted according to , represents the transpose, represents the process covariance; The adaptive Kalman filter performs gain processing on the predicted covariance to obtain a conflict degree, and updates the predicted covariance and the predicted clustering result according to the conflict degree, and outputs a final covariance and a final clustering assignment result: ; ; ; Among them, represents the conflict degree at the -th discrete time, represents the observation matrix, represents the observation noise covariance, represents the final covariance, represents the identity matrix, represents the final clustering assignment result, represents the -th discrete time belief vector.

[0013] In addition, to achieve the above object, the present invention further provides a multi-view tracking system based on diffusion clustering, wherein the multi-view tracking system based on diffusion clustering includes: A feature encoding module, configured to obtain original data corresponding to multiple views, map all the original data to a latent space by using multiple feature encoders to obtain a feature representation, and project the feature representation into multiple initial view evidences; A diffusion module, configured to construct a Dirichlet diffusion module according to all the initial view evidences, perform iterative forward diffusion through the Dirichlet diffusion module to obtain an accumulated diffusion rate, perform backward diffusion according to the accumulated diffusion rate, and output the backward-transformed Dirichlet prediction parameter; An optimization module, configured to obtain a binary indication of a set of true labels, construct a UPCE loss function according to the Dirichlet prediction parameter and the binary indication, and optimize the constructed dynamic evidence fusion module through the UPCE loss function; A fusion module, configured to map the initial view evidence into multiple belief masses, input all the belief masses into the dynamic evidence fusion module, and output a belief vector; A clustering posterior generation module, configured to input the belief vector into an adaptive Kalman filter and output the final clustering assignment results of all the views.

[0014] In addition, to achieve the above object, the present invention further provides a terminal, where the terminal includes: a memory, a processor, and a multi-view tracking program based on diffusion clustering stored on the memory and executable on the processor. When the multi-view tracking program based on diffusion clustering is executed by the processor, the steps of the multi-view tracking method based on diffusion clustering as described above are implemented.

[0015] In addition, to achieve the above object, the present invention further provides a computer-readable storage medium, where the computer-readable storage medium stores a multi-view tracking program based on diffusion clustering. When the multi-view tracking program based on diffusion clustering is executed by a processor, the steps of the multi-view tracking method based on diffusion clustering as described above are implemented.

[0016] In the present invention, the original data corresponding to multiple views is obtained, all the original data is mapped into a latent space by using multiple feature encoders to obtain feature representations, and the feature representations are projected into multiple initial view evidences; according to all the initial view evidences, a Dirichlet diffusion module is constructed, and iterative forward diffusion is performed through the Dirichlet diffusion module to obtain a cumulative diffusion rate. Reverse diffusion is performed according to the cumulative diffusion rate, and the Dirichlet prediction parameters after reverse transformation are output; a binary indicator of the true label set is obtained, a UPCE loss function (Uncertainty-aware Cross Entropy Loss) is constructed according to the Dirichlet prediction parameters and the binary indicator, and the constructed dynamic evidence fusion module is optimized through the UPCE loss function; the initial view evidences are mapped into multiple belief masses, all the belief masses are input into the dynamic evidence fusion module, and a belief vector is output; the belief vector is input into an adaptive Kalman filter, and the final clustering assignment results of all the views are output. The present invention provides a robust, imputation-free, and multi-view clustering and tracking solution that can directly model and process uncertainty, constructs an end-to-end deep clustering framework, solves the problems of data missing, noise propagation, and insufficient uncertainty quantification in the prior art, avoids the propagation of reconstruction noise, enhances the clustering performance under missing and noisy observations, and can effectively integrate multi-source information. Description of the Drawings

[0017] Figure 1 is a flowchart of a preferred embodiment of the multi-view tracking method based on diffusion clustering of the present invention; Figure 2 is a framework diagram of a preferred embodiment of the multi-view tracking method based on diffusion clustering of the present invention; Figure 3 is a convergence curve diagram of a preferred embodiment of the multi-view tracking method based on diffusion clustering of the present invention; Figure 4 is a bar chart of the clustering performance of a preferred embodiment of the multi-view tracking method based on diffusion clustering of the present invention; Figure 5 is a first performance schematic diagram of hyperparameters of a preferred embodiment of the multi-view tracking method based on diffusion clustering of the present invention; Figure 6 is a second performance schematic diagram of hyperparameters of a preferred embodiment of the multi-view tracking method based on diffusion clustering of the present invention; Figure 7 is a third performance schematic diagram of hyperparameters of a preferred embodiment of the multi-view tracking method based on diffusion clustering of the present invention; Figure 8 is a fourth performance schematic diagram of hyperparameters of a preferred embodiment of the multi-view tracking method based on diffusion clustering of the present invention; Figure 9 is a fifth performance schematic diagram of hyperparameters of a preferred embodiment of the multi-view tracking method based on diffusion clustering of the present invention; Figure 10 is a first schematic diagram of the performance trend of a preferred embodiment of the multi-view tracking method based on diffusion clustering of the present invention; Figure 11 is a second schematic diagram of the performance trend of a preferred embodiment of the multi-view tracking method based on diffusion clustering of the present invention; Figure 12 is a third schematic diagram of the performance trend of a preferred embodiment of the multi-view tracking method based on diffusion clustering of the present invention; Figure 13 is a fourth schematic diagram of the performance trend of a preferred embodiment of the multi-view tracking method based on diffusion clustering of the present invention; Figure 14 is a fifth schematic diagram of the performance trend of a preferred embodiment of the multi-view tracking method based on diffusion clustering of the present invention; Figure 15 is a structural diagram of a preferred embodiment of the multi-view tracking system based on diffusion clustering of the present invention; Figure 16 is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed implementation manners

[0018] To make the objectives, technical solutions and advantages of the present invention more clear and definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only for explaining the present invention and are not used to limit the present invention.

[0019] The diffusion clustering method based on the Dirichlet diffusion module according to a preferred embodiment of the present invention, as Figure 1 shown, the multi-view tracking method based on diffusion clustering includes the following steps: Step S10: Obtain the original data corresponding to multiple views, map all the original data to the latent space to obtain multiple feature representations, and project the multiple feature representations into multiple initial view evidences.

[0020] Among them, for traditional shallow IMVC methods, joint matrix factorization (Joint Matrix Factorisation), incomplete multiple-kernel learning (Incomplete Multiple-Kernel Learning), or graph-regularized alignment methods (Graph-regularised Alignment Methods) are usually adopted; although these shallow models provide elegant algebraic formulas, their linear representation ability is limited and it is difficult to capture the non-linear associations commonly existing in modern sensing data. In addition, when dealing with mixed data types such as images, texts, and videos, their optimization processes often become cumbersome and complex. Therefore, the present application proposes a Dirichlet-Sharma-Mittal Diffusion Track (DSMDT, deep clustering without imputation) framework, as Figure 2 shown (where represents the dimension at the th discrete time), and its core lies in realizing multi-view data processing, uncertainty modeling, diffusion process, and evidence fusion in a unified end-to-end paradigm.

[0021] Specifically, obtain the original data corresponding to multiple views, and input all the original data into multiple feature encoders respectively to obtain the hidden vector corresponding to each original data: ; where represents the index of the view, represents the hidden vector of the th view, represents the feature encoder of the th view, represents the original data of the th view, represents the set of real numbers, represents the The dimension of the th view, a set of real numbers of dimension ; Define an affine transformation for each of the said views, and perform batch normalization on all the said latent vectors according to each of the said affine transformations to obtain a feature representation corresponding to each of the said original data in the latent space: ; where represents the feature representation of the th view, represents the affine transformation of the th view, represents batch normalization, represents the shared projection of the th view, represents a set of real numbers of dimension, represents the dimension of ; Map all the said feature representations to initial view evidences corresponding to each of the said views and output: ; where represents the initial view evidence of the th view, represents a smooth non - linear activation function, represents the shared projection, represents the translation vector,

[0022] where, for all the obtained views, their respective feature encoders (e.g., a convolutional neural network for image data, a Transformer model for text data, etc.) are used to map their original data to a common latent space to generate view - specific feature representations (i.e., Figure 2 in , and ); Subsequently, project these feature representations into initial view evidences; This module is responsible for converting the original, heterogeneous multi - view data (multi - modal data) into unified numerical features that can be processed in the subsequent diffusion process, providing high - quality input for subsequent clustering analysis, where represents the process of reverse diffusion of the input data for the th step of iteration, represents the posterior distribution obtained after reverse diffusion of the input data for the th step of iteration. represents the input data for the th iteration step, represents the input data for the th iteration step, represents the posterior distribution for the th iteration step.

[0023] Step S20: Based on all the initial view evidences, construct a Dirichlet diffusion module, perform iterative forward diffusion through the Dirichlet diffusion module to obtain the cumulative diffusion rate, and perform backward diffusion based on the cumulative diffusion rate to output the Dirichlet prediction parameters after reverse transformation.

[0024] Among them, the operation process of the Dirichlet diffusion module mainly includes forward diffusion and backward diffusion on the Dirichlet distribution. At the same time, the Sharma-Mittal (SM) divergence (a generalized information-theoretic divergence measure) can be introduced as a regularizer for diffusion training. The SM divergence is a two-parameter generalization of the Kullback-Leibler divergence (relative entropy, an asymmetric measure of the difference between two probability distributions) and the Rényi divergence (originating from the Rényi entropy, used to measure uncertainty). It allows controlling the "thickness" of the distribution tails and the sensitivity to outliers.

[0025] Specifically, perform normalization processing on each of the initial view evidences to obtain the clustering assignment vector corresponding to each view, model each of the clustering assignment vectors as a multi-dimensional probability simplex, and use the multi-dimensional probability simplex as a Dirichlet random variable to construct a Dirichlet diffusion module: ; ; ; where, represents the initial view evidence of the th view, represents the concentration parameter of the first dimension, represents the th dimension's concentration parameter, represents the transpose, represents the th clustering probability, represents the th dimension's total evidence, represents the th-dimensional probability simplex, represents the normalized clustering probability, represents the th-dimensional real number set, represents the clustering probability of the d - dimension; construct a diffusion time schedule, perform iterative forward diffusion on the Dirichlet diffusion module according to the diffusion time schedule to obtain an accumulated diffusion rate, construct a single - step posterior distribution according to the accumulated diffusion rate, and predict Dirichlet prediction parameters through a reverse denoising network according to the single - step posterior distribution.

[0026] Based on existing deep - learning - based IMVC methods, autoencoders or generative adversarial networks (GANs) are used to impute missing views, and then clustering is performed after reconstructing the complete data. For example, some diffusion - model - based methods also adopt a similar strategy, first imputing missing views through a diffusion process and then performing subsequent processing. Although deep networks bring the advantage of representation ability, the overall performance of these methods highly depends on the fidelity of the reconstructed views. Inaccurate imputation will spread noise to the clustering stage, resulting in a decline in clustering results. At the same time, the uncertainty of the imputed content is rarely quantified, and the huge difference in the proportion of missing views between different modalities may make the consensus representation biased towards the views with the dominant data volume. In addition, such methods usually operate in Euclidean space, ignoring the geometric property that clustering assignments naturally exist on the probability simplex.

[0027] Therefore, this application maps the initial view evidence to a clustering assignment vector and models the clustering assignment vector as a Dirichlet random variable on a multi - dimensional probability simplex; introduces a diffusion model that directly operates on the Dirichlet simplex in the IMVC field, completely avoiding the explicit reconstruction of missing views, which fundamentally solves the problems of noise propagation and uncertainty quantification in existing imputation methods; and by modeling the clustering assignment vector as a Dirichlet random variable, directly captures the inherent uncertainty caused by missing data without synthesizing missing views, making the clustering assignment itself contain confidence information.

[0028] Furthermore, construct a diffusion time schedule and perform multiple forward diffusions according to the diffusion time schedule: ; where the diffusion time schedule is , represents the total discrete time, represents the index of the discrete time, represents the forward diffusion kernel of the step iteration, represents the The single-step size of the t-th iteration Denote the Final input data of the t-th iteration Denote the Input data of the t-th iteration Denote the Dirichlet diffusion module Denote the t-th forward process Denote the pure noise center; According to the single-step size in each forward diffusion, construct the cumulative diffusion rate, and construct the posterior scaling factor, and construct the single-step posterior distribution according to the posterior scaling factor: ; ; ; Wherein Denote the Cumulative diffusion rate of the t-th step Denote the Cumulative diffusion rate of the (t-1)-th step Denote Index of Denote the k-th single-step size Denote the posterior scaling factor Denote the single-step posterior distribution Denote the initial input data Denote the original view evidence; Determine the final input data according to the single-step posterior distribution, input the final input data into the time conditional encoder in the Dirichlet diffusion module, output the Dirichlet prediction parameters, and input the prediction parameters into the reverse denoising network to output the reverse denoised Dirichlet distribution: ; ; Wherein Denote the Final input data of the t-th iteration Denote the Dirichlet distribution of the t-th iteration Denote the Dirichlet prediction parameters Denote the time conditional encoder Denote the smooth non-linear activation function Denote the subscript of the Dirichlet diffusion module Denote the constant

[0029] The same IMVC method based on deep learning simply concatenates or linearly weights view-specific features and then applies an Euclidean regularizer; this static transformation is difficult to model the cross-view non-linear relationship. When the number of observable instances varies significantly across views, low-quality common embeddings are often produced. More importantly, operating in the Euclidean space contradicts the fact that soft clustering assignments naturally exist on the probability simplex. Also, existing static fusion rules cannot adapt to the evolution of view availability over time, limiting their robustness in dynamic environments.

[0030] Therefore, the present invention directly performs forward and reverse diffusion on the probability simplex, ensuring the effectiveness and consistency of clustering assignments and avoiding unnecessary mapping and distortion. Intuitively, each step of diffusion preserves a proportion of prior evidence while injecting a proportion of uncertainty that is uniformly distributed across categories. Through iterative forward diffusion, a closed-form marginal distribution can be obtained.

[0031] Furthermore, a time-conditioned Transformer encoder is learned to parameterize the score function of reverse diffusion, and a denoising network is used to predict the Dirichlet parameters of reverse transformation. For the training process, the posterior distribution of a single step is calculated, and this closed-form posterior distribution eliminates the Monte Carlo variance during training.

[0032] Among them, can ensure that the parameter is positive, while the constant is less than , which can ensure that the Dirichlet prediction parameter is greater than 0.

[0033] Furthermore, the SM divergence is used as the regularizer for diffusion training (where and represent the two parameters of the SM divergence, and represent two probability distributions), generalizing the Kullback-Leibler and Rényi divergences. Based on this, the objective function of diffusion training can be defined as: ; Among them, represents the objective function, represents the subscript of the Dirichlet diffusion module, Represents the expectation of the final input data; since both parameters of the SM divergence are Dirichlet distributions, a closed-form gradient can be obtained, enabling stable optimization without the need for score normalization techniques; using the SM divergence as a flexible regularizer extends the scope of the traditional KL objective function (Kullback-Leibler Divergence Objective Function), providing a controllable balance between pattern finding and distribution matching behaviors, thus significantly improving the robustness of the model under different missing view ratios.

[0034] Step S30: Obtain the binary indication of the true label set, construct the UPCE loss function based on the Dirichlet prediction parameter and the binary indication, and optimize the constructed dynamic evidence fusion module through the UPCE loss function.

[0035] Specifically, set the true label set and obtain the binary indication of the true label set: ; When At that time ; Among them, Represents the binary indication, Represents the index of the true label set, Represents the th binary indication, Represents the dimension of the multi-dimensional probability simplex, Represents the true label set; input the Dirichlet prediction parameter and the binary indication into the constructed UPCE loss module, and output the UPCE loss function: ; Among them, Represents the UPCE loss function, Represents the UPCE loss module, Represents the Dirichlet prediction parameter, Represents the subscript of the Dirichlet diffusion module, Represents the Dirichlet diffusion module, Represents the random variable of the Dirichlet prediction parameter distribution, Represents the dimensional random variable, Represents the distribution of the Dirichlet prediction parameter; simplify the UPCE loss function through the Beta random variable, and optimize the constructed dynamic evidence fusion module through the simplified UPCE loss function: ; ; Among them, represents the Beta distribution, represents a Beta random variable, represents the total evidence of the correct label, represents the original view evidence, represents the total evidence of the represents in the form of the negative logarithm matrix of represents in the form of the negative logarithm matrix of

[0036] Among them, the present invention proposes a UPCE loss function for dealing with incomplete (multi-label or single-label) supervision, which can effectively utilize weak supervision or multi-label supervision information, and by penalizing predictions with overly high confidence, ensures that the model respects its own uncertainty during supervised learning; the loss function directly considers the uncertainty encoded by the Dirichlet distribution and naturally down-weights samples with high uncertainty (low evidence), making the training process more stable and robust; the analytical gradient of the loss function avoids Monte Carlo noise, making the optimization process more efficient and stable.

[0037] Step S40: Map multiple pieces of the initial view evidence into multiple belief masses, and input all the belief masses into the dynamic evidence fusion module to output a belief vector.

[0038] Specifically, for each view, normalize the corresponding initial view evidence to obtain the belief mass and the corresponding ignorance degree for each view: ; ; ; Among them, represents the belief mass of the th view for the th category, represents the ignorance degree of the th view for the th category, represents the proposition that "the first category is true", represents the proposition that "the th category is true", represents the proposition that "the th category is true", represents index of represents the dimension of the multi-dimensional probability simplex, Indicates the initial view evidence corresponding to the th view, Indicates the original view evidence corresponding to the th view, Indicates the identification framework; Input all the belief masses into the optimized dynamic evidence fusion module. The dynamic evidence fusion module fuses two of the belief masses multiple times until all the belief masses are fused into a belief vector and outputs: ; ; Wherein, represents the first belief mass, represents the second belief mass, represents the combination using Dempster's combination rule and , represents any subset in which non-zero belief masses are assigned, represents any subset in which non-zero belief masses are assigned, represents and the intersection of, represents the conflict coefficient, represents the empty set; ; Wherein, represents the belief vector at the th discrete time, represents the belief mass for the first class, represents the belief mass for the th class.

[0039] Among them, when the fusion module processes the Dirichlet parameters from M views, it proceeds in three stages: Dirichlet to BPA (Basic Probability Assignment) mapping, Dempster (Dempster's combinational rule, used to fuse the degrees of belief of multiple independent evidence sources to support uncertainty reasoning and decision-making) synthesis, and subsequent adaptive Kalman filter.

[0040] Among them, for the belief mass of each class, the corresponding ignorance degree will be quantified, and a high evidence value means a low ignorance degree. This mapping process is unified for all views without omission.

[0041] Furthermore, when fusing two independent BPAs, the more missing values there are, the smaller the conflict coefficient. Based on this, during the Dempster synthesis process, evidence from different views is weighted and combined, effectively handling the problems of view missing and information imbalance, and avoiding the dominance of a single view in the consensus representation.

[0042] Step S50: Input the belief vector into an adaptive Kalman filter to output the final clustering assignment results for all the views.

[0043] Among them, the Bayesian deep network estimates uncertainty through Monte Carlo Dropout (a regularization technique for neural networks) or predictive entropy. The Dempster - Shafer theory (DST) provides a formal calculus for combining belief masses, and the Kalman filter is a standard method for processing sequential data. However, the estimation of the Bayesian deep network cannot directly handle conflicting multi - sensor evidence, and traditional Kalman filter variants usually assume Gaussian noise and linear dynamics, which will limit their performance when dealing with complex, non - linear multimodal data. Although there are some improved Kalman variants, they still lack a mechanism to perfectly combine with evidence in non - Euclidean spaces (such as simplices) and adaptively adjust the noise covariance according to evidence conflict.

[0044] Therefore, the present invention combines the Dempster - Shafer evidence theory for inter - view evidence aggregation and applies an adaptive Kalman filter for temporal smoothing (where the noise covariance is adaptively adjusted according to evidence conflict), thereby effectively alleviating the view imbalance problem and ensuring the temporal consistency of clustering assignment.

[0045] Specifically, input the belief vector into an adaptive Kalman filter for preliminary prediction to obtain a predicted clustering result: ; ; where represents the predicted clustering result at the -th discrete time, represents the state transition matrix, represents the clustering assignment result at the -th discrete time, represents the covariance at the -th discrete time, represents the predicted covariance predicted according to , represents the transpose, represents the process covariance; The adaptive Kalman filter processes the predicted covariance to obtain a conflict degree, and updates the predicted covariance and the predicted clustering result according to the conflict degree, and outputs a final covariance and a final clustering assignment result: ; ; ; where, represents the conflict degree at the -th discrete time, represents the observation matrix, represents the observation noise covariance, represents the final covariance, represents the identity matrix, represents the final clustering assignment result, represents the -th discrete time belief vector.

[0046] Among them, the synthesized belief vector is used as a measurement value and input into the adaptive Kalman filter, and the prediction and update processes of clustering classification are carried out therein. Its adaptive process is mainly reflected in the conflict degree: the higher the conflict degree, the larger the observation noise covariance will be increased, which makes the historical data more credible. On the contrary, more reliance is placed on the current predicted value, making the fusion process smooth and robust; by adaptively adjusting the observation noise covariance of the Kalman filter according to the conflict mass generated by the Dempster synthesis rule, the reliability of different view evidences can be intelligently evaluated, so as to reduce the dependence on the measurement value when the evidence conflict is high, significantly improving the robustness of the fusion, and the finally generated clustering posterior is time-consistent and self-consistent, and can remain stable even in the case of up to 70% view loss.

[0047] Furthermore, the total training objective function of the present invention integrates all the above modules together: ; where, , and are hyperparameters, which are respectively used to balance diffusion matching, evidence-aware supervision and fusion consistency. This objective function is trained end-to-end through an optimizer such as Adam, realizing end-to-end optimization of the entire process from the original data to the final clustering assignment, enabling all parts of the model to work together and jointly improving the performance.

[0048] During inference, given an independent and identically distributed noise sample, the model iteratively performs the reverse diffusion process. The training objective takes into account the accuracy of the diffusion process, the utilization of weak supervision information, and the consistency of multi-view fusion, ensuring that the final clustering assignment is uncertainty-calibrated.

[0049] Subsequently, a single pass is made through the dynamic evidence fusion module to obtain a temporally smoothed clustering posterior. This integrated optimization strategy enables the model to exhibit excellent robustness under various missing and noisy conditions.

[0050] Furthermore, in the embodiments disclosed in the present invention, experiments were conducted on five widely used multi-view datasets (Caltech-7 (a reduced version of the Caltech 101 dataset), Citeseer (derived from the CiteSeer publishing platform), Cora (a benchmark dataset for graph neural network research), Handwritten (Multimodal Handwritten Digit Dataset), and Reuters (Reuters News Corpus Dataset)). Table 1 below shows the clustering performance of the present invention on the above five datasets under different missing view ratios (0.1, 0.3, 0.5, 0.7), and a comparison with the prior art was made. The evaluation metrics include accuracy (ACC), normalized mutual information (NMI), and purity (Purity): Table 1: Evaluation comparison table

[0051] Among them, the first method is ; the second method is ; the third method is ; the fourth method is ; the fifth method is ‌; the sixth method ; the six methods are different methods in the prior art.

[0052] The method of the present invention shows continuous and significant superiority in the three key evaluation metrics of ACC, NMI, and Purity under various missing view ratios of the above five datasets (from 0.1 to 0.7, i.e., up to 70% data missing). Especially on the Caltech-7, Citeseer, Cora, and Handwritten datasets, the performance of the present invention far exceeds the compared prior art advanced methods (such as C IMUFS, LSIMVC, BGIMVSC, PGP, AGDIMC, DIMVC). This fully verifies the excellent robustness and effectiveness of the present invention in dealing with incomplete and noisy multi-view data. Even when the data missing rate is as high as 70%, the present invention can still maintain high clustering performance, which benefits from its imputation-free Dirichlet diffusion modeling, the robustness of Sharma-Mittal divergence, and the adaptive evidence fusion mechanism.

[0053] The visualization effect achieved by the present invention is as Figure 3 shown ( Figure 3 in (a) of Figure 3 and (b) of

[0054] respectively show the objective functions in different datasets), showing the objective function convergence curves under different missing view ratios on the Handwritten (a general term for data resources related to handwritten character recognition) and Reuters datasets (a dataset covering 46 news topics and used for multi-label classification tasks). The results show that the present invention achieves fast and stable convergence by using a closed-form gradient, which verifies the efficiency and stability of its optimization process and avoids the common gradient instability problem in traditional methods. Figure 4 shown ( Figure 4 in (a) of Figure 4 in (b) of Figure 4 in (c) of Figure 4 in (d) of Figure 4 and (e) of

[0055] respectively represent the clustering performance under different missing view ratios), presenting the ACC, NMI, and Purity performances of the present invention on five datasets (Caltech-7, Citeseer, Cora, Handwritten, and Reuters) under different missing view ratios, further corroborating the quantitative results in Table 1 and clearly demonstrating the excellent performance and stability of the present invention under various challenging conditions. Figures 5 to 9 shown, demonstrating the accuracy sensitivity analysis of the present invention for the core hyperparameters and ; among them, is used for evidence-aware supervision, and is used to balance diffusion matching. The results show that the present invention can maintain relatively stable high performance within a wide range of hyperparameters, which reflects its low sensitivity to hyperparameter selection and good generalization ability.

[0056] Furthermore, as Figures 10 to 14As shown, the changing trend of ACC with the increasing proportion of missing views of the method in this paper and six existing comparison methods (including CIMUFS, LSIMVC, BGIMVSC, PGP, AGDIMC, and DIMVC) on five datasets (Caltech7 ( Figure 10 ), Citeseer ( Figure 11 ), Cora ( Figure 12 ), Handwritten ( Figure 13 ), and Reuters ( Figure 14 )) is presented. It can be observed that as the missing proportion increases, the performance of each method generally decreases, while the method in this paper maintains a significantly leading clustering accuracy on all datasets, demonstrating good robustness and adaptability, further verifying the effectiveness and stability of the present invention in dealing with incomplete multi-view scenarios.

[0057] The present invention provides a robust, imputation-free multi-view clustering and tracking solution that can directly model and handle uncertainties, constructs an end-to-end deep clustering framework, solves the problems of data missing, noise propagation, and insufficient uncertainty quantification in the prior art, avoids the propagation of reconstruction noise, enhances the clustering performance under missing and noisy observations, and can effectively integrate multi-source information.

[0058] Furthermore, as shown in Figure 15 , based on the above-mentioned multi-view tracking method based on diffusion clustering, the present invention also correspondingly provides a multi-view tracking system based on diffusion clustering, wherein the multi-view tracking system based on diffusion clustering includes: A feature encoding module 51, configured to obtain the original data corresponding to multiple views, map all the original data to a latent space by using multiple feature encoders to obtain feature representations, and project the feature representations into multiple initial view evidences; A diffusion module 52, configured to construct a Dirichlet diffusion module according to all the initial view evidences, perform iterative forward diffusion through the Dirichlet diffusion module to obtain a cumulative diffusion rate, perform backward diffusion according to the cumulative diffusion rate, and output the Dirichlet prediction parameters of the backward transformation; An optimization module 53, configured to obtain a binary indication of the set of true labels, construct a UPCE loss function according to the Dirichlet prediction parameters and the binary indication, and optimize the constructed dynamic evidence fusion module through the UPCE loss function; A fusion module 54, configured to map the initial view evidences into multiple belief masses, input all the belief masses into the dynamic evidence fusion module, and output a belief vector; The clustering posterior generation module 55 is configured to input the belief vector into an adaptive Kalman filter and output the final clustering assignment results of all the views.

[0059] Further, as Figure 16 shown, based on the above multi-view tracking method and system based on diffusion clustering, the present invention also correspondingly provides a terminal, which includes a processor 10, a memory 20, and a display 30. Figure 16 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.

[0060] The memory 20 may be an internal storage unit of the terminal in some embodiments, such as a hard disk or a memory of the terminal. The memory 20 may also be an external storage device of the terminal in other embodiments, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the terminal. Further, the memory 20 may also include both the internal storage unit and the external storage device of the terminal. The memory 20 is used to store application software installed on the terminal and various types of data, such as program codes for installing the terminal. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, a multi-view tracking program 40 based on diffusion clustering is stored on the memory 20, and the multi-view tracking program 40 based on diffusion clustering can be executed by the processor 10, thereby implementing the multi-view tracking method based on diffusion clustering in the present application.

[0061] The processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chips in some embodiments, and is used to run the program codes stored in the memory 20 or process data, such as executing the multi-view tracking method based on diffusion clustering.

[0062] The display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. in some embodiments. The display 30 is used to display information on the terminal and for displaying a visual user interface. The components of the terminal communicate with each other through a system bus.

[0063] In one embodiment, when the processor 10 executes the multi-view tracking program 40 based on diffusion clustering in the memory 20, the steps of the above multi-view tracking method based on diffusion clustering are implemented.

[0064] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a multi-view tracking program based on diffusion clustering, and when the multi-view tracking program based on diffusion clustering is executed by a processor, the steps of the multi-view tracking method based on diffusion clustering as described above are implemented.

[0065] In summary, the present invention provides a multi-view tracking method and related devices based on diffusion clustering. The method includes: obtaining original data corresponding to multiple views, using multiple feature encoders to map all the original data into a latent space to obtain feature representations, and projecting the feature representations into multiple initial view evidences; constructing a Dirichlet diffusion module according to all the initial view evidences, and performing iterative forward diffusion through the Dirichlet diffusion module to obtain a cumulative diffusion rate, performing backward diffusion according to the cumulative diffusion rate, and outputting Dirichlet prediction parameters of reverse transformation; obtaining a binary indication of a set of true labels, constructing a UPCE loss function according to the Dirichlet prediction parameters and the binary indication, and optimizing the constructed dynamic evidence fusion module through the UPCE loss function; mapping the initial view evidences into multiple belief masses, inputting all the belief masses into the dynamic evidence fusion module, and outputting a belief vector; inputting the belief vector into an adaptive Kalman filter, and outputting final clustering assignment results of all the views. The present invention provides a robust, imputation-free, multi-view clustering and tracking solution that can directly model and process uncertainties, constructs an end-to-end deep clustering framework, solves the problems of data missing, noise propagation, and insufficient uncertainty quantification in the prior art, avoids the propagation of reconstruction noise, enhances the clustering performance under missing and noisy observations, and can effectively integrate multi-source information.

[0066] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or terminal including the element.

[0067] Of course, those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium readable by a computer. When the program is executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.

[0068] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A multi-view tracking method based on diffusion clustering, characterized in that The multi-view tracking method based on diffusion clustering includes: Obtain the original data corresponding to multiple views, map all the original data to the latent space to obtain multiple feature representations, and project the multiple feature representations into multiple initial view evidences; According to all the initial view evidences, construct a Dirichlet diffusion module, perform iterative forward diffusion through the Dirichlet diffusion module to obtain a cumulative diffusion rate, perform backward diffusion according to the cumulative diffusion rate, and output the Dirichlet prediction parameters after inverse transformation; Obtain the binary indication of the true label set, construct a UPCE loss function according to the Dirichlet prediction parameters and the binary indication, and optimize the constructed dynamic evidence fusion module through the UPCE loss function; Map the multiple initial view evidences to multiple belief masses, and input all the belief masses into the dynamic evidence fusion module to output a belief vector; Input the belief vector into an adaptive Kalman filter to output the final clustering assignment results of all the views.

2. The multi-view tracking method based on diffusion clustering according to claim 1, wherein, The obtaining the original data corresponding to multiple views, mapping all the original data to the latent space to obtain multiple feature representations, and projecting the multiple feature representations into multiple initial view evidences specifically includes: Obtain the original data corresponding to multiple views, and respectively input all the original data into multiple feature encoders to obtain the hidden vector corresponding to each original data: ; Among them, represents the index of the view, represents the th hidden vector of the view, represents the th feature encoder of the view, represents the th original data of the view, represents the set of real numbers, represents the th dimension of the view, represents the th -dimensional set of real numbers; Define the affine transformation of each view, and perform batch normalization on all the hidden vectors according to each affine transformation to obtain the feature representation corresponding to each original data in the latent space: ; Among them, represents the feature representation of the th view, represents the affine transformation of the th view, represents batch normalization processing, represents the shared projection of the th view, represents the translation vector of the th view, represents a set of real numbers of dimension represents the dimension of; Map all the feature representations to the initial view evidences corresponding to each view and output: ; Among them, represents the initial view evidence of the th view, represents a smooth non-linear activation function, represents a shared projection, represents a translation vector, represents a constant.

3. The multi-view tracking method based on diffusion clustering according to claim 1, wherein The constructing a Dirichlet diffusion module according to all the initial view evidences, performing iterative forward diffusion through the Dirichlet diffusion module to obtain a cumulative diffusion rate, performing backward diffusion according to the cumulative diffusion rate, and outputting the Dirichlet prediction parameters after inverse transformation specifically includes: Perform normalization processing on each initial view evidence to obtain the clustering assignment vector corresponding to each view, model the clustering assignment vector as a multi-dimensional probability simplex, and use the multi-dimensional probability simplex as a Dirichlet random variable to construct a Dirichlet diffusion module: ; ; ; Among them, represents the initial view evidence of the th view, represents the concentration parameter of the first dimension, represents the concentration parameter of the th dimension, represents the transpose, represents the th cluster probability, represents the total evidence of the th dimension, represents dimensional probability simplex, represents the normalized cluster probability, represents dimensional real number set, represents the th dimensional cluster probability; Construct a diffusion schedule, perform iterative forward diffusion on the Dirichlet diffusion module according to the diffusion schedule to obtain a cumulative diffusion rate, construct a single-step posterior distribution according to the cumulative diffusion rate, and predict the Dirichlet prediction parameters through a reverse denoising network according to the single-step posterior distribution.

4. The multi-view tracking method based on diffusion clustering according to claim 3, wherein The constructing a diffusion schedule, performing iterative forward diffusion on the Dirichlet diffusion module according to the diffusion schedule to obtain a cumulative diffusion rate, constructing a single-step posterior distribution according to the cumulative diffusion rate, and predicting the Dirichlet prediction parameters through a reverse denoising network according to the single-step posterior distribution specifically includes: Construct a diffusion schedule and perform multiple forward diffusions according to the diffusion schedule: ; Among them, the diffusion time table is , represents the total discrete time, represents the index of the discrete time, represents the forward diffusion kernel of the -th step iteration, represents the index of the number of iteration steps, represents the single-step size of the -th step iteration, represents the final input data of the -th step iteration, represents the input data of the -th step iteration, represents the Dirichlet diffusion module, represents the -th forward process, represents the pure noise center; Construct an accumulated diffusion rate based on the single-step size in each forward diffusion, construct a posterior scaling factor, and construct a single-step posterior distribution according to the posterior scaling factor: ; ; ; Among them, represents the cumulative diffusion rate at the th step, represents the cumulative diffusion rate at the th step, represents 's index, represents the th single-step size, represents the posterior scaling factor, represents the single-step posterior distribution, represents the initial input data, represents the original view evidence; Determine the final input data according to the single-step posterior distribution, input the final input data into the time conditional encoder in the Dirichlet diffusion module, output the Dirichlet prediction parameters, and input the prediction parameters into the reverse denoising network to output the reverse denoised Dirichlet distribution: ; ; Among them, represents the final input data of the th iteration, represents the Dirichlet distribution of the th iteration, represents the Dirichlet prediction parameter, represents the time conditional encoder, represents the smooth non-linear activation function, represents the subscript of the Dirichlet diffusion module, represents a constant.

5. The multi-view tracking method based on diffusion clustering according to claim 1, wherein Obtain the binary indication of the true label set, construct a UPCE loss function according to the Dirichlet prediction parameters and the binary indication, and optimize the constructed dynamic evidence fusion module through the UPCE loss function, specifically including: Set the true label set and obtain the binary indication of the true label set: ; When time ; Among them, represents a binary indicator, represents the index of the set of true labels, represents the th binary indicator, represents the dimension of the multi-dimensional probability simplex, represents the set of true labels; Input the Dirichlet prediction parameters and the binary indication into the constructed UPCE loss module to output the UPCE loss function: ; Among them, represents the UPCE loss function, represents the UPCE loss module, represents the Dirichlet prediction parameter, represents the subscript of the Dirichlet diffusion module, represents the Dirichlet diffusion module, represents the random variable of the Dirichlet prediction parameter distribution, represents the dimensional random variable, represents the distribution of the Dirichlet prediction parameter; Simplify the UPCE loss function through a Beta random variable, and optimize the constructed dynamic evidence fusion module through the simplified UPCE loss function: ; ; Among them, represents the Beta distribution, represents the Beta random variable, represents the total evidence of the correct label, represents the evidence of the original view, represents the -dimensional total evidence, represents in the form of the negative logarithm matrix of represents in the form of the negative logarithm matrix of 6. The multi-view tracking method based on diffusion clustering according to claim 1, characterized in that, Mapping the multiple initial view evidences to multiple belief masses and inputting all the belief masses into the dynamic evidence fusion module to output a belief vector specifically includes: For each view, normalize the corresponding initial view evidence to obtain the belief mass and the corresponding ignorance degree corresponding to each view: ; ; ; Among them, represents the belief mass for the th category within the th view, represents the ignorance degree for the th category within the th view, represents the proposition that "the first category is true", represents the proposition that "the th category is true", represents the proposition that "the th category is true", represents 's index, represents the dimension of the multi-dimensional probability simplex, represents the initial view evidence for the th view corresponding to the th category, represents the original view evidence corresponding to the th view, represents the frame of discernment; Input all the belief masses into the optimized dynamic evidence fusion module. The dynamic evidence fusion module fuses two belief masses multiple times until all the belief masses are fused into a belief vector and outputs: ; ; Among them, represents the first belief mass, represents the second belief mass, represents the combination using Dempster's combination rule of and , represents any subset in that is assigned a non-zero belief mass, represents any subset in that is assigned a non-zero belief mass, represents the intersection of represents the conflict coefficient, represents the empty set; ; Among them, represents the belief vector at the -th discrete time, represents the belief mass for the first category, represents the belief mass for the -th category.

7. The multi-view tracking method based on diffusion clustering according to claim 1, characterized in that Input the belief vector into an adaptive Kalman filter to output the final clustering assignment results of all views, specifically including: Input the belief vector into the adaptive Kalman filter for preliminary prediction to obtain a predicted clustering result: ; ; Among them, represents the predicted clustering result at the th discrete time, represents the state transition matrix, represents the clustering assignment result at the th discrete time, represents the covariance at the th discrete time, represents the predicted covariance predicted according to , represents the transpose, represents the process covariance; The adaptive Kalman filter performs gain processing on the predicted covariance to obtain a conflict degree, and updates the predicted covariance and the predicted clustering result according to the conflict degree, and outputs the final covariance and the final clustering assignment result: ; ; ; Among them, represents the conflict degree at the -th discrete time, represents the observation matrix, represents the observation noise covariance, represents the final covariance, represents the identity matrix, represents the final clustering assignment result, represents the -th discrete time belief vector.

8. A multi-view tracking system based on diffusion clustering, characterized in that, The multi-view tracking system based on diffusion clustering includes: A feature encoding module for obtaining the original data corresponding to multiple views, mapping all the original data to the latent space by using multiple feature encoders to obtain feature representations, and projecting the feature representations into multiple initial view evidences; A diffusion module for constructing a Dirichlet diffusion module according to all the initial view evidences, performing iterative forward diffusion through the Dirichlet diffusion module to obtain an accumulated diffusion rate, and performing reverse diffusion according to the accumulated diffusion rate to output the reverse-transformed Dirichlet prediction parameters; An optimization module, configured to obtain a binary indication of a set of true labels, construct a UPCE loss function according to the Dirichlet prediction parameter and the binary indication, and optimize the constructed dynamic evidence fusion module through the UPCE loss function; A fusion module, configured to map the initial view evidence into a plurality of belief masses, input all the belief masses into the dynamic evidence fusion module, and output a belief vector; A clustering posterior generation module, configured to input the belief vector into an adaptive Kalman filter and output the final clustering assignment results of all the views.

9. A terminal, characterized in that, The terminal includes: a memory, a processor, and a multi-view tracking program based on diffusion clustering stored on the memory and executable on the processor. When the multi-view tracking program based on diffusion clustering is executed by the processor, the steps of the multi-view tracking method based on diffusion clustering according to any one of claims 1-7 are implemented.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a multi-view tracking program based on diffusion clustering. When the multi-view tracking program based on diffusion clustering is executed by a processor, the steps of the multi-view tracking method based on diffusion clustering according to any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Method and system for clustering incomplete views

    CN118135279A

  • Multi-view clustering method based on uncertainty and graph neural network

    CN119206279A