Label-guided representation learning method for incomplete multi-view multi-label classification

CN122471174BActive Publication Date: 2026-09-08NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610968896.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-01
Publication Date
2026-09-08
Estimated Expiration
2046-07-01

AI Technical Summary

Technical Problem

[0007]本发明提供了一种面向不完整多视图多标签分类的标签引导表征学习方法,其目的是解决在对多模态异构数据进行分类时,因现有技术未能将标签语义系统性地贯穿于表征学习和多视图融合的全过程,从而导致在面临视图缺失与标签不完整的双重挑战时模型的实际分类准确率和鲁棒性难以保证的技术问题

Benefits of technology

与现有技术相比,本发明提供的一种面向不完整多视图多标签分类的标签引导表征学习方法,通过为每个类别构建可学习的原型分布来显式编码标签语义信息,并基于共享表征与各原型分布之间的相似性构建语义对齐损失,从而将标签共现模式和类间关系直接嵌入到潜在空间的几何结构中,解决了现有技术在表征学习阶段缺乏标签感知能力的问题;同时,本发明将每个类别的原型分布作为语义专家,与每个可用视图的单视图后验分布进行第二乘积专家融合,得到类别特定的条件后验分布并采样得到类别特定表征,再通过专用分类器进行预测,使得不同类别能够根据自身语义特性自适应地融合不同视图(如图像、文本、音频、传感器数值等异构数据模态)的特征信息,克服了现有融合机制将多视图信息聚合为单一共享表征而忽视类别间异质性需求的缺陷;此外,通过联合优化任务相关损失、语义对齐损失和分类损失,本发明系统性地将标签信息从输出层监督信号贯穿至表征学习、多视图融合以及最终分类预测的全过程,在训练阶段构建了具有语义结构的潜在空间,在预测阶段利用训练完备的模型对实际不完整输入数据进行精准推理,最终在视图和标签双重不完整的现实条件下,显著提升了模型对多模态异构数据的分类准确率和鲁棒性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122471174B_ABST
    Figure CN122471174B_ABST
Patent Text Reader

Abstract

The application relates to a label-guided representation learning method for incomplete multi-view multi-label classification, which comprises the following steps: constructing a learnable prototype distribution for each class; encoding each available view to obtain a single-view posterior; performing first product expert fusion on the single-view posteriors of all available views to sample a shared representation; constructing a semantic alignment loss based on the shared representation and the prototype distribution; for each class, performing second product expert fusion on the prototype distribution and each single-view posterior to sample a class-specific representation; inputting the class-specific representation into a corresponding classifier to obtain a prediction; calculating a classification loss according to the prediction and an observed label; jointly optimizing the semantic alignment loss and the classification loss to train a classification model; inputting input data of multiple views of a sample to be classified into the trained classification model to obtain a final classification result of the sample. The label semantics are used throughout the whole process, and the challenge of view loss and label incompleteness is effectively coped with.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically to a label-guided representation learning method for incomplete multi-view, multi-label classification. Background Technology

[0002] Multi-view multi-label classification aims to accurately classify objects belonging to multiple categories by integrating complementary information from multiple heterogeneous data sources (such as images, text, and audio). It has wide applications in multimedia content understanding, biometric recognition, and other fields. However, in practical applications, obtaining complete data for all views and labels is often very difficult due to factors such as sensor failure, privacy protection, and high annotation costs, leading to the incomplete multi-view multi-label classification problem. This problem faces the dual challenges of missing views and incomplete labels: missing views impair the model's ability to model dependencies between different views; while incomplete labels not only provide insufficient supervision signals but also obscure the inherent correlation structure between labels, severely impacting the model's classification performance.

[0003] To address the above challenges, existing technologies are mainly researched from the following three directions: The first approach focuses on robust feature extraction under incomplete view conditions. For example, early matrix factorization methods such as iMVWL (Incomplete Multi-View Weak-Label Learning) enhance robustness to missing views by learning shared subspaces. Recent deep learning methods such as DICNet (Deep Incomplete Multi-View Network) employ instance-level contrastive learning to enhance consistent representations across different views. However, these methods primarily focus on the reconstruction and alignment of input data, failing to effectively utilize label information to guide representation learning.

[0004] The second approach focuses on designing better multi-view fusion strategies. For example, LMVCAT (Label embedded Multi-view learning with Co-regularized Attention Transformer) utilizes the Transformer architecture for cross-view information aggregation, and MFD (Multi-view Factorization with Dualcomponents) decomposes representations into views. Figure 1While both view-specific components are addressed, SIP (Sparse Information Bottleneck) constructs a fusion scheme from the perspective of information bottlenecks. However, most of these methods fuse all view information into a single shared representation, implicitly assuming that all categories equally benefit from the same combination of features. This assumption often fails in practice because different categories (such as "sky" and "cars") may depend on discriminative features in different views.

[0005] The third approach explores the modeling of label semantics. For example, NAIM3L (Not-All-Iterative Multi-view Multi-label Learning) combines global high-rank and local low-rank constraints to capture label relevance, CSA (Cross-view Semantic Alignment) proposes enhanced contrastive learning for cross-view semantic alignment, and SSP (Semantic Structure Preservation) employs graph constraint learning to maintain semantic structural consistency among samples. Although these methods attempt to utilize label information, their utilization remains at a relatively shallow level. Specifically, existing technologies suffer from the following three shortcomings: at the label level, the label co-occurrence patterns and inter-class semantic relationships implied in some annotations are not explicitly modeled and utilized; at the representation level, label information is only used as the final supervision signal of the output layer, and the latent representation learning process itself lacks label awareness, resulting in learned representations that, while retaining input information, lack explicit semantic structures aligned with label relationships; at the fusion level, current fusion mechanisms ignore the heterogeneous requirements of different categories' contributions to different views, thus producing suboptimal representations for classification.

[0006] In summary, how to more fully mine and utilize the semantic information contained in the labels under the condition of incomplete views and labels, and systematically integrate it into the entire process of representation learning and multi-view fusion to improve the classification accuracy and robustness of the model, is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0007] This invention provides a label-guided representation learning method for incomplete multi-view, multi-label classification. Its purpose is to solve the technical problem that when classifying multimodal heterogeneous data, existing technologies fail to systematically integrate label semantics throughout the entire process of representation learning and multi-view fusion, resulting in difficulty in guaranteeing the actual classification accuracy and robustness of the model when facing the dual challenges of missing views and incomplete labels.

[0008] To achieve the above objectives, the first aspect of the present invention provides a label-guided representation learning method for incomplete multi-view multi-label classification, comprising the following steps: A learnable prototype distribution is constructed for each category, and the prototype distribution is used to encode the semantic information of the corresponding category; The input data of multiple views of the sample is obtained. For each available view, the single-view posterior distribution of that available view is obtained through the encoder. The single-view posterior distributions of all available views are subjected to first product expert fusion to obtain multi-view fusion posterior, and shared representations are obtained by sampling from the multi-view fusion posterior. Based on the similarity between the shared representation and each prototype distribution, a semantic alignment loss is constructed so that the shared representation attracts the prototype distribution corresponding to its positive label and repels the prototype distribution corresponding to its negative label. For each category, the prototype distribution of the category is fused with the single-view posterior distribution of each available view using a second product expert fusion to obtain the conditional posterior distribution of the category, and the category-specific representation of the category is sampled from the conditional posterior distribution. The category-specific representation of each category is input into the corresponding category's dedicated classifier to obtain the prediction result for each category; The classification loss is calculated based on the prediction results and the observed labels of the samples; The semantic alignment loss and the classification loss are jointly optimized to train a classification model; Input data is obtained from multiple views of a sample to be classified, wherein each view of the sample to be classified corresponds to different modal data, and the modal data includes at least two of the following: image modality, text modality, audio modality, and numerical sensor data modality. The sample to be classified is input into the trained classification model to obtain the final classification result of the sample to be classified.

[0009] Furthermore, the prototype distribution is a Gaussian distribution, and its mean and variance parameters are obtained by mapping the learnable category embeddings through a neural network. The learnable category embeddings are initialized as one-hot vectors, and each category embedding is transformed by a mean encoding network and a variance encoding network to generate the mean and variance parameters of the corresponding prototype distribution, wherein the variance parameter is transformed by the softplus function to ensure that it is a positive value.

[0010] Furthermore, the semantic alignment loss is constructed based on contrastive learning, in The normalized shared representation and the normalized prototype distribution mean are calculated to maximize the similarity between the shared representation and the prototype distribution mean corresponding to the positive label, and minimize the similarity with the prototype distribution mean corresponding to the negative label; the semantic alignment loss is:

[0011] in, For semantic alignment loss; For the positive label subset of the observed label set; Category indexing, traversing the positive label subset of samples Each positive label category in; The set of observation labels for the samples; For category indexing, iterate through the set of observed labels of the samples in the summation term of the denominator. All categories; For normalized shared representation; For category The normalized prototype distribution mean; For category The normalized prototype distribution mean vector; For temperature parameters; This is the vector transpose symbol.

[0012] Furthermore, the joint optimization also includes task-related learning objectives, which include feature-level constraints and label-level constraints; wherein: The feature-level constraints are used to balance the sufficiency of each view representation in reconstructing its input data with cross-view constraints. Figure 1 To the point of being compatible; The label-level constraints are used to balance the sufficiency of label predictions for each view representation with the consistency of label predictions between views. The task-related learning objective is jointly optimized with the semantic alignment loss and the classification loss.

[0013] Furthermore, the feature-level constraints are achieved by minimizing the feature-level loss for each available view:

[0014] in, Available View Feature-level loss; Available view; For reconstruction loss; The KL divergence between the single-view posterior and the multi-view fusion posterior; For expectation operators; Available View The posterior distribution of the encoder output, with parameters as follows. Given input Time-latent representation Conditional distribution; Let log-likelihood be the output of the decoder, representing the value derived from the latent representation. Reconstruct the original input The logarithm of the probability; These are weighting coefficients that control the strength of consistency constraints; The Kullback-Leibler divergence; For multi-view fusion posterior distribution; For currently available views The posterior distribution of a single view; The label-level constraint is achieved by minimizing the label-level loss:

[0015] in, Available View Label-level loss; For classification cross-entropy; To calculate the KL divergence between the predicted distributions of the fused posterior and the single-view posterior; The log-likelihood of the dedicated classifier output, with parameters as follows: , indicating from the available views The representation Predicted Labels The logarithm of the probability; For shared representations obtained from posterior sampling of multi-view fusion The predicted distribution of labels obtained by the classifier; From the available views The representation obtained by single-view posterior sampling The predicted distribution of labels obtained by the classifier; This is a weighting coefficient that controls the strength of the label consistency constraint; The task-related learning objective is a weighted sum of the feature-level loss and the label-level loss across all views:

[0016] in, Learning objectives related to the task; The set of available views for the current sample; For view indexing, iterate through all available views; It is a balancing factor.

[0017] Furthermore, the second product expert fusion employs a precision-weighted approach, where the mean and precision of the conditional posterior distribution are calculated for each category:

[0018]

[0019] in, For category The precision matrix of the conditional posterior distribution; It is a category The accuracy matrix of the prototype distribution; For category The mean vector of the prototype distribution; Available View The precision matrix output by the encoder; Available View The mean vector output by the encoder; For category The mean vector of the conditional posterior distribution; It is the inverse of the precision matrix, i.e., the covariance matrix.

[0020] Furthermore, the method also includes a multi-source prediction aggregation step: taking the prediction results obtained by a dedicated classifier from the category-specific representations corresponding to each view, and the prediction results obtained by a dedicated classifier from the shared representations, as multiple prediction sources; calculating confidence weights based on the prediction certainty of each prediction source, and weighting and aggregating the prediction results of each prediction source to obtain the final classification prediction; the confidence weights are calculated in the following way:

[0021] in, This is the final aggregated classification prediction vector; For the index of the prediction source, Predictions corresponding to shared representations Predictions of category-specific representations for each view; The number of available views, i.e., the number of views actually available for the current sample; For the first The prediction vector output by each prediction source; For the first The prediction vector output by each prediction source; Parameters for controlling weighted sharpness; The index variable in the summation of the denominator; Here is a confidence metric function used to evaluate the predictive certainty of a prediction source, defined as:

[0022] in, This is a confidence metric function used to evaluate the predicted vector. The degree of certainty; For prediction vectors; To iterate through all categories; For the prediction vector, the first... The predicted value of the class.

[0023] Furthermore, the joint optimization employs end-to-end training, and the total loss function is:

[0024] in, This is the total loss function; Learning objectives related to the task; For semantic alignment loss; For classification loss; Hyperparameters for balancing semantic alignment strength.

[0025] Furthermore, in the process of constructing the feature-level constraints and label-level constraints, the information-theoretic objective is transformed into an optimizable loss term based on variational inference; For feature sufficiency and label sufficiency, variational lower bounds are derived by introducing variational decoders and classifiers, respectively, which are transformed into reconstruction loss and classification cross-entropy loss in the feature-level loss and label-level loss, respectively. For feature consistency and label consistency, variational upper bounds are derived by using the KL divergence between the single-view posterior and the multi-view fusion posterior, and the KL divergence between the corresponding prediction distributions, respectively. These are then converted into KL divergence constraint terms in the feature-level loss and label-level loss.

[0026] Furthermore, the classification model includes a training phase and a testing phase: in the training phase, the prototype distributions are sampled to jointly optimize the classification model parameters, enabling end-to-end joint optimization of the category embedding, encoder, and dedicated classifier; in the testing phase, the deterministic mean of each prototype distribution mapped by the mean encoding network is used as a fixed semantic anchor, and the output of each view encoder is a deterministic representation, which participates in the accuracy-weighted calculation and final prediction of the second product expert fusion.

[0027] The beneficial effects of this invention are: Compared with existing technologies, this invention provides a label-guided representation learning method for incomplete multi-view multi-label classification. It explicitly encodes label semantic information by constructing a learnable prototype distribution for each category and constructs a semantic alignment loss based on the similarity between shared representations and each prototype distribution. This directly embeds label co-occurrence patterns and inter-class relationships into the geometric structure of the latent space, solving the problem of lacking label awareness in the representation learning stage of existing technologies. Simultaneously, this invention uses the prototype distribution of each category as a semantic expert and performs a second product expert fusion with the single-view posterior distribution of each available view to obtain a category-specific conditional posterior distribution. Category-specific representations are then sampled and predicted using a dedicated classifier, enabling different categories to predict according to their own... This invention adaptively fuses feature information from different views (such as images, text, audio, sensor values, and other heterogeneous data modalities) based on semantic characteristics, overcoming the shortcomings of existing fusion mechanisms that aggregate multi-view information into a single shared representation while ignoring the heterogeneity requirements between categories. Furthermore, by jointly optimizing task-related loss, semantic alignment loss, and classification loss, this invention systematically integrates label information from the output layer supervision signal to the entire process of representation learning, multi-view fusion, and final classification prediction. During the training phase, a latent space with semantic structure is constructed, and during the prediction phase, a fully trained model is used to perform accurate reasoning on actual incomplete input data. Ultimately, under the realistic condition of both incomplete views and labels, the model's classification accuracy and robustness for multimodal heterogeneous data are significantly improved. Attached Figure Description

[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below.

[0029] Figure 1 This is a framework diagram of a label-guided representation learning system disclosed in an embodiment of the present invention.

[0030] Figure 2 This is a flowchart of a label-guided representation learning method disclosed in an embodiment of the present invention. Detailed Implementation

[0031] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0032] According to embodiments of the present invention, it should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the following methods, in some cases the steps shown or described may be executed in a different order than that shown here.

[0033] The method of this invention can be applied to various practical scenarios, such as multi-lesion identification and classification in medical image analysis. In medical image diagnosis, the image data of the same patient may come from multiple views (such as CT, MRI, ultrasound, etc.), and due to equipment limitations or patient conditions, some views may be missing; at the same time, due to the high cost of annotation, doctors may only annotate some lesions, resulting in incomplete labels. This invention, through a label-guided representation learning method, can effectively fuse multi-view information, utilize the semantics of partially annotated labels, accurately identify multiple lesions, assist doctors in diagnosis, and improve the accuracy and efficiency of diagnosis. Another example is multi-attribute prediction of players in sports data analysis. By fusing multi-perspective data of players (such as video, physiological sensors, and match statistics), even when data collection is incomplete, it can accurately predict multiple attributes of players (such as physical fitness, technique, and psychological state), providing data support for coaches to develop training plans.

[0034] This invention proposes a label-guided representation learning (LGRL) method, which integrates label semantics into the entire process of representation learning and multi-view fusion. For example... Figure 1 As shown, LGRL consists of three core components: (1) a semantically aware hybrid prior, which encodes the semantics of labels into the latent spatial geometry through learnable category prototypes; and (2) a task-related learning objective, which balances reconstruction sufficiency and cross-view. Figure 1 Consistency; (3) Bayesian fusion mechanism based on category-specific conditional posterior, where the prototype acts as an expert in hypothesis-driven reasoning. The method will be described in detail below.

[0035] like Figure 2 As shown, this invention provides a label-guided representation learning method for incomplete multi-view, multi-label classification, comprising the following steps: Step S100: Construct a learnable prototype distribution for each category, wherein the prototype distribution is used to encode the semantic information of the corresponding category; In this step, to integrate the semantic information of the labels into the entire representation learning process, this invention first constructs a learnable "prototype distribution" for each category. This distribution serves as an idealized representation template for that category, encoding its core semantic features. Specifically, for a total number of categories... Tasks for each category ( Let a dedicated learnable prototype distribution be defined, which is a Gaussian distribution:

[0036] in, This indicates that the sample is assumed to belong to a category. Under these conditions, its potential characterization The probability distribution it follows, i.e., the category The prototype distribution. The mean vector of this distribution. and diagonal covariance matrix (from variance vector) The composition is not fixed, but is learned through the network; Let be the latent representation vector, with dimension . ; Indicates that the sample belongs to the category ; It follows a Gaussian distribution (normal distribution).

[0037] The specific implementation method is as follows: First, maintain a trainable class embedding for each class. It can be initialized as a one-hot vector, where It is the dimension of the embedding vector. Each embedding... The mean and variance parameters of the prototype of this category are generated by transforming the data through two independent two-layer multilayer perceptrons (MLPs).

[0038]

[0039] here, It is a mean-coded network; For category Trainable embedding vectors, It is a variance-coded network; is the activation function. Since the variance must be positive, therefore... The output was applied Functions to ensure variance parameters The values ​​are always positive. In this way, the prototype distribution is not only learnable, but also, during training, the category embeddings are jointly optimized with the observed labels, enabling the learned prototype distribution to implicitly capture semantic correlations between different categories, such as label co-occurrence patterns. This replacement of the informationless standard normal prior with a structured, label-information-guided hybrid prior lays the foundation for subsequently constructing a semantically structured latent space.

[0040] Step S200: Obtain input data for multiple views of the sample. For each available view, obtain the single-view posterior distribution of that available view through the encoder.

[0041] The method of this invention aims to handle situations where the view is incomplete. For a view containing... A dataset of samples, each sample may be associated with There are multiple views, but due to practical reasons (such as sensor malfunction), only a portion of the view can be observed. For any given sample... Its available set of views is denoted as ,in, This represents the total number of views. In this step, for the sample... Each available view , its original input data Send in a view-specific encoder. The goal of this encoder is not to output a fixed feature vector, but rather a probability distribution, i.e., a one-view posterior distribution. Used to characterize a given input At that time, its potential representation The possibility of this. In this invention, the posterior distribution is modeled as a Gaussian distribution that is independent in each dimension:

[0042] in, For view The posterior distribution of the encoder output, i.e., given the input Time-latent representation The probability distribution; This is the set of parameters for the encoder network; For view The latent representation vector; For view The original input data; It follows a Gaussian distribution; For view The mean vector output by the mean network; For view The standard deviation vector output by the variance network; This represents the diagonal covariance matrix composed of standard deviation vectors. The physical meaning of this distribution is: given input... latent representation The possible values ​​of follow the order of Centered on, with The uncertainty is a Gaussian distribution, and the larger the variance, the lower the confidence of the model in that dimension of the feature.

[0043] and They are all neural networks; they are encoders. A portion of this is used to calculate the mean vector and diagonal standard deviation vector of the Gaussian distribution. To train this model, which includes a random sampling process, using gradient descent, a reparameterization technique is employed. Specifically, the representation... The sampling process can be represented as:

[0044] in, View obtained from sampling Potential representations; This indicates element-wise multiplication. This is a random noise vector sampled from a standard normal distribution; It follows a standard multivariate normal distribution with a mean of 0 and a covariance equal to the identity matrix. In this way, randomness is transferred to input-independent noise. This allows the model to... and The gradient is backpropagated.

[0045] Step S300: Perform first product expert fusion on the single-view posterior distributions of all available views to obtain multi-view fusion posterior, and sample from the multi-view fusion posterior to obtain a shared representation.

[0046] After obtaining the single-view posterior distribution for each available view, they need to be fused into a shared representation that integrates multi-source information. This invention employs the Product Expert (PoE) framework for fusion, which is naturally suited for handling the problem of missing views. For the current sample, the multi-view fused posterior... It is obtained by multiplying the single-view posteriors (as "experts") of all available views and then normalizing them. Since all posteriors are Gaussian distributed, the fused result is also Gaussian distributed, and its parameters have a closed-form solution:

[0047] in, For multi-view fusion posterior distribution, integrate information from all available views; The input set for all available views of the current sample; This is the mean vector of the fused posterior; This is the variance vector of the fused posterior.

[0048] The fused precision (the reciprocal of the variance) is the sum of the precisions of each view, while the fused mean is the precision-weighted average of the means of each view. The specific calculation formula is as follows:

[0049]

[0050] in, The posterior precision vector (the inverse of the variance) is fused and operated on element-wise; For view index; The set of available views for the current sample; For view The precision vector; It is a unit vector (or identity matrix) with all 1s in the corresponding dimension, serving as a weak standard normal prior to ensure numerical stability; This represents the variance vector of the fused posterior. For view The mean vector.

[0051] The additional identity matrix added to the formula This can be viewed as a weakened standard normal prior to ensure numerical stability. The core advantage of this fusion mechanism lies in its adaptive weighting based on the uncertainty (variance) of each view across each feature dimension: the smaller the variance (i.e., the higher the confidence) of a view in a certain dimension, the greater its contribution to the fusion result. Simultaneously, missing views are not included in the summation, thus eliminating the need for complex missing value imputation operations, simplifying the process and avoiding the propagation of imputation errors. Finally, a shared representation incorporating information from multiple views can be sampled from this fusion posterior. .

[0052] After obtaining the multi-view fusion posterior and shared representations, this invention further constructs a task-related learning objective. This objective needs to balance sufficiency and consistency, where sufficiency preserves task-related information used for reconstruction and classification; consistency forces the representations across different views to be consistent. Unlike existing information bottleneck methods that optimize a unified objective based on shared information, this invention decomposes it into two components: feature-level and label-level, thereby achieving independent control over reconstruction quality and classification performance.

[0053] For feature-level learning, sufficiency terms Measuring single-view representation Its input The degree of information retention ensures that the representation has sufficient reconstruction capability; consistency terms The metric, given the current view, represents Suppressing other view-specific information remaining in the image can eliminate view-specific artifacts. Feature-level objectives can be represented as:

[0054] in, For the first The feature-level learning objective for each view is used to measure the reconstructability of the view representation and the consistency between views; Mutual information represents the input data. Its single-view representation The degree of interdependence between them; These are the weighting coefficients for feature-level consistency constraints, used to balance the importance of sufficiency and consistency; Conditional mutual information indicates that, given the current view... Under the conditions, characterization Other views still included in The amount of information; This represents the input data for all other views besides the view itself.

[0055] For label-level learning, since feature-level learning may preserve low-level details while discarding semantic patterns crucial for classification, this invention further incorporates label information. (Sufficiency term) Ensure that the representation of each view is predictive of the labels; consistency item The metric measures the unique contribution of the current view representation to label prediction given other view representations; suppressing this contribution forces different views to provide consistent evidence for the labels. The label-level objective is represented as:

[0056] in, For the first The label-level learning objective for each view is used to measure the predictive ability and label consistency of the view representation. Mutual information, representing tags With view representation The degree of interdependence between them; These are the weighting coefficients for label-level consistency constraints, used to balance label sufficiency and label consistency; For conditional mutual information, it represents the representation of other views that are known. Under the condition that the current view represents For tags Its unique information contribution; Indicates except for views Potential representations of all views other than those in the view.

[0057] By integrating feature-level and label-level objectives and summing them over all available views, a unified optimization objective is obtained:

[0058] in, It is a balancing factor used to adjust the relative importance of feature information and label information in the overall objective.

[0059] Step S400: Based on the similarity between the shared representation and each prototype distribution, construct a semantic alignment loss so that the shared representation attracts the prototype distribution corresponding to its positive label and repels the prototype distribution corresponding to its negative label.

[0060] In order to learn shared representations To reflect the semantic structure constructed in step S100, this invention designs a semantic alignment loss. This loss is based on the idea of ​​contrastive learning, aiming to pull the shared representations of samples toward the prototype distribution corresponding to their positive label categories and push them away from the prototype distribution corresponding to their negative label categories. Before calculation, the mean vectors of the shared representations and prototype distributions are first processed... Normalization process: , ,in, The shared representation after L2 normalization (obtained from posterior sampling of multi-view fusion); For shared representation; It is an L2 norm; L2 normalized class Prototype mean vector; For category Prototype mean vector.

[0061] For a set of observation labels The sample, let Let represent the subset of positive labels. Then the semantic alignment loss is defined as:

[0062] in, For semantic alignment loss; For the positive label subset of the observed label set; For category indexing, iterate through the subset of positive labels of the samples. Each positive label category in; The set of observed labels for the sample (including positive and negative labels); For category indexing, iterate through the set of observed labels of the samples in the summation term of the denominator. All categories; For normalized shared representation; For category The normalized prototype distribution mean; For category The normalized prototype distribution mean vector; For temperature parameters; This is the vector transpose symbol.

[0063] For each positive category This loss encourages their similarity. The similarity relative to all observed categories of the sample (including positive and negative labels) and As large as possible. This is equivalent to maximizing the similarity of the representation to the positive class prototype while minimizing the similarity to the negative class prototype. By minimizing The geometry of the latent space is effectively organized: samples with similar labels will cluster near their corresponding shared prototypes, while samples with different labels will move away from each other, thus forming a clear clustering structure in the latent space that corresponds to the semantics of the labels.

[0064] Step S500: For each category, perform a second product expert fusion between the prototype distribution of the category and the single-view posterior distribution of each available view to obtain the conditional posterior distribution of the category, and sample the category-specific representation of the category from the conditional posterior distribution.

[0065] Existing methods typically fuse multi-view information into a category-independent shared representation, implicitly assuming that all categories benefit from the same combination of features. This invention, however, argues that different categories may exhibit heterogeneity in their dependence on views. Therefore, it introduces a "hypothesis-driven" fusion mechanism. For each category... A hypothesis is proposed: "Assume that the sample belongs to category..." Then, fusion is performed under this assumption. Specifically, for categories... Its conditional posterior distribution By using the categories obtained in step S100 prototype distribution The posterior distribution of the single view is fused with the "expert" data from all available views:

[0066] in, To assume that the sample belongs to category Under the condition of fusing all available view information and the posterior distribution after category prior; Proportional to; For category The prototype distribution is used as a priori; For view The posterior distribution of a single view; Since all distributions are Gaussian, the fusion result is also Gaussian, and its parameters can be calculated using a precision-weighted method:

[0067]

[0068] in, For category The precision matrix of the conditional posterior distribution; It is a category The accuracy matrix of the prototype distribution; For category The mean vector of the prototype distribution; Available View The precision matrix output by the encoder; Available View The mean vector output by the encoder; For category The mean vector of the conditional posterior distribution; It is the inverse of the precision matrix, i.e., the covariance matrix.

[0069] Understandably, when assessing whether a sample belongs to a category... At the same time, not only the observational evidence provided by each view is considered ( ), and also categories Expected characteristic distribution This is taken into consideration as important prior information. Fusion results and That is, from category This characterizes the "category-specific representation" of the sample and its uncertainty from this perspective. Sampling can be performed from this conditional posterior to obtain the sample's category-specific representation. Category-specific representation .

[0070] Step S600: Input the category-specific representation of each category into the dedicated classifier of the corresponding category to obtain the prediction result for each category.

[0071] After obtaining the results for each category Category-specific representation Next, it is necessary to evaluate the hypothesis that "the sample belongs to the category". "Is this true? To address this, the present invention provides a dedicated classifier for each category." To represent category-specific features The input is fed into its corresponding classifier, and the Sigmoid activation function is applied. This will give you the category to which the sample belongs. Probability prediction:

[0072] in, Specific representation for a given category When, the sample belongs to category The probability of; For class Category-specific representations of conditional posterior sampling; For a dedicated classifier, output a scalar logit; The Sigmoid activation function maps the logit to (0,1).

[0073] because Categories have been incorporated during the generation process. The prototype prior is naturally biased towards the category. The relevant feature subspace makes subsequent classifiers It allows for a more focused and effective evaluation of whether the evidence supports the hypothesis.

[0074] Step S700: Calculate the classification loss based on the prediction results and the observed labels of the samples.

[0075] Since this is an incomplete multi-label scenario, only the labels already observed in the samples can be used. To calculate the loss. Final classification prediction. (The calculation method will be detailed in step S800.) The difference between the observed and true labels is measured by a binary cross-entropy loss. This loss is calculated only for the observed labels:

[0076] in, The binary cross-entropy classification loss is used. It is a category The true label (0 or 1). This is the final prediction result. Corresponding category The probability value. It is the set of observed labels for the current sample. It is its size. Unobserved labels ( It does not participate in loss calculation, ensuring that the model can be effectively trained even when the supervision signal is incomplete.

[0077] Step S800: Jointly optimize the semantic alignment loss and the classification loss to train the classification model.

[0078] To further improve the model's robustness under conditions of missing views, this invention also designs a multi-source prediction aggregation strategy and jointly optimizes it with multiple loss functions. These loss functions together constitute the model's overall training objective.

[0079] During the prediction phase, there are multiple "prediction sources": Shared representation prediction source: from the multi-view fusion posterior of step S300 Medium-sampled shared representation Then input it into all From the specialized classifiers for each category, a set of prediction vectors is obtained. (in (This indicates the source).

[0080] Single-view conditional posterior prediction source: for each available view Construct its single-view conditional posterior Samples are taken from these samples and input into a dedicated classifier to obtain another set of prediction vectors. (in (Corresponding to each individual view).

[0081] Thus, the total obtained is There are several prediction sources. The reliability of different prediction sources may vary across different categories. Therefore, a confidence-based weighted aggregation strategy is designed to dynamically assign weights to each prediction source to obtain the final classification prediction. :

[0082] in, This is the final aggregated classification prediction vector; For the index of the prediction source, Predictions corresponding to shared representations Predictions of category-specific representations for each view; The number of available views, i.e., the number of views actually available for the current sample; For the first The prediction vector output by each prediction source; For the first The prediction vector output by each prediction source; Parameters for controlling weighted sharpness; The index variable in the summation of the denominator; Here is a confidence metric function used to evaluate the predictive certainty of a prediction source, defined as:

[0083] in, This is a confidence metric function used to evaluate the predicted vector. The degree of certainty; For prediction vectors; To iterate through all categories; For the prediction vector, the first... The predicted value for each class. This function prefers sources that give deterministic predictions close to 0 or 1 for each class, because when When approaching the boundary value, The value will decrease. In this way, the model will place more trust in the predictive sources that make definite judgments when making the final decision, thereby enhancing its robustness to incomplete views.

[0084] Secondly, define task-related learning objectives. The objective is to guide the representation learning process, balancing feature reconstruction with classification performance. It consists of feature-level loss and label-level loss, applied to each available view, respectively. superior.

[0085] For each available view Feature-level loss Defined as:

[0086] in, For feature-level loss; For decoder network parameters; first item It is the reconstruction loss, which measures the loss from the latent representation. via decoder Reconstruct the original input The log-likelihood of the second term ensures the sufficiency of the representation by maximizing this likelihood (i.e., minimizing this term). It is a single-view posterior Post-fusion with multi-view Minimizing the KL divergence between views encourages the single-view posterior to approximate the fused posterior, thereby suppressing view-specific information residues and achieving consistency. It is a weighting coefficient that balances reconstruction and consistency.

[0087] For each available view Label-level loss Defined as:

[0088] in, Label-level loss; For classifier network parameters (shared classifier, outputting C-dimensional predictions); the first term It is the classification cross-entropy loss, which measures the loss from a single-view representation. Classifier Predicted Labels The log-likelihood of the view ensures the view representation's predictive power (sufficiency) for the label. The second term... It is the predicted distribution of the fused posterior. Predicted distribution with single-view posterior Minimizing the KL divergence between views forces the predictions of each view to converge, eliminating view-specific biases. It is a weighting coefficient that balances predictive power and consistency.

[0089] The task-related learning objective is obtained by weighted summing of the feature-level loss and label-level loss across all available views.

[0090] in, Learning objectives related to the task; The set of available views for the current sample; For view indexing, iterate through all available views; To balance the weights of feature-level and label-level losses.

[0091] The design of the aforementioned loss function is not arbitrary, but rather stems from a variational derivation of the mutual information objective. Since direct computation and optimization of mutual information are infeasible, this invention derives variational lower bounds for each sufficiency term and variational upper bounds for each consistency term, thereby transforming the optimization problem into a feasible form of loss minimization.

[0092] Specifically, for the reconstruction sufficiency term, a parameterized variational decoder is introduced. We can obtain a lower bound on mutual information:

[0093] in, For input With representation Mutual information; To calculate the expectation of the data distribution; To calculate the expectation of the posterior distribution; To reconstruct the log-likelihood.

[0094] The physical meaning of this lower bound is that from the encoder Characterization of sampling via decoder Reconstructing the log-likelihood expectation of the original input, maximizing this lower bound is equivalent to improving the reconstruction quality, which is consistent with... The first reconstruction loss corresponds to this.

[0095] For the encoder consistency term, an upper bound can be constructed using the KL divergence between the fused posterior and the single-view posterior:

[0096] in, For a given Under certain conditions, the mutual information represented by other views and the current view; To fuse the KL divergence between the posterior and the single-view posterior; This is the posterior for multi-view fusion based on product experts. Minimizing this upper bound causes each individual view posterior to approximate the fusion posterior, thereby eliminating view-specific information remnants, which is consistent with... The second term in the equation corresponds to the KL divergence.

[0097] Similarly, for the label sufficiency term, a parameterized classifier is introduced. Deriving the lower bound:

[0098] in, Label With View Mutual information in representation; For the classification log-likelihood.

[0099] correspond The first term in the loss is classification. For the classifier consistency term, an upper bound is constructed by comparing the differences between the label prediction distributions generated by the fused representation and the single-view representation:

[0100] in, Given other view representations, the current view representation provides unique information about the label; The KL divergence is calculated for the fusion prediction and single-view prediction distributions.

[0101] correspond The second term in the equation is the KL divergence. Through this series of variational derivations, this invention concretizes the abstract information-theoretic goal into a computable and optimizable neural network loss function.

[0102] Finally, the total loss function is constructed. The model's total loss function is composed of the task-related learning objectives. Semantic alignment loss and classification loss Composed of three weighted parts:

[0103] in, This is the total loss function; Learning objectives related to the task; For semantic alignment loss; For classification loss; Hyperparameters for balancing semantic alignment strength.

[0104] By jointly optimizing this total loss, the model can achieve semantic prior at the representation level ( ) and information theory constraints ( Guidance is provided at the decision-making level through confidence-weighted () The computational process enhances robustness, thereby systematically integrating label semantics throughout the entire process of representation extraction and fusion, effectively addressing the dual challenges of missing views and incomplete labels.

[0105] During training, all components (including the encoder) decoder Prototype Generative Network , Dedicated classifier All (etc.) are jointly optimized in an end-to-end manner. During the testing phase, deterministic prototype means are used. As a fixed semantic anchor, prediction is performed according to the steps described above.

[0106] Step S900: Obtain input data for multiple views of the sample to be classified, wherein the multiple views of the sample to be classified correspond to different modal data, and the modal data includes at least two of image modal, text modal, audio modal and numerical sensor data modal; In this step, input data from multiple views of the sample to be classified is first acquired. This input data corresponds to various heterogeneous data modalities. Specifically, the image data can be RGB color images or infrared thermal images captured by a camera, or computed tomographic images such as X-rays, CT scans, and MRI scans acquired by medical imaging equipment; the text data can be unstructured text data such as web page content composed of natural language, social media comments, medical records, or diagnostic reports; the audio data can be voice interaction signals, ambient background sounds, or bioacoustic signals; and the sensor numerical data can be multi-channel time-series numerical signals acquired at continuous time points by devices such as accelerometers, gyroscopes, temperature sensors, pressure sensors, or electrocardiogram monitoring equipment.

[0107] It should be noted that the above list of data types is only an illustrative example of the applicable scenarios of the present invention. The scope of protection of the present invention is not limited to these specific data carriers or physical signal types. As long as the digital data can provide multi-source heterogeneous information for classification tasks, it belongs to the view input data referred to in the present invention.

[0108] Step S1000: Input the sample to be classified into the trained classification model to obtain the final classification result of the sample to be classified.

[0109] In this step, after the sample to be classified is input into the pre-trained classification model, for each available view of the sample, the input data of that view is encoded using the pre-trained encoder to generate the corresponding single-view posterior distribution. Subsequently, the single-view posterior distributions of all available views are fused using a first product expert fusion to obtain the multi-view fused posterior corresponding to the shared representation that integrates the information of all available views, and the shared representation is sampled from the multi-view fused posterior.

[0110] The obtained category-specific representations for each category of the sample to be classified are input into the corresponding trained specialized classifiers. Each specialized classifier judges the received category-specific representations to evaluate whether the hypothesis that "the sample to be classified belongs to the corresponding category" holds true, and outputs the predicted probability value of the sample belonging to the corresponding category. After repeating this prediction operation for all categories, the classification prediction result of the sample to be classified in the entire category space is obtained, and the trained classification model outputs the final multi-label classification result of the sample to be classified.

[0111] Furthermore, as a more preferred embodiment of the present invention, when outputting the final result, the trained classification model can also be invoked to perform a multi-source prediction aggregation step, and the prediction results obtained from the shared representation prediction source and each single-view conditional posterior prediction source can be aggregated with confidence weighting to generate the final classification prediction result with optimal robustness.

[0112] This invention constructs an end-to-end unified training framework. This framework jointly optimizes three objectives: semantic alignment loss, task-related learning loss, and classification loss. At the representation level, it guides the geometric structure through semantic priors; at the information flow level, it constrains sufficiency and consistency through task-related learning; and at the decision level, it aggregates multi-source predictions through confidence-weighted aggregation. This design systematically integrates label semantics throughout the entire process of representation extraction and fusion, enabling the model to utilize the semantic structural information contained in some labels, overcoming the limitation of existing methods that only use labels as output layer supervision signals.

[0113] Unlike standard variational frameworks that use no-information priors (standard normal distribution), the semantic-aware hybrid prior of this invention gives the latent space an explicit semantic geometric structure. Unlike existing methods that optimize a unified objective based on shared information, the task-related learning objective of this invention decomposes the information bottleneck into two components: feature-level and label-level, enabling independent and fine-grained control over reconstruction quality and classification performance. Unlike existing methods that aggregate multi-view information into a single shared representation, the category-specific conditional posterior fusion mechanism of this invention allows different categories to use different feature combinations for reasoning.

[0114] Example 1: In industrial intelligent manufacturing applications, the data object described in this invention refers to multi-source sensor time-series data deployed on large mechanical equipment (such as wind turbines and industrial robots). The views include, but are not limited to, acceleration time-domain signals from vibration sensors, operating condition curve data from temperature / pressure sensors, and thermal / visible light images of rotating components captured by industrial cameras. These three views describe the equipment's operating status from three dimensions: mechanical motion, environmental thermodynamics, and visual appearance. The labels to be classified represent different types of equipment failures, such as rotor imbalance, bearing wear, gear tooth breakage, and poor lubrication.

[0115] In industrial settings, due to sensor aging or network fluctuations, incomplete views and labels are common occurrences. For example, during a routine inspection, a vibration sensor malfunction might result in the loss of vibration signals, leaving only temperature data and thermal images. Simultaneously, maintenance records may only list bearing wear as a fault, omitting the concurrent lubrication issues.

[0116] When applying the method of this invention, the training phase is implemented as follows: First, learnable prototype distributions are constructed for fault categories such as "rotor imbalance" and "bearing wear". Then, a neural network is used to map the embedding vectors of the fault categories to Gaussian distribution parameters.

[0117] Next, for samples containing vibration signals, temperature data, and thermal imaging, their respective latent features are extracted using a dedicated encoder to obtain their individual single-view posterior distributions. The posteriors of the available views are then subjected to a first product expert fusion to sample a shared representation that characterizes the overall operating state of the equipment.

[0118] Then, the semantic alignment loss is calculated by using the similarity between the shared representation and the prototype of each fault category, which forces the shared representation to be closer to its actual fault prototype and farther away from the fault prototype that has not occurred in the latent space, thereby constructing the fault semantic space.

[0119] Subsequently, for the fault category of "bearing wear", the prototype distribution of this category is fused with the currently available temperature posterior and thermal imaging posterior by a second product expert fusion to sample the category-specific characterization of "bearing wear", which is then input into the corresponding dedicated classifier to obtain the probability prediction of bearing wear occurring in the device.

[0120] Finally, the semantic alignment loss, task relevance loss, and classification loss are jointly optimized to complete the training of the industrial fault diagnosis model.

[0121] During the classification and prediction phase, when real-time operational data of an industrial robot lacking vibration signals (only temperature curves and thermal imaging) is collected, the trained model automatically completes data encoding and fusion with various fault prototypes, ultimately outputting the probability values ​​of various fault types such as "rotor imbalance", "bearing wear", "gear tooth breakage" and "poor lubrication", providing reliable technical data support for preventive maintenance.

[0122] Example 2: In the application scenario of constructing product graphs on e-commerce platforms, the data object described in this invention is the product details page information of the e-commerce platform. The views include, but are not limited to, the visual features (images) of the product's main image, the descriptive text (text) of the product details page, and the product's technical parameter table (structured numerical / sequence data). These three views characterize the product from the dimensions of visual appearance, semantic description, and physical specifications, respectively. The tags to be classified are the product's multi-dimensional attributes, such as clothing style (casual, business, sports), applicable occasions (daily, commuting, outdoor), and main materials (cotton, synthetic fiber, blended fabrics), etc.

[0123] In actual e-commerce data processing, incompleteness is common. For example, some merchants fail to fill out product parameter forms (missing structured data); or the uploaded product images are unclear / invalid (missing visual data), with only text descriptions provided.

[0124] When applying the method of this invention, the training phase is as follows: First, construct prototype distributions for each attribute category, such as "casual style" and "outdoor occasion".

[0125] Next, for samples containing visual, textual, and parameter tables, single-view posteriors are extracted using their respective encoders, and shared representations are obtained through first PoE fusion. Semantic alignment loss is used to cluster "commuter and blended" product representations near their corresponding prototypes.

[0126] Next, for the "business style" tag, its prototype distribution is fused with the currently available views (such as only images and text, missing parameter tables) to sample the category-specific representation of whether the product is "business style".

[0127] Because the prototype prior of "business style" acts as an "attention guide," the category-specific representation adaptively relies more on the keywords "suit" and "tie" in the text description, as well as the silhouette of the suit in the image, rather than the missing parameter table. Predictions are obtained through a dedicated classifier, and the model is trained by combining the observed labels.

[0128] In the classification prediction stage, for a new product that only contains images and text and lacks a technical parameter table, the trained model and attribute prototype can accurately output the multi-attribute classification results of the product in different dimensions such as style, occasion, and material, which greatly improves the construction efficiency of e-commerce knowledge graph.

[0129] The specific applications of this invention are not limited to the two situations described above, and this invention does not limit them.

[0130] According to another aspect of the embodiments of this application, an electronic device is also provided, including a processor and a memory, wherein the processor is configured to implement the steps of the method when executing a computer program stored in the memory.

[0131] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0132] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0133] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0134] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0135] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A label-guided representation learning method for incomplete multi-view, multi-label classification, characterized in that, The steps include the following: A learnable prototype distribution is constructed for each category, and the prototype distribution is used to encode the semantic information of the corresponding category; The input data of multiple views of the sample is obtained. For each available view, the single-view posterior distribution of that available view is obtained through the encoder. The single-view posterior distributions of all available views are subjected to first product expert fusion to obtain multi-view fusion posterior, and shared representations are obtained by sampling from the multi-view fusion posterior. Based on the similarity between the shared representation and each prototype distribution, a semantic alignment loss is constructed so that the shared representation attracts the prototype distribution corresponding to its positive label and repels the prototype distribution corresponding to its negative label. For each category, the prototype distribution of the category is fused with the single-view posterior distribution of each available view using a second product expert fusion to obtain the conditional posterior distribution of the category, and the category-specific representation of the category is sampled from the conditional posterior distribution. The category-specific representation of each category is input into the classifier of the corresponding category to obtain the prediction result for each category; The classification loss is calculated based on the prediction results and the observed labels of the samples; The semantic alignment loss and the classification loss are jointly optimized to train a classification model; Input data is obtained from multiple views of a sample to be classified, wherein each view of the sample to be classified corresponds to different modal data, and the modal data includes at least two of the following: image modality, text modality, audio modality, and numerical sensor data modality. The sample to be classified is input into the trained classification model to obtain the final classification result of the sample to be classified. The second product expert fusion uses a precision-weighted approach, where the mean and precision of the conditional posterior distribution are calculated for each category: in, For category The precision matrix of the conditional posterior distribution; It is a category The accuracy matrix of the prototype distribution; For category The mean vector of the prototype distribution; Available View The precision matrix output by the encoder; Available View The mean vector output by the encoder; For category The mean vector of the conditional posterior distribution; It is the inverse of the precision matrix, i.e., the covariance matrix.

2. The label-guided representation learning method for incomplete multi-view, multi-label classification as described in claim 1, characterized in that, The prototype distribution is a Gaussian distribution, and its mean and variance parameters are obtained by mapping the learnable category embeddings through a neural network. The learnable category embeddings are initialized as one-hot vectors. Each category embedding is transformed by a mean encoding network and a variance encoding network to generate the mean and variance parameters of the corresponding prototype distribution. The variance parameter is transformed by the softplus function to ensure that it is a positive value.

3. The label-guided representation learning method for incomplete multi-view, multi-label classification as described in claim 1, characterized in that, The semantic alignment loss is constructed based on contrastive learning. The normalized shared representation and the normalized prototype distribution mean are calculated to maximize the similarity between the shared representation and the prototype distribution mean corresponding to the positive label, and minimize the similarity with the prototype distribution mean corresponding to the negative label; the semantic alignment loss is: in, For semantic alignment loss; For the positive label subset of the observed label set; For category indexing, iterate through the subset of positive labels of the samples. Each positive label category in; The set of observation labels for the samples; For category indexing, iterate through the set of observed labels of the samples in the summation term of the denominator. All categories; For normalized shared representation; For category The normalized prototype distribution mean; For category The normalized prototype distribution mean vector; For temperature parameters; This is the vector transpose symbol.

4. The label-guided representation learning method for incomplete multi-view multi-label classification as described in claim 1, characterized in that, The joint optimization also includes task-related learning objectives, which include feature-level constraints and label-level constraints; wherein: The feature-level constraints are used to balance the sufficiency of each view representation in reconstructing its input data with cross-view consistency. The label-level constraints are used to balance the sufficiency of label predictions for each view representation with the consistency of label predictions between views. The task-related learning objective is jointly optimized with the semantic alignment loss and the classification loss.

5. The label-guided representation learning method for incomplete multi-view, multi-label classification as described in claim 4, characterized in that, The feature-level constraints are achieved by minimizing the feature-level loss for each available view: in, Available View Feature-level loss; Available view; For reconstruction loss; The KL divergence between the single-view posterior and the multi-view fusion posterior; For expectation operators; Available View The posterior distribution of the encoder output, with parameters as follows. Given input Time-latent representation Conditional distribution; Let log-likelihood be the output of the decoder, representing the value derived from the latent representation. Reconstruct the original input The logarithm of the probability; These are weighting coefficients that control the strength of consistency constraints; The Kullback-Leibler divergence; For multi-view fusion posterior distribution; For currently available views The posterior distribution of a single view; The label-level constraint is achieved by minimizing the label-level loss: in, Available View Label-level loss; For classification cross-entropy; To calculate the KL divergence between the predicted distributions of the fused posterior and the single-view posterior; The log-likelihood of the classifier output, with parameters as follows: , indicating from the available views The representation Predicted Labels The logarithm of the probability; For shared representations obtained from posterior sampling of multi-view fusion The predicted distribution of labels obtained by the classifier; From the available views The representation obtained by single-view posterior sampling The predicted distribution of labels obtained by the classifier; This is a weighting coefficient that controls the strength of the label consistency constraint; The task-related learning objective is a weighted sum of the feature-level loss and the label-level loss across all views: in, Learning objectives related to the task; The set of available views for the current sample; For view indexing, iterate through all available views; It is a balancing factor.

6. The label-guided representation learning method for incomplete multi-view, multi-label classification as described in claim 1, characterized in that, The method further includes a multi-source prediction aggregation step: the prediction results obtained by classifying the category-specific representations corresponding to each view and the prediction results obtained by classifying the shared representations are used as multiple prediction sources; Confidence weights are calculated based on the predictive certainty of each prediction source. The prediction results from each source are then weighted and aggregated to obtain the final classification prediction. The confidence weights are calculated in the following way: in, This is the final aggregated classification prediction vector; For the index of the prediction source, Predictions corresponding to shared representations Predictions of category-specific representations for each view; The number of available views, i.e., the number of views actually available for the current sample; For the first The prediction vector output by each prediction source; For the first The prediction vector output by each prediction source; Parameters for controlling weighted sharpness; The index variable in the summation of the denominator; Here is a confidence metric function used to evaluate the predictive certainty of a prediction source, defined as: in, This is a confidence metric function used to evaluate the predicted vector. The degree of certainty; For prediction vectors; To iterate through all categories; For the prediction vector, the first... The predicted value of the class.

7. The label-guided representation learning method for incomplete multi-view multi-label classification as described in claim 1, characterized in that, The joint optimization employs end-to-end training, and the total loss function is: in, This is the total loss function; Learning objectives related to the task; For semantic alignment loss; For classification loss; Hyperparameters for balancing semantic alignment strength.

8. The label-guided representation learning method for incomplete multi-view multi-label classification as described in claim 5, characterized in that, In the process of constructing the feature-level constraints and label-level constraints, the information-theoretic objective is transformed into an optimizable loss term based on variational inference; For feature sufficiency and label sufficiency, variational lower bounds are derived by introducing variational decoders and classifiers, respectively, which are transformed into reconstruction loss and classification cross-entropy loss in the feature-level loss and label-level loss, respectively. For feature consistency and label consistency, variational upper bounds are derived by using the KL divergence between the single-view posterior and the multi-view fusion posterior, and the KL divergence between the corresponding prediction distributions, respectively. These are then converted into KL divergence constraint terms in the feature-level loss and label-level loss.

9. The label-guided representation learning method for incomplete multi-view multi-label classification as described in claim 2, characterized in that, The classification model includes a training phase and a testing phase: In the training phase, the prototype distributions are sampled to jointly optimize the classification model parameters, enabling end-to-end joint optimization of the category embedding, encoder, and classifier; In the testing phase, the deterministic mean of each prototype distribution mapped by the mean encoding network is used as a fixed semantic anchor, and the output of each view encoder is a deterministic representation, which participates in the accuracy-weighted calculation and final prediction of the second product expert fusion.

Citation Information

Patent Citations

  • Self-adaptive enhanced adversarial hash unsupervised cross-modal retrieval method

    CN117851659A

  • Unsupervised multi-view clustering method for problem of noisy clustering information

    CN119068224A