Incomplete Multi-View Multi-Label Classification Method and System Based on Deep Contrastive Learning
The incomplete multi-view multi-label classification network model constructed through in-depth comparison learning solves the problem of missing views and labels, improves classification accuracy and generalization capabilities, and is suitable for real-time inference of multi-view data.
Patent Information
- Application Number
- CN202211593949.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-13
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-12-13
AI Technical Summary
In the incomplete multi-view multi-label classification task, it is difficult to effectively deal with the problem of missing views and labels. Traditional methods rely on manual design feature extraction rules, are difficult to generalize, and performance depends heavily on parameter settings.
Using a method based on deep contrast learning, an incomplete multi-view multi-label classification network model is constructed, including a specific view representation learning framework, an incomplete instance-level contrast learning module and a weighted fusion and multi-label classification module. The autoencoder is used to extract features and enhance consistency through contrast learning, and unsupervised contrast learning is introduced to balance the importance of view.
It can effectively handle the number and missing views of multi-view data, improve classification accuracy, adapt to the random lack of multi-label supervision information, and has real-time reasoning capabilities, which is suitable for production environments.
Smart Images

Figure CN115994317B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pattern recognition, and in particular to an incomplete multi-view multi-label classification method and system based on deep contrast learning. Background Art
[0002] In recent years, with the explosive growth of data collection and feature extraction methods, it has become difficult to meet the increasingly complex comprehensive analysis requirements by only describing, analyzing, and processing samples from a single perspective. Multi-view data collected from multiple sources can describe the observed object more comprehensively and accurately. Some methods use adversarial loss and label loss to learn the shared semantics of multiple views. Other methods obtain predicted labels by maximizing the correlation between the latent space, feature space, and label space. Another class of matrix factorization-based methods align the semantic space by maximizing the dependence of the basis matrices of different views in the kernel space. It should be noted that all of these methods are invariably based on an unreasonable premise of complete data. However, in real practice, the data used for multi-view multi-label classification is often incomplete. On the one hand, the feature data collected from multiple sources may have missing views due to various reasons. For example, the media forms of files in some archives may include text, audio, video, etc. These information media regarded as different views do not exist universally in all archives. Therefore, the multi-view feature data extracted from them naturally has missing views. On the other hand, since it is difficult and costly to manually label all labels, label information in real data often has varying degrees of missing, which is more common in datasets with a large number of highly correlated labels. Based on this, different from existing methods that only consider missing views or missing labels, it aims to handle the problem of double missing of labels and views, that is, the problem of random multi-view feature data missing and multi-class label missing. Some scholars combine the incomplete multi-view learning model based on matrix factorization and the multi-label prediction model based on label correlation, connect the feature space and semantic space by learning a common representation, and impose a low-rank constraint on the label correlation matrix to enhance the robustness of the prediction model.
[0003] Although these traditional methods have achieved certain results in the field of incomplete multi-view multi-label, however, this learning mode that requires manual design of feature extraction rules and is difficult to generalize limits the further development of incomplete multi-view and multi-label learning. Deep neural networks are increasingly applied to feature extraction and data analysis tasks. On the one hand, traditional methods based on matrix factorization, spectral clustering, or kernel learning only act on exploring shallow features of data, while complex data analysis tasks often require capturing higher-level semantic representations relative to the original data. On the other hand, the performance of traditional multi-view learning models highly depends on the setting of parameters, and usually requires searching for the optimal parameter combination for different datasets. Summary of the Invention
[0004] In view of the above problems, the present invention provides an incomplete multi-view multi-label classification method, system and storage medium based on deep contrast learning. In addition to using a deep neural network as a framework, the classification method improves the discriminative ability of the extracted features through contrast learning, thereby improving the network classification performance.
[0005] In a first aspect of the present invention, there is provided an incomplete multi-view multi-label classification method based on deep contrast learning, the method comprising the following steps:
[0006] Construct an incomplete multi-view multi-label classification network model;
[0007] Train the incomplete multi-view multi-label classification network model;
[0008] Input test data into the trained incomplete multi-view multi-label classification network model for inference, and output predicted labels;
[0009] Among them, the incomplete multi-view multi-label classification network model includes three sub-modules: a specific view representation learning framework, an incomplete instance-level contrast learning module, and a weighted fusion and incomplete multi-label classification module. The specific view representation learning framework uses an autoencoder to extract features and reconstruct the original data. The autoencoder includes an encoder and a decoder. The encoder is used to extract features, and the decoder is used to reconstruct the original data. The incomplete instance-level contrast learning module is used to apply an incomplete instance-level contrast loss to the features extracted by the encoder to enhance the consistency of multi-view representations. The weighted fusion and incomplete multi-label classification module is used to perform weighted fusion of multi-views and calculate multi-label classification scores using the weighted fusion results to obtain the inference results of multi-label classification.
[0010] A further technical solution of the present invention is that in the specific view representation learning framework, for the view input data X (v) , the encoder E (v) extracts the corresponding feature Z (v) = E (v) (X (v) ), and the decoder D (v) decodes Z (v) to obtain the reconstructed
[0011] Apply a squared loss function at the output end of the decoder to make the reconstructed have a small error with the original X (v) :
[0012]
[0013] Where l represents the number of views, n represents the number of samples, and m v represents the feature dimension of view v, represents the reconstructed features of sample i, represents the original view input data of sample i, represents the a priori missing view indicator matrix, Indicates that the v-th view of the i-th sample is missing, Indicates that the v-th view of the i-th sample is available.
[0014] A further technical solution of the present invention is that the incomplete instance-level contrastive learning module uses a contrastive learning method to guide the encoder to extract consistent features, specifically including:
[0015] For l views, there are l×n instances, and any of them There are l×n-1 instances and instances Constitute an instance pair, and instance Instance pairs belonging to the same sample is a positive instance pair, and the rest are negative instance pairs i and j both represent the number of samples, v and u both represent the number of views, and the incomplete instance-level contrastive learning module is used to shorten the distance between positive instance pairs and expand the distance between negative instance pairs. The distance metric function Using cosine similarity, the expression is as follows:
[0016]
[0017] Where <·> represents the dot product operation. For any two views, the incomplete contrastive learning loss function is for:
[0018]
[0019] Where l represents the number of views, n represents the number of samples, τ represents the degree of diffusion of the control distribution, represents the a priori missing view indicator matrix, Indicates that the v-th view of the i-th sample is missing, Indicates that the vth view of the i-th sample is available, combined with Incomplete contrastive loss function for all views for:
[0020]
[0021] A further technical solution of the present invention is: in the weighted fusion and incomplete multi-label classification module, the fusion representation H of l views is calculated, and for the fusion representation h of each sample i , the specific expression is:
[0022]
[0023]
[0024] where l represents the number of views, v represents the view number, represents the feature representation obtained after the instance of the v-th view of sample i passes through the feature encoder E (v) , represents the prior missing view indication matrix, indicates that the v-th view of the i-th sample is missing, indicates that the v-th view of the i-th sample is available, and i represents the number of samples.
[0025] A further technical solution of the present invention is: the weighted fusion and incomplete multi-label classification module calculates the multi-label classification score by using the weighted fusion result, specifically including:
[0026] Perform a linear activation operation on the fusion representation H, and the activation function is the Sigmoid function:
[0027]
[0028] where ω represents the parameter of the fully connected layer , P represents the predicted class score, and the multi-label classification loss function used is the weighted BCE loss, and the expression is as follows:
[0029]
[0030] where is the introduced label missing indication matrix, represents the label missing indication category, n represents the number of samples, c is the number of label categories, and Y i,j represents the predicted label category of the j-th view of the i-th sample, and P i,j represents the predicted class score of the j-th view of the i-th sample.
[0031] In the second aspect of the present invention, an incomplete multi-view multi-label classification system based on deep contrast learning includes:
[0032] A network model construction unit for constructing an incomplete multi-view multi-label classification network model;
[0033] A network model training unit for training the incomplete multi-view multi-label classification network model;
[0034] A prediction unit for inputting test data into the trained incomplete multi-view multi-label classification network model for inference and outputting predicted labels;
[0035] Wherein, the incomplete multi-view multi-label classification network model includes three sub-modules: a specific view representation learning framework, an incomplete instance-level contrast learning module, and a weighted fusion and incomplete multi-label classification module. The specific view representation learning framework uses an autoencoder to extract features and reconstruct the original data. The autoencoder includes an encoder and a decoder. The encoder is used to extract features, and the decoder is used to reconstruct the original data. The incomplete instance-level contrast learning module is used to impose an incomplete instance-level contrast loss on the features extracted by the encoder to enhance the consistency of multi-view representations. The weighted fusion and incomplete multi-label classification module is used to perform weighted fusion of multiple views and calculate multi-label classification scores using the weighted fusion results to obtain the inference results of multi-label classification.
[0036] In a third aspect of the present invention, there is provided an incomplete multi-view multi-label classification system based on deep contrast learning, including: a processor; and a memory. Wherein, a computer executable program is stored in the memory, and when the computer executable program is executed by the processor, the above-mentioned incomplete multi-view multi-label classification method based on deep contrast learning is executed.
[0037] In a fourth aspect of the present invention, a storage medium stores a program, and when the program is executed by a processor, the processor is caused to execute the above-mentioned incomplete multi-view multi-label classification method based on deep contrast learning.
[0038] An incomplete multi-view multi-label classification method, system and storage medium provided by the present invention propose a deep contrast network for the dual incomplete multi-view multi-label classification problem. Different from traditional methods, the present invention focuses on using a deep neural network to extract high-level semantic representations of samples and uses an autoencoder to construct an end-to-end multi-view feature extraction framework to learn the representation vectors of samples. In addition, in order to further improve the representation ability of the model, the present invention introduces unsupervised contrast learning to guide the encoder to extract high-level representation information of multiple views according to the consistency hypothesis. At the same time, the present invention proposes a weighted fusion method to balance the importance of different views.
[0039] In summary, the beneficial effects of the present invention are mainly as follows:
[0040] 1) The incomplete multi-view multi-label classification network model proposed by the present invention has no additional restrictions on the number of views and the missing situation of multi-view data, that is, it can process multi-view data sets with a large number of views and any missing situations, and can also adapt to the situation where multi-label supervision information is randomly missing.
[0041] 2) The incomplete instance-level contrast loss proposed by the present invention can effectively aggregate cross-view features, enabling instances of the same sample in different views to satisfy the multi-view consistency assumption, thereby enhancing the high-level feature representation ability and improving the classification accuracy.
[0042] 3) The incomplete multi-view multi-label classification network model proposed by the present invention has good application characteristics. After training, it can be deployed in a production environment, and can immediately give inference results for the input incomplete multi-view test data. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is a schematic flowchart of an incomplete multi-view multi-label classification method based on deep contrast learning in Embodiment 1 of the present invention;
[0044] Figure 2 is a schematic structural diagram of an incomplete multi-view multi-label classification network model in Embodiment 1 of the present invention;
[0045] Figure 3 is a schematic flowchart of the training and inference of an incomplete multi-view multi-label classification network model in Embodiment 1 of the present invention;
[0046] Figure 4 is a schematic structural diagram of an incomplete multi-view multi-label classification system based on deep contrast learning in Embodiment 2 of the present invention;
[0047] Figure 5 is the architecture of a computer device in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the present invention, rather than limiting the present invention. In addition, it should be noted that only parts related to the present invention rather than all structures are shown in the drawings for the sake of convenience of description.
[0049] Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts depict the steps as sequential processes, many of the steps can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the figures. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0050] Embodiments of the present invention are directed to an incomplete multi-view multi-label classification method, system, and storage medium based on deep contrast learning, and provide the following embodiments:
[0051] Based on Embodiment 1 of the present invention
[0052] This embodiment is used to illustrate the incomplete multi-view multi-label classification method based on deep contrast learning. Refer to Figure 1 , which is a schematic diagram of the process of the incomplete multi-view multi-label classification method based on deep contrast learning, and specifically includes the following steps:
[0053] S110. Construct an incomplete multi-view multi-label classification network model;
[0054] S120. Train the incomplete multi-view multi-label classification network model;
[0055] S130. Input the test data into the trained incomplete multi-view multi-label classification network model for inference, and output the predicted labels;
[0056] Among them, as Figure 2 shown, the incomplete multi-view multi-label classification network model includes three sub-modules: a specific view representation learning framework, an incomplete instance-level contrast learning module, and a weighted fusion and incomplete multi-label classification module. The specific view representation learning framework uses an autoencoder to extract features and reconstruct the original data. The autoencoder includes an encoder and a decoder. The encoder is used to extract features, and the decoder is used to reconstruct the original data. The incomplete instance-level contrast learning module is used to apply an incomplete instance-level contrast loss to the features extracted by the encoder to enhance the consistency of the multi-view representation. The weighted fusion and incomplete multi-label classification module is used to perform weighted fusion of multiple views and calculate the multi-label classification score using the weighted fusion result to obtain the inference result of multi-label classification.
[0057] In the specific implementation process, first define the problem as follows. Given data Prior missing view indication matrix and prior missing label indication matrix Predict label Y ∈ {0, 1} n×cwhere \(l\) is the number of views, \(n\) is the number of samples, \(m\) v is the feature dimension of view \(v\), and \(c\) is the number of label categories. The matrix element indicates that the \(j\)-th view of the \(i\)-th sample is missing, indicating that the corresponding view is available. Similarly, indicates that the \(j\)-th label of the \(i\)-th sample is available, and vice versa for the missing category. The incomplete multi-view multi-label classification network model requires optimizing the model parameters in the training set with incomplete labels and inferring and predicting labels for the test set data.
[0058] Specifically, in the specific view representation learning framework, an autoencoder is used to extract high-level features. The autoencoder consists of a set of encoders and a set of decoders, which are used to extract high-level features and reconstruct the original data respectively. Each view has a corresponding encoder-decoder pair, which is used to independently capture the high-level discriminative features of the specific view. For the input data \(X\) of a certain view (v) , the encoder \(E\) (v) extracts the corresponding high-level features \(Z\) (v) = \(E\) (v) (\(X\) (v) ). Specifically, the encoder-decoder consists of several fully connected layers (FC), activation layers (ReLU activation function), and batch normalization layers (Batch Normalization, BN). For the encoder, its composition sequence is: FC; ReLU; FC; ReLU; FC; ReLU; FC; BN. For the decoder, its composition sequence is: FC; ReLU; FC; ReLU; FC; ReLU; FC. The decoder decodes \(Z\) (v) to obtain the reconstructed To make the reconstructed as close as possible to the original \(X\) (v) , a squared loss function is applied at the output end of the decoder to make the reconstructed have a small error with the original \(X\) (v) :
[0059]
[0060] where \(l\) represents the number of views, \(n\) represents the number of samples, \(m\) v represents the feature dimension of view \(v\), represents the reconstructed feature of sample \(i\), represents the original view input data of sample \(i\), represents the prior missing view indicator matrix, indicates that the \(v\)-th view of the \(i\)-th sample is missing, Indicates that the vth view of the i-th sample is available. This loss function introduces The significance of is to avoid the additional error caused by missing views, which enables the representation learning framework to handle arbitrary missing view data.
[0061] Furthermore, in order to increase the consistency of the extracted high-level representation, an incomplete instance-level contrastive learning loss is proposed. Specifically, the same sample has different representations in different views, that is, different instances. Figure 1 Consistency requires that these different instances should have consistent semantic expressions. Based on this, the contrastive learning method is used in the incomplete instance-level contrastive learning module to guide the encoder to extract consistent features, specifically including:
[0062] For l views, there are l×n instances, and any of them There are l×n-1 instances and instances Constitute an instance pair, and instance Instance pairs belonging to the same sample is a positive instance pair, and the rest are negative instance pairs i and j both represent the number of samples, v and u both represent the number of views, and the incomplete instance-level contrastive learning module is used to shorten the distance between positive instance pairs and expand the distance between negative instance pairs. The distance metric function Using cosine similarity, the expression is as follows:
[0063]
[0064] Where <·> represents the dot product operation. For any two views, the incomplete contrastive learning loss function is for:
[0065]
[0066] Where l represents the number of views, n represents the number of samples, τ represents the degree of diffusion of the control distribution, represents the a priori missing view indicator matrix, Indicates that the v-th view of the i-th sample is missing, indicates that the v-th view of the i-th sample is available, The purpose of introducing is to eliminate the negative impact of missing instances. Incomplete contrastive loss function for all views for:
[0067]
[0068] Furthermore, in the weighted fusion and incomplete multi-label classification module, the fusion representation H of l views is calculated for the fusion representation h of each sample i , and the specific expression is as follows:
[0069]
[0070]
[0071] where l represents the number of views, v represents the view number, represents the feature representation obtained after the instance of the v-th view of sample i passes through the feature encoder E (v) . represents the prior missing view indication matrix, indicates that the v-th view of the i-th sample is missing, indicates that the v-th view of the i-th sample is available, and i represents the number of samples.
[0072] Furthermore, the weighted fusion and incomplete multi-label classification module calculates the multi-label classification score using the weighted fusion result, specifically including:
[0073] Performing a linear activation operation on the fusion representation H, and the activation function is the Sigmoid function:
[0074]
[0075] where ω represents the parameters of the fully connected layer , P represents the predicted class score. It should be noted that the multi-label classification method adopted in the present invention comes from the common binary cross entropy loss (BCE). Different from the ordinary BCE loss, in order to adapt to the incomplete labels, the multi-label classification loss function is the weighted BCE loss, and the expression is as follows:
[0076]
[0077] where is the introduced label missing indication matrix, and its purpose is to avoid the influence of non-existent labels on the BCE loss calculation and the model backpropagation, represents the label missing indication category, n represents the number of samples, c is the number of label categories, Y i,j represents the predicted label category of the j-th view of the i-th sample, and P i,j represents the predicted class score of the j-th view of the i-th sample.
[0078] To sum up, the overall loss function is:
[0079]
[0080] where β and γ are penalty coefficients.
[0081] The following gives a specific example of the application of an incomplete multi-view multi-label classification network model, as Figure 3 shown below:
[0082] Model training:
[0083] 1. Training preparation stage:
[0084] 1) Prepare multi-view data Missing view indicator matrix Missing label indicator matrix Training set label Y;
[0085] 2) Set hyperparameters τ, β, γ and training stop threshold σ;
[0086] 3) Fill all missing views and missing labels with '0';
[0087] 4) Initialize network model parameters;
[0088] 5) Set the previous round of loss
[0089] 2. Training stage
[0090] 1) The encoder calculates the specific view representation and calculates the loss according to Equation (1)
[0091] 2) Calculate the incomplete instance-level contrast loss according to Equations (2), (3), and (4)
[0092] 3) Calculate the fused representation feature H according to Equation (5);
[0093] 4) Calculate the prediction result P according to Equation (6) and calculate the multi-label classification loss according to Equation (7)
[0094] 5) Calculate the total loss according to Equation (8) If is less than σ, then go to step 6), otherwise return to step 1);
[0095] 6) Output the prediction result P.
[0096] Model testing:
[0097] 3. Testing preparation stage:
[0098] 1) Prepare multi-view data Missing View Indicator Matrix
[0099] 2) Fill all missing views with '0'.
[0100] 3) Load the trained network model parameters.
[0101] 4. Testing Phase
[0102] 1) The encoder calculates the specific view representation
[0103] 2) Calculate the fused representation feature H according to Equation (5).
[0104] 3) Calculate the prediction result P according to Equation (6).
[0105] 4) Output the prediction result P.
[0106] Based on Embodiment 2 of the present invention
[0107] An incomplete multi-view multi-label classification system 400 provided by Embodiment 2 of the present invention can execute the incomplete multi-view multi-label classification method provided by Embodiment 1 of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. This device can be implemented in the form of software and / or hardware (integrated circuit), and is generally integrated in a server or a terminal device. Figure 4 It is a schematic structural diagram of an incomplete multi-view multi-label classification system 400 in Embodiment 2 of the present invention. Refer to Figure 4 , the incomplete multi-view multi-label classification system 400 based on deep contrast learning in the embodiments of the present invention specifically may include:
[0108] A network model construction unit 410, configured to construct an incomplete multi-view multi-label classification network model;
[0109] A network model training unit 420, configured to train the incomplete multi-view multi-label classification network model;
[0110] A prediction unit 430, configured to input test data into the trained incomplete multi-view multi-label classification network model for inference and output prediction labels;
[0111] Among them, the incomplete multi-view multi-label classification network model includes three sub-modules: a specific view representation learning framework, an incomplete instance-level contrast learning module, and a weighted fusion and incomplete multi-label classification module. The specific view representation learning framework uses an autoencoder to extract features and reconstruct the original data. The autoencoder includes an encoder and a decoder. The encoder is used to extract features, and the decoder is used to reconstruct the original data. The incomplete instance-level contrast learning module is used to apply an incomplete instance-level contrast loss to the features extracted by the encoder to enhance the consistency of multi-view representations. The weighted fusion and incomplete multi-label classification module is used to perform weighted fusion of multi-views and calculate multi-label classification scores using the weighted fusion results to obtain the inference results of multi-label classification.
[0112] In addition to the above units, the incomplete multi-view multi-label classification system 400 based on deep contrast learning may also include other components. However, since these components are not relevant to the content of the embodiments of the present disclosure, their illustrations and descriptions are omitted here.
[0113] The specific working process of the incomplete multi-view multi-label classification system 400 based on deep contrast learning refers to the description of Embodiment 1 of the incomplete multi-view multi-label classification method based on deep contrast learning above, and will not be elaborated here.
[0114] Based on Embodiment III of the present invention
[0115] The system according to the embodiments of the present invention can also be implemented with the aid of Figure 5 the architecture of the computing device shown. Figure 5 The architecture is shown. As Figure 5 shown, a computer system 501, a system bus 503, one or more CPUs 504, an input / output 502, a memory 505, etc. The memory 505 can store various data or files used for computer processing and / or communication and program instructions executed by the CPU, including the method of Embodiment 1. Figure 5 The architecture shown is only exemplary, and when implementing different devices, adjust according to actual needs Figure 5One or more components therein. The memory 505 serves as a computer-readable storage medium and can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the incomplete multi-view multi-label classification method based on deep contrast learning in the embodiments of the present invention (for example, the network model construction unit 410, the network model training unit 420, and the prediction unit 430 in the incomplete multi-view multi-label classification system 400 based on deep contrast learning). One or more CPUs 304 execute various functional applications and data processing of the system of the present invention by running the software programs, instructions, and modules stored in the memory 505, that is, implement the above-mentioned incomplete multi-view multi-label classification method based on deep contrast learning, and this method includes:
[0116] Construct an incomplete multi-view multi-label classification network model;
[0117] Train the incomplete multi-view multi-label classification network model;
[0118] Input test data into the trained incomplete multi-view multi-label classification network model for inference, and output predicted labels;
[0119] Wherein, the incomplete multi-view multi-label classification network model includes three sub-modules: a specific view representation learning framework, an incomplete instance-level contrast learning module, and a weighted fusion and incomplete multi-label classification module. The specific view representation learning framework uses an autoencoder to extract features and reconstruct the original data. The autoencoder includes an encoder and a decoder. The encoder is used to extract features, and the decoder is used to reconstruct the original data. The incomplete instance-level contrast learning module is used to apply an incomplete instance-level contrast loss to the features extracted by the encoder to enhance the consistency of multi-view representations. The weighted fusion and incomplete multi-label classification module is used to perform weighted fusion of multi-views and calculate multi-label classification scores using the weighted fusion result to obtain the inference result of multi-label classification.
[0120] Certainly, the processor of the server provided by the embodiments of the present invention is not limited to executing the method operations as described above, and can also execute related operations in the incomplete multi-view multi-label classification method based on deep contrast learning provided by any embodiment of the present invention.
[0121] The memory 505 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the terminal and the like. In addition, the memory 505 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 505 may further include a memory remotely provided with respect to one or more CPUs 504, and these remote memories may be connected to the device through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0122] The input / output 502 may be used to receive input digital or character information, and generate key signal inputs related to user settings and function controls of the device. The input / output 502 may also include a display device such as a display screen.
[0123] Based on Embodiment 4 of the present invention
[0124] Embodiments of the present invention may also be implemented as a computer-readable storage medium. A computer program is stored on the computer-readable storage medium according to Embodiment 4. When the computer program is executed by a processor, the method for incomplete multi-view multi-label classification based on deep contrast learning according to Embodiment 1 of the present invention described with reference to the above drawings may be executed.
[0125] Of course, for a storage medium containing computer-executable instructions provided by embodiments of the present invention, the computer-executable instructions are not limited to the method operations as described above, and may also execute related operations in the method for incomplete multi-view multi-label classification based on deep contrast learning provided by any embodiment of the present invention.
[0126] The computer-readable storage medium of the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0127] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0128] The program code contained on the storage medium may be transmitted by any appropriate medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the above.
[0129] The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or terminal. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0130] In summary, as can be seen from the embodiments, a method, system, and storage medium for incomplete multi-view multi-label classification based on deep contrast learning provided by the present invention address the problem of dual incomplete multi-view multi-label classification by proposing a deep contrast network. Different from traditional methods, the present invention focuses on using a deep neural network to extract high-level semantic representations of samples and constructs an end-to-end multi-view feature extraction framework using an autoencoder to learn the representation vectors of samples. In addition, to further improve the representation ability of the model, the present invention introduces unsupervised contrast learning to guide the encoder to extract high-level representation information of multi-views based on the consistency assumption. At the same time, the present invention proposes a weighted fusion method to balance the importance of different views. In summary, the beneficial effects of the present invention are mainly as follows: The incomplete multi-view multi-label classification network model proposed by the present invention has no additional restrictions on the number of views and missing situations of multi-view data, that is, it can process multi-view data sets with a large number of views and any missing situations, and can also adapt to the situation where multi-label supervision information randomly disappears; The incomplete instance-level contrast loss proposed by the present invention can effectively aggregate cross-view features, enabling instances of the same sample in different views to satisfy the multi-view consistency assumption, thereby enhancing the high-level feature representation ability and improving the classification accuracy; The incomplete multi-view multi-label classification network model proposed by the present invention has good application characteristics and can be deployed in a production environment after training, and can immediately give inference results for the input incomplete multi-view test data.
[0131] Note that the above is only the preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. An incomplete multi-view multi-label classification method based on deep contrast learning, characterized in that, The method includes the following steps: Construct an incomplete multi-view multi-label classification network model; Train the incomplete multi-view multi-label classification network model; Input test data into the trained incomplete multi-view multi-label classification network model for inference and output predicted labels; Among them, the incomplete multi-view multi-label classification network model includes three sub-modules: a specific view representation learning framework, an incomplete instance-level contrast learning module, and a weighted fusion and incomplete multi-label classification module. The specific view representation learning framework uses an autoencoder to extract features and reconstruct the original data. The autoencoder includes an encoder and a decoder. The encoder is used to extract features, and the decoder is used to reconstruct the original data. The incomplete instance-level contrast learning module is used to impose an incomplete instance-level contrast loss on the features extracted by the encoder to enhance the consistency of multi-view representations; The weighted fusion and incomplete multi-label classification module is used to perform weighted fusion of multi-views and calculate multi-label classification scores using the weighted fusion results to obtain the inference results of multi-label classification; The incomplete multi-view data is selected from text, audio, and video; The incomplete instance-level contrast learning module uses a contrast learning method to guide the encoder to extract consistent features, specifically including: For l views, there are l×n instances, and any of them There are l×n-1 instances and instances Constitute an instance pair, and instance Instance pairs belonging to the same sample is a positive instance pair, and the rest are negative instance pairs i and j both represent the number of samples, v and u both represent the number of views, and the incomplete instance-level contrastive learning module is used to shorten the distance between positive instance pairs and expand the distance between negative instance pairs. The distance metric function Using cosine similarity, the expression is as follows: where <·> represents the dot product operation, and for any two views, the incomplete contrastive learning loss function is as follows: where \(l\) represents the number of views, \(n\) represents the number of samples, \(\tau\) represents the diffusion degree of the control distribution, represents the prior missing view indicator matrix, indicating that the \(v\)-th view of the \(i\)-th sample is missing, indicating that the \(v\)-th view of the \(i\)-th sample is available, combined with the incomplete contrast loss function of all views is:
2. The incomplete multi-view multi-label classification method based on deep contrast learning according to claim 1, wherein In the specific view representation learning framework, for the view input data X (v) , the encoder E (v) extracts the corresponding feature Z (v) = E (v) (X (v) ), and the decoder D (v) decodes Z (v) to obtain the reconstructed feature Apply a squared loss function at the output of the decoder for making the reconstructed have a small error with the original X (v) : small error Among them, l represents the number of views, n represents the number of samples, and m v represents the feature dimension of view v, represents the reconstructed feature of sample i, represents the original view input data of sample i, represents the prior missing view indicator matrix, indicates that the v-th view of the i-th sample is missing, indicates that the v-th view of the i-th sample is available.
3. The incomplete multi-view multi-label classification method based on deep contrast learning according to claim 1, characterized in that, In the weighted fusion and incomplete multi-label classification module, the fusion representation H of l views is calculated, and for the fusion representation h of each sample i , the specific expression is as follows: where \(l\) represents the number of views, \(v\) represents the view number, denotes the feature representation obtained after the instance of the \(v\)-th view of sample \(i\) passes through the feature encoder \(E\) (v) ; denotes the prior missing view indicator matrix, indicating that the \(v\)-th view of the \(i\)-th sample is missing, indicating that the \(v\)-th view of the \(i\)-th sample is available, \(i\) represents the sample number, and \(n\) represents the number of samples.
4. The incomplete multi-view multi-label classification method based on deep contrastive learning according to claim 3, wherein The weighted fusion and incomplete multi-label classification module calculates multi-label classification scores using the weighted fusion results, specifically including: performing a linear activation operation on the fused representation H, and the activation function is the Sigmoid function: Among them, ω represents the parameters of the fully connected layer , P represents the predicted class scores, and the multi-label classification loss function used is the weighted BCE loss, and the expression is as follows: Among them, is the introduced label missing indication matrix, represents the label missing indication category, n represents the number of samples, c is the number of label categories, and Y i,j represents the predicted label category of the j-th view of the i-th sample, and P i,j represents the predicted category score of the j-th view of the i-th sample.
5. An incomplete multi-view multi-label classification system based on deep contrast learning, characterized in that, including: A network model construction unit for constructing an incomplete multi-view multi-label classification network model; A network model training unit for training the incomplete multi-view multi-label classification network model; A prediction unit for inputting test data into the trained incomplete multi-view multi-label classification network model for inference and outputting predicted labels; Among them, the incomplete multi-view multi-label classification network model includes three sub-modules: a specific view representation learning framework, an incomplete instance-level contrast learning module, and a weighted fusion and incomplete multi-label classification module. The specific view representation learning framework uses an autoencoder to extract features and reconstruct the original data. The autoencoder includes an encoder and a decoder. The encoder is used to extract features, and the decoder is used to reconstruct the original data. The incomplete instance-level contrast learning module is used to impose an incomplete instance-level contrast loss on the features extracted by the encoder to enhance the consistency of multi-view representations. The weighted fusion and incomplete multi-label classification module is used to perform weighted fusion of multi-views and calculate multi-label classification scores using the weighted fusion results to obtain the inference results of multi-label classification; The incomplete multi-view data is selected from text, audio, and video; The incomplete instance-level contrast learning module uses a contrast learning method to guide the encoder to extract consistent features, specifically including: For l views, there are l×n instances, and any of them There are l×n-1 instances and instances Constitute an instance pair, and instance Instance pairs belonging to the same sample is a positive instance pair, and the rest are negative instance pairs i and j both represent the number of samples, v and u both represent the number of views, and the incomplete instance-level contrastive learning module is used to shorten the distance between positive instance pairs and expand the distance between negative instance pairs. The distance metric function Using cosine similarity, the expression is as follows: where <·> represents the dot product operation. For any two views, the incomplete contrastive learning loss function is as follows: where \(l\) represents the number of views, \(n\) represents the number of samples, \(\tau\) represents the diffusion degree of the control distribution, denotes the prior missing view indicator matrix, indicating that the \(v\)-th view of the \(i\)-th sample is missing, indicating that the \(v\)-th view of the \(i\)-th sample is available, combined with the incomplete contrast loss function of all views is:
6. An electronic device, characterized in that, including: A processor; and a memory, wherein a computer-executable program is stored in the memory, and when the computer-executable program is executed by the processor, the incomplete multi-view multi-label classification method based on deep contrast learning according to any one of claims 1-4 is executed.
7. A storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, the incomplete multi-view multi-label classification method based on deep contrast learning according to any one of claims 1-4 is implemented.
Citation Information
Patent Citations
Incomplete multi-view clustering method and system based on local structure and balance perception
CN115311483A
Overlapping trace norms for multi-view learning
US20160026925A1