Training method and device of conversion model for incomplete multi-view data
By training a transformation model to extract and transform features from multi-view data, the problem of poor clustering accuracy caused by missing view data is solved, and efficient clustering analysis under missing view data is achieved.
Patent Information
- Application Number
- CN202111583624.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-22
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2041-12-22
AI Technical Summary
In practical engineering applications, since it is impossible to collect complete multi-view data of all things, cluster analysis based on data with missing views results in poor accuracy of cluster analysis.
By training a transformation model, the first and second transformation models are used to extract and transform features from multi-view data, generate data for missing views, and update model parameters to improve the accuracy of cluster analysis.
This technology enables the automatic generation of missing view data, thereby improving the accuracy and effectiveness of cluster analysis.
Smart Images

Figure CN114254709B_ABST
Abstract
Description
Technical Field
[0001] This application relates to computer technology, and more particularly to a training method and apparatus for a conversion model for incomplete multi-view data. Background Technology
[0002] With the continuous development of data acquisition technology, the data people obtain often has multiple perspectives, forming multi-view data.
[0003] Cluster analysis is the process of dividing a set into multiple clusters based on the relationships between data objects, grouping closely spaced data objects into the same cluster and distributing widely spaced data objects into different clusters. In practical engineering applications, it is generally impossible to collect complete multi-view data for all things.
[0004] Because multi-view data can comprehensively and accurately describe data objects, clustering analysis based on data from missing views will result in poor accuracy of the clustering analysis. Summary of the Invention
[0005] This application provides a training method and apparatus for a transformation model for incomplete multi-view data, and an information clustering method and apparatus for incomplete multi-view data, to overcome the problem of poor accuracy in cluster analysis.
[0006] In a first aspect, embodiments of this application provide a training method for a transformation model oriented towards incomplete multi-view data, comprising:
[0007] Obtain first and second description information of the sample object, wherein the first description information is of a first information type and the second description information is of a second information type;
[0008] The first conversion model is used to perform feature extraction and feature conversion processing on the first descriptive information to obtain first feature information and first conversion information. The first conversion model is used to convert the descriptive information of the first information type into the descriptive information of the second information type.
[0009] The second transformation model is used to perform feature extraction and feature transformation on the second description information to obtain second feature information and second transformation information. The second transformation model is used to convert the description information of the second information type into the description information of the first information type.
[0010] The model parameters of the first conversion model and the second conversion model are updated based on the first feature information, the first conversion information, the second feature information, the second conversion information, the first description information, and the second description information.
[0011] In one possible design, updating the model parameters of the first transformation model and the second transformation model based on the first feature information, the first transformation information, the second feature information, the second transformation information, the first description information, and the second description information includes:
[0012] The first feature information is obtained by restoring the first feature information using a first restoration model, and the second feature information is obtained by restoring the second feature information using a second restoration model.
[0013] The model parameters of the first conversion model and the second conversion model are updated based on the first feature information, the first conversion information, the second feature information, the second conversion information, the first description information, the second description information, the first restoration information, and the second restoration information.
[0014] In one possible design, the first conversion model includes a first feature extraction model and a first feature conversion model; the second conversion model includes a second feature extraction model and a second feature conversion model.
[0015] Based on the first feature information, the first transformation information, the second feature information, the second transformation information, the first description information, the second description information, the first restoration information, and the second restoration information, update the model parameters of the first transformation model and the second transformation model, including:
[0016] Based on the first description information, the first restoration information, the second description information, and the second restoration information, update the model parameters of the first feature extraction model, the first restoration model, the second feature extraction model, and the second restoration model;
[0017] Update the model parameters of the first feature extraction model and the second feature extraction model based on the first feature information and the second feature information;
[0018] The model parameters of the first feature transformation model and the second feature transformation model are updated based on the first feature information, the first transformation information, the second feature information, and the second transformation information.
[0019] In one possible design, updating the model parameters of the first feature extraction model, the first restoration model, the second feature extraction model, and the second restoration model based on the first description information, the first restoration information, the second description information, and the second restoration information includes:
[0020] Based on the first description information, the first restoration information, the second description information, and the second restoration information, determine the first loss;
[0021] Based on the first loss, update the model parameters of the first feature extraction model, the first restoration model, the second feature extraction model, and the second restoration model.
[0022] In one possible design, determining the first loss based on the first description information, the first restoration information, the second description information, and the second restoration information includes:
[0023] The first description information, the first restoration information, the second description information, and the second restoration information are processed according to the first preset loss function to determine the first loss.
[0024] In one possible design, updating the model parameters of the first feature extraction model and the second feature extraction model based on the first feature information and the second feature information includes:
[0025] Obtain the mutual information between the first feature information and the second feature information;
[0026] Based on the first feature information, the second feature information, and the mutual information, a second loss between the first feature information and the second feature information is determined;
[0027] Based on the second loss, update the model parameters of the first feature extraction model and the second feature extraction model.
[0028] In one possible design, the model parameters of the first feature transformation model and the second feature transformation model are updated based on the first feature information, the first transformation information, the second feature information, and the second transformation information, including:
[0029] Based on the first conversion information, the second feature information, the second conversion information, and the first feature information, a third loss is determined;
[0030] Based on the third loss, update the model parameters of the first feature transformation model and the second feature transformation model.
[0031] In one possible design, the method further includes:
[0032] When the sum of the first loss, the second loss, and the third loss is minimized, the first transformation model and the second transformation model are determined to be converged.
[0033] Secondly, embodiments of this application provide an information clustering method for incomplete multi-view data, including:
[0034] Obtain first description information of the first object, wherein the first description information is a first information type;
[0035] The first feature information of the first description information is obtained through the conversion model, and the second feature information corresponding to the first feature information is obtained through the conversion model, wherein the second feature information is the feature information corresponding to the description information of the second information type;
[0036] The first object is clustered based on the first feature information and the second feature information.
[0037] In one possible design, the transformation model includes a feature extraction model and a feature transformation model; obtaining first feature information of the first descriptive information through the transformation model, and obtaining second feature information corresponding to the first feature information through the transformation model, includes:
[0038] The first feature information is obtained by performing feature extraction processing on the first descriptive information using the feature extraction model.
[0039] The first feature information is transformed using the feature transformation model to obtain the second feature information.
[0040] Thirdly, embodiments of this application provide a training apparatus for a transformation model oriented towards incomplete multi-view data, comprising:
[0041] The acquisition module is used to acquire first description information and second description information of a sample object, wherein the first description information is a first information type and the second description information is a second information type;
[0042] The processing module is used to perform feature extraction and feature transformation processing on the first description information through a first transformation model to obtain first feature information and first transformation information. The first transformation model is used to convert the description information of the first information type into the description information of the second information type.
[0043] The processing module is further configured to perform feature extraction and feature transformation processing on the second description information through the second transformation model to obtain second feature information and second transformation information. The second transformation model is configured to convert the description information of the second information type into the description information of the first information type.
[0044] The update module is used to update the model parameters of the first conversion model and the second conversion model based on the first feature information, the first conversion information, the second feature information, the second conversion information, the first description information, and the second description information.
[0045] In one possible design, the update module is specifically used for:
[0046] The first feature information is obtained by restoring the first feature information using a first restoration model, and the second feature information is obtained by restoring the second feature information using a second restoration model.
[0047] The model parameters of the first conversion model and the second conversion model are updated based on the first feature information, the first conversion information, the second feature information, the second conversion information, the first description information, the second description information, the first restoration information, and the second restoration information.
[0048] In one possible design, the first conversion model includes a first feature extraction model and a first feature conversion model; the second conversion model includes a second feature extraction model and a second feature conversion model.
[0049] The update module is specifically used for:
[0050] Based on the first description information, the first restoration information, the second description information, and the second restoration information, update the model parameters of the first feature extraction model, the first restoration model, the second feature extraction model, and the second restoration model;
[0051] Update the model parameters of the first feature extraction model and the second feature extraction model based on the first feature information and the second feature information;
[0052] The model parameters of the first feature transformation model and the second feature transformation model are updated based on the first feature information, the first transformation information, the second feature information, and the second transformation information.
[0053] In one possible design, the update module is specifically used for:
[0054] Based on the first description information, the first restoration information, the second description information, and the second restoration information, determine the first loss;
[0055] Based on the first loss, update the model parameters of the first feature extraction model, the first restoration model, the second feature extraction model, and the second restoration model.
[0056] In one possible design, the update module is specifically used for:
[0057] The first description information, the first restoration information, the second description information, and the second restoration information are processed according to the first preset loss function to determine the first loss.
[0058] In one possible design, the update module is specifically used for:
[0059] Obtain the mutual information between the first feature information and the second feature information;
[0060] Based on the first feature information, the second feature information, and the mutual information, a second loss between the first feature information and the second feature information is determined;
[0061] Based on the second loss, update the model parameters of the first feature extraction model and the second feature extraction model.
[0062] In one possible design, the update module is specifically used for:
[0063] Based on the first conversion information, the second feature information, the second conversion information, and the first feature information, a third loss is determined;
[0064] Based on the third loss, update the model parameters of the first feature transformation model and the second feature transformation model.
[0065] In one possible design, the processing module is further configured to:
[0066] When the sum of the first loss, the second loss, and the third loss is minimized, the first transformation model and the second transformation model are determined to be converged.
[0067] Fourthly, embodiments of this application provide an information clustering apparatus for incomplete multi-view data, comprising:
[0068] The first acquisition module is used to acquire first description information of the first object, wherein the first description information is a first information type.
[0069] The second acquisition module is used to acquire first feature information of the first description information through a conversion model, and to acquire second feature information corresponding to the first feature information through the conversion model, wherein the second feature information is feature information corresponding to description information of the second information type;
[0070] The clustering module is used to perform clustering processing on the first object based on the first feature information and the second feature information.
[0071] In one possible design, the transformation model includes a feature extraction model and a feature transformation model; the second acquisition module is specifically used for:
[0072] The first feature information is obtained by performing feature extraction processing on the first descriptive information using the feature extraction model.
[0073] The first feature information is transformed using the feature transformation model to obtain the second feature information.
[0074] Fifthly, embodiments of this application provide a training device for a transformation model oriented towards incomplete multi-view data, comprising:
[0075] Memory, used to store programs;
[0076] A processor for executing the program stored in the memory, wherein, when the program is executed, the processor is configured to perform the method described in the first aspect above and any of the various possible designs of the first aspect.
[0077] Sixthly, embodiments of this application provide an information clustering device for incomplete multi-view data, comprising:
[0078] Memory, used to store programs;
[0079] A processor for executing the program stored in the memory, wherein, when the program is executed, the processor is configured to perform the method described in the second aspect above and any of the various possible designs of the second aspect.
[0080] In a seventh aspect, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the methods described in the first aspect and various possible designs of the first aspect, or in the second aspect and any of the various possible designs of the second aspect.
[0081] Eighthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect above and various possible designs of the first aspect, or any of the methods described in the second aspect above and various possible designs of the second aspect.
[0082] This application provides a training method and apparatus for a conversion model for incomplete multi-view data. The method includes: acquiring first and second descriptive information of a sample object, wherein the first descriptive information is of a first information type and the second descriptive information is of a second information type. The first descriptive information is processed by feature extraction and feature transformation using a first conversion model to obtain first feature information and first conversion information. The first conversion model is used to convert the descriptive information of the first information type into descriptive information of the second information type. The second descriptive information is processed by feature extraction and feature transformation using a second conversion model to obtain second feature information and second conversion information. The second conversion model is used to convert the descriptive information of the second information type into descriptive information of the first information type. The model parameters of the first and second conversion models are updated based on the first feature information, first conversion information, second feature information, second conversion information, first descriptive information, and second descriptive information. By acquiring first and second descriptive information from multiple perspectives of the sample object, then converting the information types of the first and second descriptive information according to the conversion model, and updating the model parameters of the conversion model based on the converted data and corresponding descriptive information, a conversion model for converting information types can be trained, enabling the automatic generation of data for missing views based on data from existing views.
[0083] Furthermore, this application provides an information clustering method and apparatus for incomplete multi-view data. The method extracts features through a transformation model to obtain feature information of the existing view data. Then, it processes the feature information of the existing view data through the transformation model to obtain feature information of the missing view, thereby completing the feature information of the missing view. Clustering can then be performed based on the first feature information and the second feature information, or the description information of the missing view can be restored based on the completed feature information of the missing view. Clustering can then be performed based on the first description information and the second description information. Therefore, in this embodiment, when data of some views is missing, the missing view data can be completed based on the data of the existing view. After completion, clustering is performed based on the information of multiple views, thereby effectively improving the accuracy of clustering. Attached Figure Description
[0084] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0085] Figure 1This is a schematic diagram illustrating the implementation of multi-view data provided in an embodiment of this application;
[0086] Figure 2 A flowchart illustrating the training method for the conversion model provided in this application embodiment;
[0087] Figure 3 The flowchart of the training method for the conversion model provided in the embodiments of this application Figure 2 ;
[0088] Figure 4 This is a schematic diagram of the structure of the training conversion model provided in the embodiments of this application;
[0089] Figure 5 This is a schematic diagram illustrating the implementation of mutual information provided in an embodiment of this application;
[0090] Figure 6 This is a schematic diagram illustrating the optimization objective of the conversion model provided in the embodiments of this application;
[0091] Figure 7 A flowchart illustrating the information clustering method provided in this application embodiment;
[0092] Figure 8 Schematic diagram of the conversion model provided in the embodiments of this application Figure 1 ;
[0093] Figure 9 Schematic diagram of the conversion model provided in the embodiments of this application Figure 2 ;
[0094] Figure 10 A schematic diagram illustrating the implementation of the restored view provided in an embodiment of this application;
[0095] Figure 11 A schematic diagram illustrating the clustering effect provided in an embodiment of this application;
[0096] Figure 12 This is a schematic diagram of the structure of the training device for the conversion model provided in the embodiments of this application;
[0097] Figure 13 This is a schematic diagram of the information clustering device provided in the embodiments of this application;
[0098] Figure 14 A schematic diagram of the hardware structure of the training device for the conversion model provided in the embodiments of this application;
[0099] Figure 15 This is a schematic diagram of the hardware structure of the information clustering device provided in the embodiments of this application. Detailed Implementation
[0100] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0101] To better understand the technical solution of this application, the relevant concepts involved in this application will be introduced below.
[0102] Autoencoders: An autoencoder is an artificial neural network that learns an efficient representation of input data through unsupervised learning. This efficient representation of the input data is called coding, and its dimensionality is generally much smaller than the input data, making autoencoders useful for dimensionality reduction. More importantly, autoencoders can serve as powerful feature detectors for pre-training deep neural networks. Furthermore, autoencoders can randomly generate data similar to the training data, which is called a generative model. For example, an autoencoder can be trained using images of faces and can generate new images.
[0103] Contrastive learning: Contrastive learning methods learn feature representations of samples by comparing them separately with positive and negative samples in the feature space. For any data x, the goal of contrastive learning is to learn a representation mechanism f such that:
[0104] score(f(x),f(x + ))>>score(f(x),f(x - ))
[0105] Where, x + x and x are similar positive samples, x - x is a negative sample that is dissimilar to x. score is a metric function to measure the similarity between samples. f() is the representation of the processing of the corresponding sample by the representation mechanism f.
[0106] If we use the vector dot product to calculate similarity, the contrastive learning loss function can be expressed as:
[0107]
[0108] Where N represents the number of samples x, and exp is an exponential function with the natural constant e as its base in advanced mathematics. This form can obviously be calculated using the cross-entropy loss function, which aims to make samples more similar to positive samples and less similar to negative samples. This formula can also be explained using information theory, namely, increasing the mutual information between positive samples; hence, this loss is also called mutual information loss.
[0109] MNIST Dataset: The MNIST dataset is a classic dataset in the field of machine learning. It consists of 60,000 training samples and 10,000 test samples. Each sample is a 28*28 pixel grayscale handwritten digit image.
[0110] Multi-source heterogeneous data: During the process of enterprise informatization construction, due to the phased, technical, and other economic and human factors involved in the construction and implementation of data management systems for various business systems, enterprises accumulate a large amount of business data using different storage methods, including vastly different data management systems, ranging from simple file databases to complex network databases. These constitute the enterprise's heterogeneous data sources. The multi-source heterogeneous data in this application can also be referred to as multi-view data.
[0111] Based on the concepts introduced above, the relevant technical background involved in this application will be further described in detail below.
[0112] With the continuous development of data acquisition technology, the data people obtain often has multiple perspectives, forming multi-view data. Furthermore, in some practical problems, the same thing can be described from multiple different ways or angles, and these multiple descriptions constitute multiple views of the thing.
[0113] In fact, we often encounter multi-view data in our daily work and daily life. For example, for the product "Apple," its multi-view data can be referenced... Figure 1 To understand, Figure 1 This is a schematic diagram illustrating the implementation of multi-view data provided in an embodiment of this application.
[0114] Reference Figure 1 Regarding Apple, it may simultaneously possess Figure 1 Image data of an apple shown in 101, and Figure 1 The textual description data about apples shown in 102 are collectively referred to as the two views of the apple.
[0115] For example, for big data on web pages, data can be obtained through text or web page links, thus forming multi-view data with two views of the web page data.
[0116] based on Figure 1 As demonstrated by the examples above, multi-view data provides a more multi-dimensional and richer description of things. Data from different views can reflect different characteristics of things. Therefore, data analysis based on multi-view data has better completeness and practical application capabilities. Since most multi-view data is unlabeled, how to complete better data analysis tasks in the absence of labels is also an important goal.
[0117] However, in practical engineering applications, it's usually impossible to collect complete and comprehensive datasets for all objects. For example, for apples, there might only be image data without text data, or for pears, there might only be text data without images. Understandably, analyzing data objects based on incomplete views will lead to poor data processing results.
[0118] The following explains cluster analysis for data object processing. In today's information explosion era, the amount of data is constantly increasing. How to extract useful information from this vast amount of data has become a key focus. Data mining technology, as an important tool for big data processing and information extraction, has been widely applied. Specifically, cluster analysis is the process of dividing a set into multiple clusters based on the relationships between data objects, grouping objects that are close together into the same cluster and objects that are far apart into different clusters.
[0119] Therefore, data can be segmented based on similarity to obtain more accurate clustering results. From a machine learning perspective, clustering analysis is essentially an unsupervised learning method that can perform clustering operations on data with unknown labels to extract useful information.
[0120] Furthermore, with the increasing demands for data informatization, describing data from a single view is no longer sufficient to achieve the desired results. Therefore, clustering of multi-view data is a key research focus. Clustering data consisting of a single view is called single-view clustering, while multi-view clustering uses clustering methods to process multi-view data.
[0121] Based on the above introduction, it can be determined that in practical engineering applications, it is usually impossible to collect a complete and comprehensive dataset of all things. In summary, based on the above introduction to cluster analysis, performing cluster analysis on data objects based on data with missing views will result in poor accuracy and effectiveness of the cluster analysis.
[0122] To address the problems in the existing technology, this application proposes the following technical concept: by training a transformation model, the transformation model can automatically infer and complete some information of the actual view based on the data of the existing view of the current object, and then perform cluster analysis based on the data after the view is completed, thereby effectively improving the accuracy and clustering effect of the cluster analysis.
[0123] Based on the above description, the training method of the conversion model and the information clustering method provided in this application will be introduced below with reference to specific embodiments. It is worth noting that the execution subject of each embodiment in this application can be a server, processor, microprocessor or other device with data processing function. This embodiment does not limit the specific implementation of the execution subject, as long as it has data processing function. Its specific implementation can be selected and set according to actual needs.
[0124] It should be noted that the method provided in this application consists of two parts: one part is the training method for training the transformation model, and the other part is the information clustering method for the application of the transformation model. These two parts will be introduced separately below.
[0125] First, combine Figure 2 The training methods for the conversion model are introduced. Figure 2 A flowchart illustrating the training method for the conversion model provided in this application embodiment.
[0126] like Figure 2 As shown, the method includes:
[0127] S201. Obtain the first description information and the second description information of the sample object, wherein the first description information is of the first information type and the second description information is of the second information type.
[0128] In this embodiment, the sample object is the data object corresponding to the multi-view description, such as the apple, webpage, etc. mentioned in the above example. This embodiment does not limit the specific implementation of the sample object; it can be any object, as long as it can be described using multi-view data.
[0129] In this embodiment, the first description information and the second description information are multi-view data of the sample object. The first description information is a first information type, and the second description information is a second information type. That is to say, the first description information and the second description information are used to describe the sample object from different perspectives.
[0130] In one possible example, the sample object in this embodiment may be, for example, the evaluation described above, and the first descriptive information may be, for example, the image data described above, and the second descriptive information may be, for example, the text descriptive data described above.
[0131] This embodiment does not limit the specific implementation of the sample object and the first and second description information, as long as the sample object is an object that can be described by data, and the first and second description information are multi-view data of the sample object.
[0132] It should also be noted that, since the current embodiment is for model training, the current sample object includes complete multi-view data, which means that the first description information and the second description information of the sample object can be obtained.
[0133] S202. The first description information is processed by feature extraction and feature transformation through the first transformation model to obtain the first feature information and the first transformation information. The first transformation model is used to transform the description information of the first information type into the description information of the second information type.
[0134] After determining the first and second descriptive information of the sample information, this embodiment has a first conversion model, which is used to convert the descriptive information of the first information type into the descriptive information of the second type, such as converting the image type descriptive information into the text type descriptive information.
[0135] Since the first descriptive information in this embodiment is descriptive information of a first information type, it can be processed by feature extraction and feature transformation using a first transformation model. Specifically, the first transformation model in this embodiment includes two parts: one part is used to extract features from the data to obtain the first feature information of the first descriptive information, and the other part is used to perform feature transformation after feature extraction to obtain the first transformed information of the first descriptive information. It can be understood that the first transformed information here actually corresponds to the transformed descriptive information of the second information type.
[0136] S203. The second description information is processed by feature extraction and feature transformation through the second transformation model to obtain the second feature information and the second transformation information. The second transformation model is used to transform the description information of the second information type into the description information of the first information type.
[0137] Furthermore, after determining the first and second description information of the sample information, this embodiment also includes a second conversion model, which is used to convert the description information of the second information type into the description information of the first type, such as converting the text-type description information into the image-type description information.
[0138] Since the second descriptive information in this embodiment is descriptive information of a second information type, feature extraction and feature transformation processing can be performed on the second descriptive information using the second transformation model. Specifically, the second transformation model in this embodiment includes two parts: one part is used to extract features from the data to obtain the second feature information of the second descriptive information, and the other part is used to perform feature transformation after feature extraction to obtain the second transformed information of the second descriptive information. It can be understood that the second transformed information here actually corresponds to the transformed descriptive information of the first information type.
[0139] S204. Update the model parameters of the first transformation model and the second transformation model based on the first feature information, the first transformation information, the second feature information, the second transformation information, the first description information, and the second description information.
[0140] After determining the first feature information, the first transformation information, the second feature information, and the second transformation information, the transformed description information can actually be determined. This is because in this embodiment, the original first and second description information of the sample object can be determined. Therefore, the parameters of the first transformation model and the second transformation model can be updated based on the determined content.
[0141] Understandably, after updating the parameters of the transformation model, the above steps can be repeated, and the parameters of the transformation model can be updated repeatedly, until the transformation model reaches the final optimization goal, thus obtaining the transformation model we need.
[0142] The training method for the conversion model provided in this application includes: acquiring first and second descriptive information of a sample object, wherein the first descriptive information is of a first information type and the second descriptive information is of a second information type. The first descriptive information is processed by feature extraction and feature conversion using a first conversion model to obtain first feature information and first conversion information. The first conversion model is used to convert descriptive information of the first information type into descriptive information of the second information type. The second descriptive information is processed by feature extraction and feature conversion using a second conversion model to obtain second feature information and second conversion information. The second conversion model is used to convert descriptive information of the second information type into descriptive information of the first information type. The model parameters of the first and second conversion models are updated based on the first feature information, first conversion information, second feature information, second conversion information, first descriptive information, and second descriptive information. By acquiring first and second descriptive information from multiple perspectives of the sample object, then converting the information types of the first and second descriptive information according to the conversion model, and updating the model parameters of the conversion model based on the converted data and corresponding descriptive information, a conversion model for converting information types can be trained, enabling the automatic generation of data for missing views based on data from existing views.
[0143] Based on the above embodiments, the following is combined with Figures 3 to 6 The training method of the conversion model provided in this application will be further described in detail. Figure 3 The flowchart of the training method for the conversion model provided in the embodiments of this application Figure 2 , Figure 4 This is a schematic diagram of the structure of the training conversion model provided in the embodiments of this application. Figure 5 This is a schematic diagram illustrating the implementation of mutual information provided in an embodiment of this application. Figure 6 This is a schematic diagram illustrating the optimization objective of the conversion model provided in the embodiments of this application.
[0144] like Figure 3 As shown, the method includes:
[0145] S301. Obtain the first description information and the second description information of the sample object, wherein the first description information is of the first information type and the second description information is of the second information type.
[0146] The implementation of S301 is similar to that of S201, and the specific implementation will not be described in detail here.
[0147] To facilitate the description of the implementation in this embodiment, Figure 4 The diagram below illustrates the processing procedure in this embodiment, which can be referred to as... Figure 4 A further understanding of the first and second description information in this embodiment is required.
[0148] Reference Figure 4 In one possible implementation, Figure 4 X in 1 This can be understood as the first descriptive information in this embodiment. Figure 4 X in 2 This can be understood as the second descriptive information in this embodiment, that is... Figure 4 X in 1 and X 2 It refers to multi-perspective data on sample objects.
[0149] S302. The first description information is processed by feature extraction and feature transformation through the first transformation model to obtain the first feature information and the first transformation information. The first transformation model is used to transform the description information of the first information type into the description information of the second information type.
[0150] The implementation of S302 is similar to that of S202, and the specific implementation will not be described in detail here.
[0151] Furthermore, the first conversion model in this embodiment includes a first feature extraction model and a first feature conversion model. The first feature extraction model is used to extract features from the first description information to obtain first feature information of the first description information, and the first feature conversion model is used to perform feature conversion processing on the first feature information after the first description information is extracted, thereby obtaining first converted information.
[0152] For example, you can refer to Figure 4 To understand, Figure 4 f in (1) This is the first feature extraction model in this embodiment, such as... Figure 4 As shown, the first feature extraction model f (1) The first descriptive information X can be used. 1 Feature extraction is performed to obtain the first descriptive information X. 1 First feature information Z 1 .
[0153] S303. The second description information is processed by feature extraction and feature transformation through the second transformation model to obtain the second feature information and the second transformation information. The second transformation model is used to transform the description information of the second information type into the description information of the first information type.
[0154] The implementation of S303 is similar to that of S203, and the specific implementation will not be described in detail here.
[0155] Furthermore, the second conversion model in this embodiment includes a second feature extraction model and a second feature conversion model. The second feature extraction model is used to extract features from the second descriptive information to obtain second feature information of the second descriptive information, and the second feature conversion model is used to perform feature conversion processing on the second feature information after the extraction of the second descriptive information to obtain second converted information.
[0156] For example, you can refer to Figure 4 To understand, Figure 4 f in (2) This is the second feature extraction model in this embodiment, such as... Figure 4 As shown, the second feature extraction model f (2) The second descriptive information X can be used. 2 Feature extraction is performed to obtain the second descriptive information X. 2 First feature information Z 2 .
[0157] S304. First restored information is obtained by restoring the first feature information through the first restoration model, and second restored information is obtained by restoring the second feature information through the second restoration model.
[0158] Furthermore, this embodiment also includes a first restoration model, which is used to restore the first feature information to obtain the first restored information. For example, refer to... Figure 4 , Figure 4 g in (1) This is the first restoration model in this embodiment, such as... Figure 4 As shown, the first reduction model g (1) The first feature information Z can be used 1 The restoration process is performed to obtain the first restored information.
[0159] Similarly, this embodiment also includes a second restoration model, which is used to restore the second feature information to obtain the second restored information. For example, refer to... Figure 4 , Figure 4 g in (2) This is the second restoration model in this embodiment, such as... Figure 4 As shown, the second reduction model g (2) The second feature information Z can be used 2 The restoration process is performed to obtain the second restored information.
[0160] S305. Based on the first preset loss function, process the first description information, the first restored information, the second description information, and the second restored information to determine the first loss.
[0161] It is understandable that in this embodiment, the first feature extraction model f is used. (1) For the first descriptive information X 1 Feature extraction is performed to obtain the first feature information Z. 1 And subsequently through the first reduction model g (1) For the first feature information Z 1 Perform restoration processing to obtain the first restored information. In practice, the first feature extraction model and the first reconstruction model learn the features of the first descriptive information in the middle, so as to reconstruct the first descriptive information as much as possible. Therefore, in this embodiment, the first descriptive information and the first reconstruction information need to be as close as possible to ensure the accuracy and effectiveness of the features extracted by the first feature extraction model.
[0162] Therefore, in this embodiment, the first loss can be determined based on the first description information, the first restoration information, the second description information, and the second restoration information, and then the parameters of the model can be updated based on the first loss.
[0163] In one possible implementation, for example, the first descriptive information, the first restored information, the second descriptive information, and the second restored information can be processed according to a first preset loss function to determine the first loss. The first preset loss function can, for example, satisfy the following formula:
[0164]
[0165] Where V represents the number of views, and m represents the number of sample objects. This represents the v-th descriptive information of the t-th sample (for example, when v is 1, it is the first descriptive information described above; when v is 2, it is the second descriptive information described above). f represents the feature extraction model of the vth generation. (v) Processing the v-th descriptive information actually means processing the v-th feature information of the v-th descriptive information, and... Represents the v-th reduction model g (v) The process of restoring the v-th feature information actually refers to the restored v-th feature information, which corresponds to... Figure 4 In For mean square error, L 重建 This is the first loss, which actually corresponds to the reconstruction loss in the above embodiments.
[0166] Based on the above introduction, it can be understood that the first loss in this embodiment actually describes the difference between the restored information and the initial description information. In actual implementation, the first preset loss function can satisfy Formula 1 described above, or it can be used as the first preset loss function by performing identity transformations, adding relevant parameters, coefficients, etc. on the basis of Formula 1. This embodiment does not limit the specific implementation of the first preset loss function, as long as the first preset loss function can determine the first loss based on the first description information, the first restored information, the second description information, and the second restored information, where the first loss is used to describe the difference between the restored information and the description information.
[0167] Therefore, the purpose of the reconstruction process corresponding to the first loss in this embodiment is to use X 1 and To reconstruct the loss to train the model f (1) and g (1) , so that the generated And the original X 1 The closer the approximation, the better. The model f is... (1) It can be understood as an encoder, used to convert X 1 Mapped to Z 1 And g (1) It can be understood as a decoder, used to convert Z... 1 Mapped to X 1 .
[0168] And, for model f (2) and g (2) Similarly, using X 2 and To reconstruct the loss to train the model f (2) and g (2) , so that the generated And the original X 2 The closer the approximation, the better. The model f is... (2) It can be understood as an encoder, used to convert X 2 Mapped to Z 2 And g (2) It can be understood as a decoder, used to convert Z... 2 Mapped to X 1 .
[0169] And in one possible implementation, for X 1 and X 2 The reconstruction loss can be determined separately, and each can be used to reconstruct and generate its corresponding X. 1 and X 2 Its implementation method is similar to that described above, and will not be repeated here.
[0170] S306. Based on the first loss, update the model parameters of the first feature extraction model, the first restoration model, the second feature extraction model, and the second restoration model.
[0171] It is understandable that the loss function actually represents the optimization objective of the model. The loss value determined by the loss function indicates to the model which direction is better, and the model will then evolve in that direction. The smaller the loss value, the better the model. Therefore, the goal of model optimization is to reduce the loss value as much as possible.
[0172] Therefore, after determining the first loss, the parameters of the corresponding model can be adjusted. The first loss is obtained by processing the first descriptive information, the first restored information, the second descriptive information, and the second restored information. The first loss describes the difference between the restored information and the descriptive information. The optimization objective of the current model is to make the restored information and the descriptive information as close as possible.
[0173] The first and second descriptive information are fixed, while the first restored information is obtained from the first feature extraction model and the first restored model, and the second restored information is obtained from the second feature extraction model and the second restored model. Therefore, to narrow the gap between the restored information and the descriptive information, the model parameters of the first feature extraction model, the first restored model, the second feature extraction model, and the second restored model can be updated. In other words, the parameters of the first feature extraction model, the first restored model, the second feature extraction model, and the second restored model can be updated. Figure 4 The first feature extraction model f in (1) First reduction model g (1) Second feature extraction model f (2) Second reduction model g (2) Optimization can be performed, and the specific parameter update methods can be selected and set according to actual needs, for example, referring to relevant implementations in machine learning.
[0174] S307. Obtain the mutual information between the first feature information and the second feature information.
[0175] Furthermore, in this embodiment, it is necessary to complete the missing data. Therefore, it is necessary to infer some information of the missing view based on the existing view. For example, in the example described above, it is necessary to generate the image information of the apple based on the text information of the apple, or to generate the text information of the apple based on the image information of the apple. In other words, it is necessary to realize the mutual generation of data under different views.
[0176] For example, you can refer to Figure 5 Understanding mutual information. Assume there exists a first descriptive information X. 1 And the second descriptive information X 2 , where X 1For example, it can be a view Figure 1 Information, and X 2 For example, it can be a view Figure 2 The information is as follows. Then refer to... Figure 5 , Figure 5 The solid line in the image represents the first descriptive information X in the multi-view data. 1 The included feature information Z 1 Similarly, Figure 5 The dashed lines in the text represent the second descriptive information X in the multi-view data. 2 The included feature information Z 2 .
[0177] In information theory, Figure 5 I(Z) 1 Z 2 ) represents the first descriptive information X 1 Second description information X 2 The mutual information represents the first feature information Z of the first descriptive information. 1 The second feature information Z of the second description information 2 The degree of information overlap between them. And, Figure 5 H(Z) 1 |Z 2 ) is given Z 2 In the case of Z 1 Information entropy, that is, in Z 2 Under the premise that it occurs, Z 1 The new information brought about by the occurrence of events. And, Figure 5 H(Z) 2 |Z 1 ) is given Z 1 In the case of Z 2 Information entropy, that is, in Z 1 Under the premise that it occurs, Z 2 The amount of new information brought about by the occurrence of events. Among them, entropy is a measure of the uncertainty of a random variable.
[0178] In the process of mutual deduction of information, it is usually necessary to have greater mutual information and reference. Figure 5 In order to enable mutual information I(Z) 1 Z 2 To make the conditional entropy H(Z) larger, we actually need to reduce the conditional entropy H(Z). 1 |Z 2 ) and H(Z 2 |Z 1 ), that is, Z 1 By Z 2 The decision part, and Z 2 By Z 1The decisive factor, namely minimizing the conditional entropy H, will encourage the discarding of inconsistent information across views, thereby further improving information consistency. (See reference...) Figure 5 It can be determined that in H(Z) 1 |Z 2 ) and H(Z 2 |Z 1 When the mutual information I(Z) approaches 0, the mutual information... 1 Z 2 ) is the largest.
[0179] Furthermore, because X 1 and X 2 The information contained in the data may contain redundancy and invalid noise. Therefore, in this embodiment, the optimization by extracting effective information to form feature information Z can effectively improve the final data derivation effect.
[0180] In the above Figure 5 Based on the information presented, it can be determined that the model optimization in this embodiment also has another objective: to minimize conditional entropy and maximize mutual information. Based on this, referring to... Figure 5 In this embodiment, it can be based on the first feature information Z 1 Second feature information Z 2 We conducted comparative learning and set a loss function for comparative learning in order to improve the mutual information between the feature information of the same labeled data from different perspectives.
[0181] In one possible implementation, for example, the mutual information between the first feature information and the second feature information can be obtained, that is, the I(Z) described above can be determined. 1 Z 2 The specific implementation of mutual information determination can be found in the relevant technical descriptions, and will not be elaborated here.
[0182] S308. Based on the first feature information, the second feature information, and the mutual information, determine the second loss between the first feature information and the second feature information.
[0183] After determining the mutual information, a second loss can be determined, for example, based on the first feature information, the second feature information, and the mutual information. In one possible implementation, the first feature information, the second feature information, and the mutual information can be processed according to a second preset loss function to determine the second loss.
[0184] The second preset loss function can, for example, satisfy the following formula:
[0185]
[0186] Where m is the number of sample objects. The mutual information of the first and second feature information of the t-th sample. Let the information entropy be the first feature information of the t-th sample. Let L be the information entropy of the second feature information of the t-th sample, and α be the balancing parameter. The value of α can be selected and set according to actual needs. 对比 This is the second loss.
[0187] Meanwhile, in multi-view data, assuming that in the view corresponding to the first descriptive information, the first feature information Z of the data sample... 1 It can be represented as a representation vector Z, which includes multiple elements z, where each element z is a probability output. Each element z in the representation vector Z represents the probability of a certain unknown category. In fact, this representation vector is the marginal probability output of the first descriptive information of the current data sample.
[0188] Similarly, in the view corresponding to the second descriptive information, the second feature information Z of the data sample 2 It can be represented as a representation vector Z`, which includes multiple elements z`, where each element z` is a probability output. Each element z` in the representation vector Z` represents the probability of a certain unknown category. In fact, this representation vector is the marginal probability output of the second descriptive information of the current data sample.
[0189] Based on this, the current objective of contrastive learning is to learn Z and Z' with the maximum mutual information, thereby learning the consistency between views. Therefore, in this embodiment, based on the second preset loss function described in Formula 2 above, and after derivation, the second preset loss function can also be expressed as Formula 3 as follows:
[0190]
[0191] Where D represents the dimension of the representation vector Z and the representation vector Z', P dd′ P is the joint probability value (normalized to 1) of the element in the d-th row of vector Z and the element in the d-th row of vector Z'. d P represents the value of the d-th row element of the vector Z (which will be normalized to 1). d` α is the value of the element in the d-th row of the representation vector Z` (which will be normalized to 1), and α is the balance parameter.
[0192] It is understandable that Formula 3 above is actually derived from Formula 2 above. Therefore, the second loss in this embodiment can be obtained based on both Formula 2 and Formula 3 above. The second loss here is actually... Figure 4The comparative loss is introduced in the text.
[0193] S309. Update the model parameters of the first feature extraction model and the second feature extraction model according to the second loss.
[0194] Similar to the above, the loss function can indicate the optimization objective of the model, so after determining the second loss, the parameters of the same model can be adjusted.
[0195] The second loss is obtained by processing the first feature information, the second feature information, and the mutual information. The second loss corresponding to the second preset loss function in this embodiment is used to narrow down the common parts between feature information. In fact, the first feature information and the second feature information output by the model need to expand the common parts as much as possible in order to learn the consistency between different views and then complete the subsequent data derivation.
[0196] The first feature information is obtained by the first feature extraction module, and the second feature information is obtained by the second feature extraction module. Therefore, to increase the commonalities between the first and second feature information, both the first and second feature extraction models can be updated. In other words, the... Figure 4 The first feature extraction model f in (1) Second feature extraction model f (2) Optimization can be performed, and the specific parameter update methods can be selected and set according to actual needs, for example, referring to relevant implementations in machine learning.
[0197] It is understandable that in a complete view, different view information of the same data sample must have consistent links, meaning they must have significant mutual information. Therefore, the feature information of each view should also have significant mutual information. Based on this, this embodiment introduces the idea of contrastive learning. By setting the aforementioned second preset loss function, the first and second feature extraction models are optimized based on the calculated second loss. This effectively expands the mutual information between the first feature information extracted by the first feature extraction model and the second feature information extracted by the second feature extraction model. Consequently, the feature information mentioned by the first and second feature extraction models includes as much feature information as possible from other perspectives, enhancing the effectiveness and accuracy of subsequent data inference from missing perspectives.
[0198] Therefore, in this embodiment, for the comparison process, the input is a view. Figure 1 Heshi Figure 2 The data, the target model for optimization is f (1) and f (2) The purpose of this sub-process is to use Z 1and Z 2 By comparing the loss training model f (1) and f (2) This makes Z 1 and Z 2 Consistency between them can be maximized.
[0199] S310. Determine the third loss based on the first transformation information, the second feature information, the second transformation information, and the first feature information.
[0200] Furthermore, in this embodiment, the first transformation model further includes a first feature transformation model, and the second transformation model further includes a second feature transformation model.
[0201] For example, you can refer to Figure 4 To understand, Figure 4 G in (1) This is the first feature transformation model in this embodiment, such as... Figure 4 As shown, the first feature transformation model G (1) The first feature information Z can be used 1 Feature transformation processing is performed to obtain the first transformation information. In this embodiment, the first transformation information is actually the predicted second feature information.
[0202] as well as, Figure 4 G in (2) This is the second feature transformation model in this embodiment, such as... Figure 4 As shown, the second feature transformation model G (2) The second feature information Z can be used 2 Feature transformation processing is performed to obtain the second transformation information. In this embodiment, the second transformation information is actually the predicted second feature information.
[0203] Reference Figure 4 It is understood that in this embodiment, the first feature transformation model G can be used. (1) For the first feature information Z 1 Feature transformation is performed to obtain the predicted second feature information. And can be achieved through the second feature transformation model G (2) For the second feature information Z 2 Feature transformation is performed to obtain the first predicted feature information. Understandably, in the pairwise prediction process described above, the goal of model optimization is to ensure that the predicted feature information is consistent with the original feature information, so as to guarantee the accuracy of the predicted feature information from the missing perspective.
[0204] Therefore, in this embodiment, the third loss can be determined based on the first transformation information, the second feature information, the second transformation information, and the first feature information. In one possible implementation, for example, the first transformation information, the second feature information, the second transformation information, and the first feature information can be processed according to a second preset loss function to determine the third loss.
[0205] The second prediction loss function can, for example, satisfy the following formula:
[0206]
[0207] Among them, Z 2 It is the second feature information, G (1) (Z 1 ) represents the first feature transformation model G (1) For the first feature information Z 1 The processing essentially refers to the predicted second feature information described above. That is, the first transformation information, and Z. 1 It is the first feature information, G (2) (Z 2 ) represents the second feature transformation model G (2) For the second feature information Z 2 The processing essentially refers to the predicted first feature information described above. That is, the second conversion information, L 预测 This is the third loss. The third loss is... Figure 4 The predicted loss in the middle.
[0208] Based on the above introduction, it can be determined that the third loss in this embodiment actually describes the difference between the predicted feature information and the true feature information. In actual implementation, the third preset loss function can satisfy Formula 4 as described above. Identical transformations, additions of relevant parameters and coefficients, etc., based on Formula 4 can also be used as the third preset loss function in this embodiment. This embodiment does not limit the specific implementation of the third preset loss function, as long as the third preset loss function can determine the third loss based on the first transformation information, the second feature information, the second transformation information, and the first feature information, where the third loss is used to describe the difference between the predicted feature information and the true feature information.
[0209] Therefore, in this embodiment, regarding the prediction process, model G... (1) and model G (2) It can be similar to model g in the reconstruction process described above, but the parameters are completely different. In this embodiment, the input Z... 1 It can be done through model G (1)After calculation The purpose of this subprocess is to use Z 1 , and Z 2 Model G is trained by predicting the loss. (1) , making and Z 2 The closer the better. This corresponds to formula four above. Similarly, for Z 2 Predicting Z 1 The situation is similar, which corresponds to the second term in Formula 4 above.
[0210] The goal of this prediction process is to help fill in missing views. After all three models mentioned above have been trained, as long as we have some view data of a sample, such as X... 2 You can use f (2) Its feature Z is obtained 2 Then through G (2) Get the feature Z of another view 1 Then use Z 1 You can use g (1) Get missing view data X 1 In conclusion, the three loss terms can be trained simultaneously, only with different training objectives.
[0211] S311. Update the model parameters of the first feature transformation model and the second feature transformation model according to the third loss.
[0212] Similarly, since the loss function essentially represents the optimization objective of the model, the parameters of the corresponding model can be adjusted after determining the third loss. The third loss is obtained by processing the first transformation information, the second feature information, the second transformation information, and the first feature information. The third loss describes the difference between the predicted feature information and the true feature information. The current optimization objective of the model is to make the predicted feature information as close as possible to the true feature information.
[0213] When measuring the gap, the second and first feature information are fixed, while the first transformation information is obtained from the first feature transformation model, and the second transformation information is obtained from the second feature transformation model. Therefore, to narrow the gap between the predicted and true feature information, the model parameters of both the first and second feature transformation models can be updated. In other words, the parameters of the predicted and true feature transformation models can be updated. Figure 4 The first feature transformation model G in (1) and the second feature transformation model G (2)Optimization can be performed, and the specific parameter update methods can be selected and set according to actual needs, for example, referring to relevant implementations in machine learning.
[0214] It is understandable that the above-described implementation of determining the loss value based on the loss function and updating the corresponding model parameters based on the loss value can be iterated multiple times. The ultimate goal of iterative optimization of each model in this embodiment can be, for example, to minimize the total loss, where the total loss can satisfy, for example, the following formula five:
[0215] L 总 =L 重建 +L 对比 +L 预测 Formula 5
[0216] Among them, L 总 For the total loss, L 重建 As the first loss, L 对比 For the second loss, L 预测 This is the third loss.
[0217] In other words, in this embodiment, the convergence of the first and second transformation models can be determined when the sum of the first, second, and third losses is minimized. This allows for the determination of the trained first and second transformation models.
[0218] In one possible implementation, during the iterative training of the model, to ensure stable training, the model parameters of the first feature extraction model, the first reconstruction model, the second feature extraction model, and the second reconstruction model described above can be updated based on the first loss and the second loss, so as to obtain a model that can output relatively stable feature information.
[0219] In other words, during model training, it can be done in stages. First, the first feature extraction model, the first restoration model, the second feature extraction model, and the second restoration model can be initially optimized. When the model can output relatively stable feature information, the first feature extraction model and the second feature extraction model can be further optimized. This achieves the optimization of the first feature extraction model, the first restoration model, the second feature extraction model, the second restoration model, the first feature extraction model, and the second feature extraction model. By setting up staged optimization, the efficiency and correctness of model training can be effectively guaranteed.
[0220] Alternatively, in an optional implementation, all six models described above can be optimized together; this embodiment does not impose any restrictions on this approach.
[0221] Based on the model training process described above, the following section combines... Figure 6The objectives of model training in this embodiment will be further explained.
[0222] Reference Figure 6 By optimizing models f and g for the aforementioned reconstruction process, it becomes possible to learn the view's features Z through information reconstruction. Specifically, refer to... Figure 6 For models f and g, assuming the input is the view Figure 1 Data X 1 The final output is the generated (also a view) Figure 1 (Data of type f), the intermediate processing model is f (1) and g (1) The intermediate output is the feature data Z. 1 Regarding the optimization objectives described above, the current input data X 1 and output It should be as close as possible. Also, assuming the input is a view... Figure 2 Data X 2 The final output is the generated (also a view) Figure 2 (Data of type f), the intermediate processing model is f (1) and g (1) The intermediate output is the feature data Z. 2 Regarding the optimization objectives described above, the current input data X 2 and output The closer the better.
[0223] Furthermore, in this embodiment, the first feature information Z can be used. 1 Second feature information Z 2 The comparative learning achieves consistent learning across views, aiming to increase the mutual information I(Z) between the first and second feature information. 1 Z 2 To maximize mutual information I(Z) 1 Z 2 ).
[0224] Furthermore, in this embodiment, the first feature information Z can be used. 1 Predicting the second feature information, and using the second feature information Z 2 Predict the first feature information, and use pairwise prediction to repair the missing view, referencing... Figure 6 Understanding can be achieved through model G. (1) For feature information Z 1 Feature transformation is performed to obtain the predicted feature information. Among the predicted feature information and actual feature information Z 2The more similar, the better. And this can be achieved through model G. (2) For feature information Z 2 Feature transformation is performed to obtain the predicted feature information. Among the predicted feature information and actual feature information Z 1 The more similar, the better.
[0225] The above is specifically to make the conditional entropy H(Z) 2 |Z 1 ) and H(Z 1 |Z 2 Minimizing Z can effectively reduce the amount of Z in the processed feature information. 1 By Z 2 The decision part, and Z 2 By Z 1 The decision is made by maximizing the mutual information between different perspectives, thereby effectively increasing the consistency of data from different perspectives included in the extracted features. This, in turn, can effectively improve the accuracy and effectiveness of the data derived from different perspectives in subsequent data derivation.
[0226] In summary, the training method for the conversion model provided in this application effectively increases the consistency of data from different perspectives in the feature information extracted by the first and second feature extraction models by incorporating contrastive learning and pairwise prediction during model training and adjusting the corresponding model parameters through appropriate loss functions. Furthermore, the extracted feature information reduces dependence on data from other perspectives, thereby ensuring the effectiveness and accuracy of subsequent inferences about missing perspective data. By setting a prediction model and optimizing it based on a third loss, the accuracy and effectiveness of the prediction model in predicting data from other perspectives can be effectively guaranteed. The conversion model trained based on the above process can effectively and automatically complete missing perspective data based on existing perspective data.
[0227] In the optional implementation methods, the above-described methods are all for data from two perspectives. In fact, the implementation methods for data from more than two perspectives are similar. As long as similar processing is performed between each pair, this embodiment will not elaborate further.
[0228] Understandably, after training the aforementioned transformation model, it's possible to use it to complete data from missing perspectives and perform corresponding clustering. Therefore, based on the above embodiments, the following will combine... Figures 7 to 11 This application introduces the information clustering method provided.
[0229] Figure 7 A flowchart of the information clustering method provided in the embodiments of this application. Figure 8 Schematic diagram of the conversion model provided in the embodiments of this application Figure 1 , Figure 9 Schematic diagram of the conversion model provided in the embodiments of this application Figure 2 , Figure 10 This is a schematic diagram illustrating the implementation of the restored view provided in an embodiment of this application. Figure 11 This is a schematic diagram illustrating the clustering effect provided in the embodiments of this application.
[0230] like Figure 7 As shown, the method includes:
[0231] S701. Obtain the first description information of the first object, wherein the first description information is of the first information type.
[0232] In this embodiment, the first object is the object corresponding to the current data. Similar to the sample object mentioned above, the first object in this embodiment can be, for example, an apple, a webpage, etc., as in the example above. This embodiment does not limit the specific implementation of the first object, which can be any object that can be described by data.
[0233] In this embodiment, the first object has first description information, which is a first information type, similar to that described above. For example, the first description information may be image data.
[0234] It should be noted that the current embodiment describes a situation where the first object has missing perspective data. Therefore, the first object in this embodiment only has first descriptive information, while the second descriptive information is of the second information type. For example, in one example, the current first object is an apple, the first descriptive information is image data, and the second descriptive information is text descriptive data. In this case, the current example could be that only image data of the apple is available, while text descriptive data is missing.
[0235] In actual implementation, the specific implementation of the first object and the data of the perspective that the first object lacks can be selected according to actual needs. That is to say, in another possible implementation, the first object may only have the second descriptive information and lack the first descriptive information, and its implementation is similar.
[0236] S702. Obtain first feature information of first descriptive information through a conversion model, and obtain second feature information corresponding to the first feature information through a conversion model, wherein the second feature information is feature information corresponding to descriptive information of the second information type.
[0237] Based on the above description, it can be determined that the transformation model in this embodiment can complete the data from missing perspectives. Specifically, the transformation model can perform feature extraction and feature transformation processing on the descriptive information. Therefore, the first descriptive information can be obtained through the transformation model, and feature extraction processing can be performed to obtain the first feature information of the first descriptive information.
[0238] Furthermore, the conversion model can also perform feature conversion processing on the first feature information to obtain the second feature information corresponding to the first feature information. In this embodiment, the second feature information is actually the feature information corresponding to the description information of the second information type.
[0239] Specifically, when the first data contains first descriptive information but second descriptive information is not available, the second descriptive information needs to be completed. In this embodiment, for example, it can be processed by a first conversion model. Based on the above description, it can be determined that the first conversion model may include a first feature extraction model and a first feature conversion model.
[0240] For example, you can refer to Figure 8 To understand, such as Figure 8 As shown, assume that X 1 The first descriptive information, and f (1) As the first feature extraction model, the first transformation model in the current example includes... Figure 8 The first feature extraction model f shown (1) and the first feature transformation model G (1) .
[0241] Then it can be extracted using the first feature extraction model f (1) For the first descriptive information X 1 Feature extraction is performed to obtain Figure 8 The first feature information Z shown 1 This step involves feature extraction from existing descriptive information.
[0242] Because the second descriptive information from the other perspective is currently missing, it is necessary to complete the information from the missing perspective, referring to... Figure 8 It can be achieved through the first feature transformation model G (1) For the first feature information Z 1 The transformation process is performed to obtain the second feature information Z. 2 It is understandable that the second feature information Z here... 2 In reality, it refers to the predicted feature information, and therefore can be understood as described above.
[0243] It should be noted that the first feature transformation model described in the above embodiments produces first transformation information. However, the first transformation information is actually the second feature information predicted by the first feature transformation model. Therefore, in this embodiment, the first feature transformation model can transform the first feature information to obtain the second feature information corresponding to the first feature information, thereby effectively completing the information of the missing perspective.
[0244] It is understandable that after determining the second feature information, the data features of the missing perspective have actually been recovered. In this embodiment, clustering can be performed directly based on the data features. Therefore, subsequent clustering can be performed directly based on the first and second feature information.
[0245] Alternatively, in another possible implementation, clustering based on the data could be performed, as described in [reference needed]. Figure 8 In determining the second feature information Z 2 Subsequently, for example, it can be achieved through a second reduction model g. (2) For the second feature information Z 2 Processing is performed to obtain the second restored information. The second restoration information here In essence, it is the second descriptive information X obtained through the restoration process. 2 Therefore, subsequent clustering processing can be performed, for example, based on the first descriptive information and the second descriptive information obtained from the restoration process.
[0246] The above describes the implementation where the first data contains first descriptive information but lacks second descriptive information, and the second descriptive information is restored according to the first transformation model. In another possible implementation, the first data may contain second descriptive information but lack first descriptive information; similarly, the first descriptive information can be restored according to the second transformation model.
[0247] For example, you can refer to Figure 9 To understand, such as Figure 9 As shown, assume that X 2 For the second descriptive information, and f (2) As the second feature extraction model, the second transformation model in the current example includes... Figure 9 The second feature extraction model f shown (2) Second feature transformation model G (2) .
[0248] Then the second feature extraction model f can be used. (2) For the second descriptive information X 2 Feature extraction is performed to obtain Figure 9 The second feature information Z shown 2This step involves feature extraction from existing descriptive information.
[0249] Because the initial descriptive information from the other perspective is currently missing, it is necessary to complete the information from the missing perspective, referring to... Figure 9 It can be achieved through the second feature transformation model G (2) For the second feature information Z 2 The transformation process is performed to obtain the first feature information Z. 1 It is understandable that the first feature information Z here is... 1 In reality, it refers to the predicted feature information, and therefore can be understood as described above.
[0250] It should be noted that the second feature transformation model described in the above embodiments produces second transformation information. However, the second transformation information is actually the first feature information predicted by the second feature transformation model. Therefore, in this embodiment, the second feature transformation model transforms the second feature information to obtain the first feature information corresponding to the second feature information, thereby effectively completing the information from the missing perspective.
[0251] It is understandable that after determining the first feature information, the data features of the missing perspective have actually been recovered. In this embodiment, clustering can be performed directly based on the data features. Therefore, subsequent clustering can be performed directly based on the first and second feature information.
[0252] Alternatively, in another possible implementation, clustering based on the data could be performed, as described in [reference needed]. Figure 9 In determining the first feature information Z 1 Subsequently, for example, it can be achieved through the first reduction model g. (1) For the first feature information Z 1 Processing is performed to obtain the first restored information. The first restored information here In essence, it is the first descriptive information X obtained through the restoration process. 1 Therefore, subsequent clustering processing can be performed based on the second descriptive information and the first descriptive information obtained from the restoration process.
[0253] For example, it can be combined Figure 10 To understand the data recovery results, refer to... Figure 10 For example, the missing view data recovery operation described in this embodiment can be performed using the NoiseMnist dataset (i.e., Gaussian noise added to the MNIST dataset). The recovery result is referenced. Figure 10The first and fourth rows represent the complete view, the second and fifth rows represent the missing view, and the third and sixth rows represent the results recovered from the complete view. Figure 10 As can be seen, the transformation model in this embodiment can accurately and effectively restore and complete the missing view.
[0254] S703. Perform clustering processing on the first object based on the first feature information and the second feature information.
[0255] After completing the data for missing perspectives, the first object can be clustered based on the first and second feature information. It is understood that there can be many objects participating in the clustering process. In this embodiment, the first object is just one object participating in the clustering process. For each object participating in the clustering process, the view can be completed and restored in the manner described above, and then the clustering process can be performed. This can effectively achieve data clustering based on multiple views, thereby effectively improving the accuracy of the clustering process.
[0256] Alternatively, based on the above description, the first object can be clustered according to the first descriptive information and the restored second descriptive information. Alternatively, if the first descriptive information is missing, the first object can be clustered according to the second descriptive information and the restored first descriptive information. This embodiment does not limit whether the clustering process involves features or data, as long as it can achieve clustering based on information from multiple perspectives. The specific implementation method can be selected according to actual needs.
[0257] Furthermore, in the actual implementation process, the specific implementation method of clustering can be selected and set according to actual needs. For example, K-Means clustering can be used, or mean-shift clustering, density-based clustering methods, etc. This embodiment does not limit the specific implementation of clustering. In short, clustering is an unsupervised data mining method that can help users discover the inherent class relationships in data. Generally speaking, data of the same class are close together, while data of different classes are far apart.
[0258] The specific implementation method of clustering can be selected and set according to actual needs. As long as clustering can be performed based on information from multiple views in this embodiment to obtain the results of multi-view clustering, it is acceptable.
[0259] And it can also be combined Figure 11 To understand, such as Figure 11 As shown, clustering can currently be performed on multiple datasets, and the visualization of clustering can be represented by grouping the data into different clusters. (And reference...) Figure 11As the number of iterations of the transformation model increases, the clustering effect also improves. It is assumed that normalized mutual information (NMI) can be used as a clustering metric to measure the clustering effect, where a higher NMI value indicates better clustering performance.
[0260] Then refer to Figure 11 Among them, (a) shows the clustering effect with 20 iterations and an NMI of 0.567; (b) shows the clustering effect with 50 iterations and an NMI of 0.651; (c) shows the clustering effect with 100 iterations and an NMI of 0.707; and (d) shows the clustering effect with 200 iterations and an NMI of 0.759. Therefore, based on... Figure 11 It is certain that during the training of the transformation model, the clustering effect improves with the increase of the number of iterations. Therefore, the transformation model in this embodiment can effectively complete the data from missing perspectives to achieve multi-perspective clustering processing.
[0261] It should be noted that, due to the multi-view data X 1 and X 2 Generally, data with different dimensions, such as image data and text data, differ in size and dimensionality. Therefore, directly clustering these raw data is not feasible. The technical solution of this application transforms these two different types of data into feature data Z with the same data dimensionality. 1 and Z 2 Afterwards, due to Z 1 and Z 2 The dimensions are the same, and the feature information of the source view is preserved, so the feature data Z is directly applied. 1 and Z 2 Clustering is effective and meaningful, for example, referring to Figure 11 As shown in the clustering results, with the increase of the number of iterations, the clusters of the same category become more cohesive and the NMI value becomes higher and higher. Therefore, the technical solution of this application can also effectively solve the problem of multi-view clustering.
[0262] The information clustering method provided in this application extracts features from existing viewpoint data using a transformation model. Then, the transformation model processes these features to obtain feature information from missing viewpoints, thus completing the missing viewpoint feature information. Clustering can then be performed based on the first and second feature information, or the missing viewpoint description information can be reconstructed based on the completed missing viewpoint feature information, followed by clustering based on the first and second description information. Therefore, this embodiment can complete the missing viewpoint data based on existing viewpoint data even when some viewpoint data is missing. After completion, clustering is performed based on multi-viewpoint information, effectively improving the accuracy of clustering. The improvement in clustering accuracy described in this application can also be understood as improving the overall effect of clustering.
[0263] Figure 12 This is a schematic diagram of the structure of the training device for the conversion model provided in an embodiment of this application. Figure 12 As shown, the device 120 includes: an acquisition module 1201, a processing module 1202, and an update module 1203.
[0264] The acquisition module 1201 is used to acquire first description information and second description information of a sample object, wherein the first description information is a first information type and the second description information is a second information type;
[0265] Processing module 1202 is used to perform feature extraction and feature transformation processing on the first description information through a first transformation model to obtain first feature information and first transformation information. The first transformation model is used to convert the description information of the first information type into the description information of the second information type.
[0266] The processing module 1202 is further configured to perform feature extraction and feature transformation processing on the second description information through the second transformation model to obtain second feature information and second transformation information. The second transformation model is configured to convert the description information of the second information type into the description information of the first information type.
[0267] The update module 1203 is used to update the model parameters of the first conversion model and the second conversion model based on the first feature information, the first conversion information, the second feature information, the second conversion information, the first description information, and the second description information.
[0268] In one possible design, the update module 1203 is specifically used for:
[0269] The first feature information is obtained by restoring the first feature information using a first restoration model, and the second feature information is obtained by restoring the second feature information using a second restoration model.
[0270] The model parameters of the first conversion model and the second conversion model are updated based on the first feature information, the first conversion information, the second feature information, the second conversion information, the first description information, the second description information, the first restoration information, and the second restoration information.
[0271] In one possible design, the first conversion model includes a first feature extraction model and a first feature conversion model; the second conversion model includes a second feature extraction model and a second feature conversion model.
[0272] The update module 1203 is specifically used for:
[0273] Based on the first description information, the first restoration information, the second description information, and the second restoration information, update the model parameters of the first feature extraction model, the first restoration model, the second feature extraction model, and the second restoration model;
[0274] Update the model parameters of the first feature extraction model and the second feature extraction model based on the first feature information and the second feature information;
[0275] The model parameters of the first feature transformation model and the second feature transformation model are updated based on the first feature information, the first transformation information, the second feature information, and the second transformation information.
[0276] In one possible design, the update module 1203 is specifically used for:
[0277] Based on the first description information, the first restoration information, the second description information, and the second restoration information, determine the first loss;
[0278] Based on the first loss, update the model parameters of the first feature extraction model, the first restoration model, the second feature extraction model, and the second restoration model.
[0279] In one possible design, the update module 1203 is specifically used for:
[0280] The first description information, the first restoration information, the second description information, and the second restoration information are processed according to the first preset loss function to determine the first loss.
[0281] In one possible design, the update module 1203 is specifically used for:
[0282] Obtain the mutual information between the first feature information and the second feature information;
[0283] Based on the first feature information, the second feature information, and the mutual information, a second loss between the first feature information and the second feature information is determined;
[0284] Based on the second loss, update the model parameters of the first feature extraction model and the second feature extraction model.
[0285] In one possible design, the update module 1203 is specifically used for:
[0286] Based on the first conversion information, the second feature information, the second conversion information, and the first feature information, a third loss is determined;
[0287] Based on the third loss, update the model parameters of the first feature transformation model and the second feature transformation model.
[0288] In one possible design, the processing module 1202 is further configured to:
[0289] When the sum of the first loss, the second loss, and the third loss is minimized, the first transformation model and the second transformation model are determined to be converged.
[0290] The apparatus provided in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effects are similar, and will not be described again here.
[0291] Figure 13 This is a schematic diagram of the information clustering device provided in an embodiment of this application. Figure 13 As shown, the device 130 includes: a first acquisition module 1301, a second acquisition module 1302, and a clustering module 1303.
[0292] The first acquisition module 1301 is used to acquire first description information of the first object, wherein the first description information is a first information type.
[0293] The second acquisition module 1302 is used to acquire first feature information of the first description information through a conversion model, and to acquire second feature information corresponding to the first feature information through the conversion model, wherein the second feature information is feature information corresponding to description information of the second information type;
[0294] Clustering module 1303 is used to perform clustering processing on the first object based on the first feature information and the second feature information.
[0295] In one possible design, the transformation model includes a feature extraction model and a feature transformation model; the second acquisition module 1302 is specifically used for:
[0296] The first feature information is obtained by performing feature extraction processing on the first descriptive information using the feature extraction model.
[0297] The first feature information is transformed using the feature transformation model to obtain the second feature information.
[0298] The apparatus provided in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effects are similar, and will not be described again here.
[0299] Figure 14 A schematic diagram of the hardware structure of the training device for the conversion model provided in the embodiments of this application is shown below. Figure 14 As shown, the training device 140 for the conversion model in this embodiment includes: a processor 1401 and a memory 1402; wherein
[0300] Memory 1402 is used to store computer-executed instructions;
[0301] The processor 1401 is configured to execute computer execution instructions stored in the memory to implement the various steps performed by the training method of the conversion model in the above embodiments. For details, please refer to the relevant descriptions in the foregoing method embodiments.
[0302] Alternatively, the memory 1402 can be either standalone or integrated with the processor 1401.
[0303] When the memory 1402 is set up independently, the training device for the conversion model also includes a bus 1403 for connecting the memory 1402 and the processor 1401.
[0304] In one alternative implementation, the training device for the conversion model in this embodiment can be a graphics processing unit (GPU).
[0305] Figure 15 This is a schematic diagram of the hardware structure of the information clustering device provided in the embodiments of this application, such as... Figure 15 As shown, the information clustering device 150 of this embodiment includes: a processor 1501 and a memory 1502; wherein
[0306] Memory 1502 is used to store instructions executed by the computer;
[0307] Processor 1501 is used to execute computer execution instructions stored in memory to implement the various steps performed by the information clustering method in the above embodiments. For details, please refer to the relevant descriptions in the foregoing method embodiments.
[0308] Alternatively, the memory 1502 can be either standalone or integrated with the processor 1501.
[0309] When the memory 1502 is set up independently, the information clustering device also includes a bus 1503 for connecting the memory 1502 and the processor 1501.
[0310] In one alternative implementation, the information clustering device in this embodiment can be a CPU.
[0311] This application also provides a computer-readable storage medium storing computer-executable instructions. When a processor executes the computer-executable instructions, it implements the training method of the conversion model executed by the training device of the above conversion model, or the information clustering method executed by the information clustering device.
[0312] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.
[0313] The integrated modules implemented as software functional modules described above can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods described in the various embodiments of this application.
[0314] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly manifested as execution by a hardware processor, or execution by a combination of hardware and software modules within the processor.
[0315] The memory may include high-speed RAM, and may also include non-volatile storage (NVM), such as at least one disk storage device, and may also be a USB flash drive, external hard drive, read-only memory, disk or optical disc, etc.
[0316] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0317] The aforementioned storage medium can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium can be any available medium accessible to general-purpose or special-purpose computers.
[0318] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0319] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A method for training a conversion model for incomplete multi-view data, applied to the field of multi-view clustering analysis, characterized in that, The method comprises: obtaining first description information and second description information of a sample object, the first description information being of a first information type, and the second description information being of a second information type; wherein the first description information is picture image data, and the second description information is text description data; or the first description information is text description data, and the second description information is picture image data; performing feature extraction and feature conversion processing on the first description information through a first conversion model to obtain first feature information and first conversion information, the first conversion model being used to convert the description information of the first information type into the description information of the second information type; performing feature extraction and feature conversion processing on the second description information through a second conversion model to obtain second feature information and second conversion information, the second conversion model being used to convert the description information of the second information type into the description information of the first information type; the first conversion model comprises a first feature extraction model and a first feature conversion model; the second conversion model comprises a second feature extraction model and a second feature conversion model; wherein the first feature extraction model is used to perform feature extraction on the first description information to obtain first feature information of the first description information, and the first feature conversion model is used to perform feature conversion processing on the first feature information of the first description information after extraction to obtain first conversion information; the second feature extraction model is used to perform feature extraction on the second description information to obtain second feature information of the second description information, and the second feature conversion model is used to perform feature conversion processing on the second feature information of the second description information after extraction to obtain second conversion information; performing restoration processing on the first feature information through a first restoration model to obtain first restoration information, and performing restoration processing on the second feature information through a second restoration model to obtain second restoration information; updating model parameters of the first feature extraction model, the first restoration model, the second feature extraction model, and the second restoration model according to the first description information, the first restoration information, the second description information, and the second restoration information; updating model parameters of the first feature extraction model and the second feature extraction model according to the first feature information and the second feature information; updating model parameters of the first feature conversion model and the second feature conversion model according to the first feature information, the first conversion information, the second feature information, and the second conversion information.
2. The method of claim 1, wherein, The method comprises: determining a first loss according to the first description information, the first restoration information, the second description information, and the second restoration information; updating model parameters of the first feature extraction model, the first restoration model, the second feature extraction model, and the second restoration model according to the first loss.
3. The method of claim 2, wherein, According to the first description information, the first restoration information, the second description information, and the second restoration information, a first loss is determined, including: According to a first preset loss function, the first description information, the first restoration information, the second description information, and the second restoration information are processed to determine the first loss.
4. The method according to claim 2 or 3, characterized in that, According to the first feature information and the second feature information, model parameters of a first feature extraction model and the second feature extraction model are updated, including: Mutual information of the first feature information and the second feature information is obtained. According to the first feature information, the second feature information, and the mutual information, a second loss between the first feature information and the second feature information is determined. According to the second loss, model parameters of the first feature extraction model and the second feature extraction model are updated.
5. The method of claim 4, wherein, According to the first feature information, the first conversion information, the second feature information, and the second conversion information, model parameters of the first feature conversion model and the second feature conversion model are updated, including: According to the first conversion information, the second feature information, the second conversion information, and the first feature information, a third loss is determined. According to the third loss, model parameters of the first feature conversion model and the second feature conversion model are updated.
6. The method according to any of claims 2-3, 5, characterized in that, The method further includes: When the sum of the first loss, the second loss, and the third loss is minimized, it is determined that the first conversion model and the second conversion model converge.
7. A method of information clustering for incomplete multi-view data, characterized in that, Including: First description information of a first object is obtained, the first description information being of a first information type; First feature information of the first description information is obtained through a conversion model, and second feature information corresponding to the first feature information is obtained through the conversion model, the second feature information being feature information corresponding to description information of a second information type; wherein the conversion model is trained by the method of any one of claims 1-6; The first object is processed for clustering according to the first feature information and the second feature information.
8. The method of claim 7, wherein, The conversion model includes a feature extraction model and a feature conversion model; First feature information of the first description information is obtained through a conversion model, and second feature information corresponding to the first feature information is obtained through the conversion model, including: The first description information is processed for feature extraction through the feature extraction model to obtain the first feature information; The first feature information is processed for conversion through the feature conversion model to obtain the second feature information.
9. A training device of a conversion model for incomplete multi-view data, applied to the field of multi-view clustering analysis, characterized in that, Including: An obtaining module is configured to obtain first description information and second description information of a sample object, the first description information being of a first information type, and the second description information being of a second information type; wherein the first description information is picture image data, and the second description information is text description data; or the first description information is text description data, and the second description information is picture image data; The processing module is configured to perform feature extraction and feature conversion processing on the first description information by using a first conversion model to obtain first feature information and first conversion information, the first conversion model being configured to convert description information of the first information type into description information of the second information type. The processing module is further configured to perform feature extraction and feature conversion processing on the second description information by using a second conversion model to obtain second feature information and second conversion information, the second conversion model being configured to convert description information of the second information type into description information of the first information type. The first conversion model comprises a first feature extraction model and a first feature conversion model, and the second conversion model comprises a second feature extraction model and a second feature conversion model; the first feature extraction model is configured to perform feature extraction on the first description information to obtain first feature information of the first description information, and the first feature conversion model is configured to perform feature conversion processing on the first feature information of the first description information after feature extraction to obtain first conversion information; the second feature extraction model is configured to perform feature extraction on the second description information to obtain second feature information of the second description information, and the second feature conversion model is configured to perform feature conversion processing on the second feature information of the second description information after feature extraction to obtain second conversion information. The updating module is configured to: perform restoration processing on the first feature information by using a first restoration model to obtain first restoration information, and perform restoration processing on the second feature information by using a second restoration model to obtain second restoration information; update model parameters of the first feature extraction model, the first restoration model, the second feature extraction model, and the second restoration model according to the first description information, the first restoration information, the second description information, and the second restoration information; update model parameters of the first feature extraction model and the second feature extraction model according to the first feature information and the second feature information; update model parameters of the first feature conversion model and the second feature conversion model according to the first feature information, the first conversion information, the second feature information, and the second conversion information.
10. The apparatus of claim 9, wherein, The updating module is specifically configured to: determine a first loss according to the first description information, the first restoration information, the second description information, and the second restoration information. update model parameters of the first feature extraction model, the first restoration model, the second feature extraction model, and the second restoration model according to the first loss.
11. The apparatus of claim 10, wherein, The updating module is specifically configured to: determine the first loss by processing the first description information, the first restoration information, the second description information, and the second restoration information according to a first preset loss function.
12. The apparatus of claim 10 or 11, wherein, The updating module is specifically configured to: obtain mutual information of the first feature information and the second feature information; determine a second loss between the first feature information and the second feature information according to the first feature information, the second feature information, and the mutual information. According to the second loss, update model parameters of the first feature extraction model and the second feature extraction model.
13. The apparatus of claim 12, wherein, The updating module is specifically configured to: According to the first conversion information, the second feature information, the second conversion information, and the first feature information, determine a third loss; According to the third loss, update model parameters of the first feature conversion model and the second feature conversion model.
14. The apparatus of any of claims 10-11, 13, wherein, The processing module is further configured to: When the sum of the first loss, the second loss, and the third loss is the minimum, determine that the first conversion model and the second conversion model converge.
15. An information clustering apparatus for incomplete multiview data, characterized by, Comprise: A first obtaining module is configured to obtain first description information of a first object, the first description information being of a first information type; A second obtaining module is configured to obtain first feature information of the first description information through a conversion model, and obtain second feature information corresponding to the first feature information through the conversion model, the second feature information being feature information corresponding to description information of a second information type; wherein the conversion model is trained using the method of any one of claims 1-6; A clustering module is configured to perform clustering processing on the first object according to the first feature information and the second feature information.
16. The apparatus of claim 15, wherein, The conversion model comprises a feature extraction model and a feature conversion model; the second obtaining module is specifically configured to: perform feature extraction processing on the first description information through the feature extraction model to obtain the first feature information; perform conversion processing on the first feature information through the feature conversion model to obtain the second feature information.
17. A training device of a conversion model for incomplete multi-view data, characterized in that, Comprise: A memory is configured to store a program; A processor is configured to execute the program stored in the memory, and when the program is executed, the processor is configured to execute the method of any one of claims 1-6.
18. An information clustering device for incomplete multiview data, characterized by Comprise: A memory is configured to store a program; A processor is configured to execute the program stored in the memory, and when the program is executed, the processor is configured to execute the method of any one of claims 7-8.
19. A computer-readable storage medium, characterized in that, The instructions, when executed on a computer, cause the computer to perform the method of any one of claims 1-6 or 7-8.
20. A computer program product comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the method of any one of claims 1-6 or 7-8. The computer program, when executed by a processor, implements the method of any one of claims 1-6 or 7-8.
Citation Information
Patent Citations
Image-text conversion method and device, intelligent interaction method, device and system, client, server, machine and medium
CN110598739A
Classification model training method and system, computer equipment and storage medium
CN112149705A