A multi-view clustering method and system for imperfect information and a storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-01
- Publication Date
- 2026-08-11
AI Technical Summary
[0005]针对现有技术中的上述不足,本发明提出一种面向不完美信息的多视图聚类方法、系统及存储介质,将与锚点视图正确配对的目标视图建模为潜在变量,并结合实例级配对可靠度与原型级语义传输推断潜在配对视图,从而统一处理不完整视图和噪声关联问题,提升多视图聚类的鲁棒性和准确率
[0038]本发明的有益效果为从不完美信息的统一视角出发,将目标视图配对视图作为潜在变量进行后验推断,并通过实例级配对可靠度估计和原型级语义传输同时缓解不完美多视图信息造成的多视图聚类性能下降问题,提升了多视图聚类在复杂真实数据场景中的鲁棒性与泛化能力。
Smart Images

Figure CN122551003A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multi-view clustering and multimodal data analysis, specifically to a multi-view clustering method, system, and storage medium for imperfect information. Background Technology
[0002] In real-world scenarios, data is often characterized by multiple views or modalities, such as primary color images, depth maps, audio, and text descriptions. Multi-view clustering, as a fundamental tool for heterogeneous data analysis, aims to divide unlabeled instances into clusters with different semantics by mining complementary information within each view and consistency information contained in cross-view paired data, without relying on manual annotation.
[0003] However, in real-world applications, multi-view data often suffers from imperfect information, primarily manifested as incomplete views and noisy associations. Specifically, incomplete views refer to the missing portions of some instances; noisy associations refer to incorrect pairings between different views for some instances. Both of these problems simultaneously undermine the utilization of information within a view and the alignment of information across views, thereby reducing the representation learning ability and clustering accuracy of multi-view clustering models. To address the incomplete view problem, existing methods typically learn a cross-view mapping network based on complete instances and then use observed views to recover the missing views. These methods ignore the potential cross-view noisy associations within complete instances, leading the cross-view mapping network to learn incorrect mapping relationships and ultimately reconstructing inaccurate views. To address the noisy association problem, existing methods typically rely on a sufficient number of complete instances to learn reliable cross-view alignment patterns and reconstruct cross-view pairings accordingly. However, when a large number of incomplete views exist, the insufficient number of complete instances weakens the learning effect of the cross-view alignment patterns.
[0004] In summary, existing multi-view clustering faces the problem of decreased clustering performance in scenarios with imperfect view information caused by unavailable or unreliable cross-view pairing relationships. Summary of the Invention
[0005] To address the aforementioned shortcomings in existing technologies, this invention proposes a multi-view clustering method, system, and storage medium for imperfect information. The method models the target view that is correctly paired with the anchor view as a latent variable, and combines instance-level pairing reliability with prototype-level semantic transmission to infer the latent paired view. This approach uniformly addresses the issues of incomplete views and noisy associations, thereby improving the robustness and accuracy of multi-view clustering.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: This solution provides a multi-view clustering method, system, and storage medium for imperfect information, including the following steps: S1: Determine the availability status of each instance in the multi-view data in each view, construct an imperfect information modeling method, and model the target view that is correctly matched with the anchor view as a latent variable; S2: Calculate cross-view for complete instances Figure 1 Consistent loss is used to estimate instance-level pairing reliability based on the loss distribution, identifying reliable cross-view associations and potential noisy associations. S3: Construct semantic prototypes for each view, calculate semantic attribution vectors from instances to prototypes, construct a reliability-weighted cross-view prototype co-occurrence matrix based on instance-level pairing reliability, and obtain a cross-view prototype transmission scheme through optimal transmission. S4: Combining information from the observable view with prototype-level semantic transmission results, perform posterior inference on potential paired views in the target view; S5: Construct a joint optimization objective based on the inferred potential paired views, perform robust learning on noisy associations and incomplete views, obtain a common representation of multiple views, and output clustering results.
[0007] The beneficial effects of this invention are as follows: it unifies incomplete views and noise associations into the problem of unavailable or unreliable cross-view pairing relationships, and generates reliable cross-view supervision signals under the joint constraints of instance-level reliability estimation and prototype-level semantic transmission through posterior-guided latent pairing view inference; at the same time, it improves the accuracy, normalized mutual information, and adjusted RAND index of multi-view clustering in imperfect information scenarios by using joint training objectives to constrain robust cross-view alignment of complete instances and semantic consistency of incomplete instances.
[0008] Further, step S1 includes the following sub-steps: S11: Given a multi-view dataset, determine the observation state of each instance in different views and construct a common representation space in the multi-view clustering task so that multiple instances can be divided into different semantic clusters; S12: Define imperfect information data as a set of instances with incomplete views or noise associations; S13: For the observed instances in the anchor view, model their expected paired views in the target view as latent variables and establish a posterior distribution model.
[0009] Furthermore, the imperfect information data set in step S12 is defined as follows: in, For multi-view datasets, For a set of instances contaminated by imperfect information, For the first One example, For the number of views, For view indexing, Indicates the first The first instance in Available states in each view This indicates that the view is unavailable. This indicates that the view is available. For anchor point view indexing, Index the target view. This is an indicator function that takes the value 1 when two view observations belong to the same instance, and 0 otherwise. Representation of instances The A view.
[0010] Furthermore, the formula for the posterior distribution model of the latent variables for modeling the expected paired views in the target view in step S13 is as follows: in, Describes the posterior distribution model. This indicates the expected paired view of the anchor view. This indicates the amount of information observed.
[0011] The beneficial effect of the above-mentioned further scheme is that by uniformly representing incomplete views and noise associations as problems of unavailable or unreliable cross-view pairing relationships, a unified modeling basis is provided for subsequent potential paired view inference.
[0012] Furthermore, step S2 includes the following sub-steps: S21: For a complete instance that contains both anchor view and target view observations, extract features using a view-specific encoder and compute cross-view... Figure 1 Set of consistent loss; S22: A two-component Gaussian mixture model is used to fit the loss distribution of the complete instance; S23: Based on the memory effect of deep neural networks prioritizing the fitting of simple patterns, Gaussian components with smaller means are regarded as reliable pairing components, and instance-level pairing reliability is calculated based on the posterior probability that the loss belongs to the component with the lower loss mean.
[0013] Furthermore, in step S21, the cross-view Figure 1 The formula for the set of consistency loss is: in, For measuring two views Figure 1 Consistency loss function For the first Each view has a corresponding view-specific encoder. Representation of instances The One view, Indicates the first The first instance in Available states in each view Indicates the first An instance in the anchor view and target view Both are available in Chinese.
[0014] Furthermore, the formula for fitting the loss distribution of the two-component Gaussian mixture model to the complete instance in step S22 is as follows: in, Represents probability. For measuring two views Figure 1 Consistency loss function Here are the parameters of the Gaussian mixture model, and k is a component of the Gaussian mixture model. For the first The mixing coefficient of each Gaussian component. For the first The probability density corresponding to each Gaussian component.
[0015] Furthermore, the instance-level pairing reliability formula calculated in step S23 is: in, For the first The reliability of pairing two views in a complete instance. Represents probability. It is a low-loss mean Gaussian component. For measuring two views Figure 1 Consistency loss function.
[0016] The beneficial effect of the above-mentioned further scheme is that by estimating instance-level reliability through the cross-view loss distribution of complete instances, the misleading effect of noise association on the construction of cross-view common space can be reduced.
[0017] Furthermore, step S3 includes the following sub-steps: S31: Maintain a set of view-specific semantic prototypes for each view, and assign observation instances to each semantic prototype to obtain the semantic attribution vector of the instance; S32: Introduce instance-level pairing reliability into cross-view prototype co-occurrence estimation to construct a reliability-weighted prototype co-occurrence matrix; S33: Construct the cost matrix for optimal transmission based on the prototype co-occurrence matrix; S34: Solve the optimal transmission problem of the above cost matrix to obtain the cross-view prototype transmission scheme, and transmit the semantic structure in the anchor view to the target view.
[0018] Furthermore, the semantic structure allocation formula from instance to prototype obtained in step S31 is as follows: in, For the first The first instance The view is assigned to the first The probability of a semantic prototype For example No. Features of each view For the first The first view in the A semantic prototype For temperature parameters, For similarity function, The number of semantic clusters or semantic prototypes.
[0019] Furthermore, the formula for constructing the reliability-weighted prototype co-occurrence matrix in step S32 is as follows: in, For view With View The prototype co-occurrence matrix between them For view With View Between The reliability of pairing two views for an instance. and The first An instance in the view and view The prototype semantic attribution vector in the model.
[0020] Furthermore, the formula for the optimal transmission cost matrix constructed in step S33 based on the prototype co-occurrence matrix is as follows: in, For view To view The transmission cost matrix, For view With View The prototype co-occurrence matrix between them It is an infinite norm.
[0021] Furthermore, the formula for transferring the semantic structure from the anchor view to the target view in step S34 is as follows: in, For anchor point view Transfer to target view The semantic attribution vector, For the first An instance in the view The semantic attribution vector, This is a cross-view prototype transfer scheme.
[0022] The beneficial effect of the above-mentioned further scheme is that, through reliability-weighted prototype-level semantic transmission, the semantic structure estimated from all observable samples in each view is propagated across views, avoiding excessive reliance on reliable pairing relationships and sufficient complete samples when directly recovering missing views at the instance level.
[0023] Furthermore, step S4 includes the following sub-steps: S41: Approximate the posterior mean of potential paired views as a weighted combination of instance-level candidate paired views and cross-view semantic transport paired views; S42: Calculate adaptive weights based on instance-level pairing reliability and prototype semantic allocation consistency; S43: For complete instances with reliable pairing relationships, directly use the target view observations as potential pairing views; S44: For complete instances with unreliable pairing relationships, correct the target view observations using the transmitted semantic pairing views; S45: For incomplete instances where the target view is missing, combine the candidate paired views generated by the cross-view generator with the transmitted semantic paired views.
[0024] Furthermore, the approximate formula for the posterior mean of the potential paired views in step S41 is: in, Represents the posterior expectation. This indicates the expected paired view of the anchor view. For adaptive weights, This indicates the amount of information observed. For candidate matching views in the target view, This is a semantic pairing view obtained from prototype-level semantic transfer. For the first An instance in the view Passing semantics to the view Later assigned to the The probability of a semantic prototype For view The first in A semantic prototype.
[0025] Furthermore, the adaptive weight formula obtained in step S42 based on instance-level pairing reliability and prototype semantic allocation consistency is as follows: in, For adaptive weights, For the first The reliability of pairing two views in a complete instance. This represents the vector inner product or a semantic consistency measure. For anchor point view Transfer to target view The semantic attribution vector, This is the semantic attribution vector in the target view.
[0026] Furthermore, step S43 provides the following potential pairing view for a complete instance of a fully reliable pairing relationship: in, Represents the posterior expectation. This indicates the expected paired view of the anchor view. For adaptive weights, This indicates the amount of information observed. For the first One instance in the target view The observed features.
[0027] Furthermore, step S44 provides the following potential pairing view for a complete instance of an unreliable pairing relationship: in, Represents the posterior expectation. This indicates the expected paired view of the anchor view. For adaptive weights, This indicates the amount of information observed. For the first One instance in the target view Observational features in For anchor point view Transfer to target view The semantic attribution vector, For view The semantic prototype in.
[0028] Furthermore, in step S45, the potential paired view for an incomplete instance where the target view is missing is: in, Represents the posterior expectation. This indicates the expected paired view of the anchor view. For adaptive weights, This indicates the amount of information observed. These are candidate paired views for the target view, obtained by passing the anchor view through the cross-view generator. For view To view Cross-view generator, For anchor point view Transfer to target view The semantic attribution vector, For view The semantic prototype in.
[0029] The beneficial effects of the above-mentioned further scheme are as follows: different potential paired view inference strategies are adopted for reliable complete instances, unreliable complete instances and incomplete instances, so that reliable observations are preserved, noisy observations are corrected and missing observations are semantically completed, thereby providing robust cross-view supervision signals.
[0030] Furthermore, step S5 includes the following sub-steps: S51: Construct a joint optimization objective based on the inferred potential pairing views; S52: For a complete instance, construct cross-view contrast constraints using the inferred potential paired views; S53: For incomplete instances, use the transport semantic structure to constrain the semantic consistency of the inferred potential paired views; S54: Based on the joint optimization objective, update the view-specific encoder, cross-view generator, and semantic prototype to obtain a robust multi-view common representation, and perform clustering based on the common representation to output the clustering label of each instance.
[0031] Furthermore, the joint optimization objective formula in step S51 is: in, For the total loss, For the loss term used to handle noise association, For the loss term used to handle incomplete views, To balance hyperparameters.
[0032] Furthermore, the cross-view comparison constraint formula constructed for the complete instance in step S52 is as follows: in, For the loss term used to handle noise association, For measuring two views Figure 1 Consistency loss function For view The corresponding view-specific encoder, This indicates the expected paired view of the anchor view. By replacing potentially noisy original target view observations with potential paired views, the encoder can still learn robust cross-view alignment even under noisy associations.
[0033] Furthermore, the regularization constraint formula for incomplete instances in step S53 is as follows: in, For the loss term used to handle incomplete views, Let KL divergence be the KL divergence. For anchor point view Transfer to target view The semantic attribution vector, This is the prototype semantic attribution vector of the inferred potential paired view.
[0034] Further, step S54 updates the view-specific encoder, cross-view generator, and semantic prototype based on the joint optimization objective to obtain a robust multi-view common representation, and performs clustering based on the common representation to output the clustering label of each instance.
[0035] The beneficial effects of the above further solutions are: through To suppress the negative impact of noise association on the cross-view alignment of complete instances, by... Constraining incomplete instances preserves the semantic structure of the target view, enabling the model to stably learn multi-view representations when incomplete views and noise coexist.
[0036] A computer device includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the method described above.
[0037] A computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the above-described method.
[0038] The beneficial effects of this invention are that, from the unified perspective of imperfect information, the target view paired with the view is used as a latent variable for posterior inference, and the performance degradation of multi-view clustering caused by imperfect multi-view information is alleviated by instance-level pairing reliability estimation and prototype-level semantic transmission, thereby improving the robustness and generalization ability of multi-view clustering in complex real-world data scenarios. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. It should be understood that the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0040] Figure 1 This is a flowchart illustrating the steps of a multi-view clustering method, system, and storage medium for imperfect information provided in an embodiment of the present invention. Detailed Implementation
[0041] The technical solutions of the present invention will now be described with reference to the accompanying drawings in the embodiments of the present invention.
[0042] like Figure 1 As shown, this invention provides a multi-view clustering method, system, and storage medium for imperfect information, the implementation of which is as follows: S1: Determine the availability status of each instance in the multi-view data in each view, construct an imperfect information modeling method, and model the target view that is correctly matched with the anchor view as a latent variable; S2: Calculate cross-view for complete instances Figure 1 Consistent loss is used to estimate instance-level pairing reliability based on the loss distribution, identifying reliable cross-view associations and potential noisy associations. S3: Construct semantic prototypes for each view, calculate semantic attribution vectors from instances to prototypes, construct a reliability-weighted cross-view prototype co-occurrence matrix based on instance-level pairing reliability, and obtain a cross-view prototype transmission scheme through optimal transmission. S4: Combining information from the observable view with prototype-level semantic transmission results, perform posterior inference on potential paired views in the target view; S5: Construct a joint optimization objective based on the inferred potential paired views, perform robust learning on noisy associations and incomplete views, obtain a common representation of multiple views, and output clustering results.
[0043] In this embodiment, step S1 includes the following sub-steps: S11: Given a multi-view dataset, determine the observation state of each instance in different views and construct a common representation space in the multi-view clustering task so that multiple instances can be divided into different semantic clusters; S12: Define imperfect information data as a set of instances with incomplete views or noise associations; S13: For the observed instances in the anchor view, model their expected paired views in the target view as latent variables and establish a posterior distribution model.
[0044] In this embodiment, the imperfect information data set in step S12 is defined as: in, For multi-view datasets, For a set of instances contaminated by imperfect information, For the first One example, For the number of views, For view indexing, Indicates the first The first instance in Available states in each view This indicates that the view is unavailable. This indicates that the view is available. For anchor point view indexing, Index the target view. This is an indicator function that takes the value 1 when two view observations belong to the same instance, and 0 otherwise. Representation of instances The A view.
[0045] In this embodiment, the formula for the posterior distribution model of the latent variables for modeling the expected paired views in the target view in step S13 is as follows: in, Describes the posterior distribution model. This indicates the expected paired view of the anchor view. Indicates the amount of information observed In this embodiment, step S2 includes the following sub-steps: S21: For a complete instance that contains both anchor view and target view observations, extract features using a view-specific encoder and compute cross-view... Figure 1 Set of consistent loss; S22: A two-component Gaussian mixture model is used to fit the loss distribution of the complete instance; S23: Based on the memory effect of deep neural networks prioritizing the fitting of simple patterns, Gaussian components with smaller means are regarded as reliable pairing components, and instance-level pairing reliability is calculated based on the posterior probability that the loss belongs to the component with the lower loss mean.
[0046] In this embodiment, the cross-view step S21 Figure 1 The formula for the set of consistency loss is: in, For measuring two views Figure 1 Consistency loss function For the first Each view has a corresponding view-specific encoder. Representation of instances The One view, Indicates the first The first instance in Available states in each view Indicates the first An instance in the anchor view and target view Both are available in Chinese.
[0047] In this embodiment, the formula for fitting the loss distribution of the two-component Gaussian mixture model to the complete instance in step S22 is as follows: in, Represents probability. For measuring two views Figure 1 Consistency loss function Here are the parameters of the Gaussian mixture model, and k is a component of the Gaussian mixture model. For the first The mixing coefficient of each Gaussian component. For the first The probability density corresponding to each Gaussian component.
[0048] In this embodiment, the instance-level pairing reliability formula calculated in step S23 is: in, For the first The reliability of pairing two views in a complete instance. Represents probability. It is a low-loss mean Gaussian component. For measuring two views Figure 1 Consistency loss function.
[0049] In this embodiment, step S3 includes the following sub-steps: S31: Maintain a set of view-specific semantic prototypes for each view, and assign observation instances to each semantic prototype to obtain the semantic attribution vector of the instance; S32: Introduce instance-level pairing reliability into cross-view prototype co-occurrence estimation to construct a reliability-weighted prototype co-occurrence matrix; S33: Construct the cost matrix for optimal transmission based on the prototype co-occurrence matrix; S34: Solve the optimal transmission problem of the above cost matrix to obtain the cross-view prototype transmission scheme, and transmit the semantic structure in the anchor view to the target view.
[0050] In this embodiment, the semantic structure allocation formula from instance to prototype obtained in step S31 is as follows: in, For the first The first instance The view is assigned to the first The probability of a semantic prototype For example No. Features of each view For the first The first view in the A semantic prototype For temperature parameters, For similarity function, The number of semantic clusters or semantic prototypes.
[0051] In this embodiment, the formula for constructing the reliability-weighted prototype co-occurrence matrix in step S32 is as follows: in, For view With View The prototype co-occurrence matrix between them For view With View Between The reliability of pairing two views for an instance. and The first An instance in the view and view The prototype semantic attribution vector in the model.
[0052] In this embodiment, the formula for the optimal transmission cost matrix constructed in step S33 based on the prototype co-occurrence matrix is as follows: in, For view To view The transmission cost matrix, For view With View The prototype co-occurrence matrix between them It is an infinite norm.
[0053] In this embodiment, the formula for transferring the semantic structure from the anchor view to the target view in step S34 is: in, For anchor point view Transfer to target view The semantic attribution vector, For the first An instance in the view The semantic attribution vector, This is a cross-view prototype transfer scheme.
[0054] In this embodiment, step S4 includes the following sub-steps: S41: Approximate the posterior mean of potential paired views as a weighted combination of instance-level candidate paired views and cross-view semantic transport paired views; S42: Calculate adaptive weights based on instance-level pairing reliability and prototype semantic allocation consistency; S43: For complete instances with reliable pairing relationships, directly use the target view observations as potential pairing views; S44: For complete instances with unreliable pairing relationships, correct the target view observations using the transmitted semantic pairing views; S45: For incomplete instances where the target view is missing, combine the candidate paired views generated by the cross-view generator with the transmitted semantic paired views.
[0055] In this embodiment, the approximate formula for the posterior mean of the potential paired views in step S41 is: in, Represents the posterior expectation. This indicates the expected paired view of the anchor view. For adaptive weights, This indicates the amount of information observed. For candidate matching views in the target view, This is a semantic pairing view obtained from prototype-level semantic transfer. For the first An instance in the view Passing semantics to the view Later assigned to the The probability of a semantic prototype For view The first in A semantic prototype.
[0056] In this embodiment, the adaptive weight formula obtained in step S42 based on instance-level pairing reliability and prototype semantic allocation consistency is as follows: in, For adaptive weights, For the first The reliability of pairing two views in a complete instance. This represents the vector inner product or a semantic consistency measure. For anchor point view Transfer to target view The semantic attribution vector, This is the prototype semantic attribution vector in the target view.
[0057] In this embodiment, the potential pairing view for a complete instance of a fully reliable pairing relationship in step S43 is as follows: in, Represents the posterior expectation. This indicates the expected paired view of the anchor view. For adaptive weights, This indicates the amount of information observed. For the first One instance in the target view The observed features.
[0058] In this embodiment, step S44 provides the following potential pairing view for a complete instance of an unreliable pairing relationship: in, Represents the posterior expectation. This indicates the expected paired view of the anchor view. For adaptive weights, This indicates the amount of information observed. For the first One instance in the target view Observational features in For anchor point view Transfer to target view The semantic attribution vector, For view The semantic prototype in.
[0059] In this embodiment, the potential paired view for an incomplete instance where the target view is missing in step S45 is: in, Represents the posterior expectation. This indicates the expected paired view of the anchor view. For adaptive weights, This indicates the amount of information observed. These are candidate paired views for the target view, obtained by passing the anchor view through the cross-view generator. For view To view Cross-view generator, For anchor point view Transfer to target view The semantic attribution vector, For view The semantic prototype in.
[0060] In this embodiment, step S5 includes the following sub-steps: S51: Construct a joint optimization objective based on the inferred potential pairing views; S52: For a complete instance, construct cross-view contrast constraints using the inferred potential paired views; S53: For incomplete instances, use the transport semantic structure to constrain the semantic consistency of the inferred potential paired views; S54: Based on the joint optimization objective, update the view-specific encoder, cross-view generator, and semantic prototype to obtain a robust multi-view common representation, and perform clustering based on the common representation to output the clustering label of each instance.
[0061] In this embodiment, the joint optimization objective formula in step S51 is: in, For the total loss, For the loss term used to handle noise association, For the loss term used to handle incomplete views, To balance hyperparameters.
[0062] In this embodiment, the cross-view comparison constraint formula constructed in step S52 for the complete instance is as follows: in, For the loss term used to handle noise association, For measuring two views Figure 1 Consistency loss function For view The corresponding view-specific encoder, This indicates the expected paired view of the anchor view. By replacing potentially noisy original target view observations with potential paired views, the encoder can still learn robust cross-view alignment even under noisy associations.
[0063] In this embodiment, the regularization constraint formula for incomplete instances in step S53 is as follows: in, For the loss term used to handle incomplete views, Let KL divergence be the KL divergence. For anchor point view Transfer to target view The semantic attribution vector, This is the prototype semantic attribution vector of the inferred potential paired view.
[0064] In this embodiment, step S54 updates the view-specific encoder, cross-view generator, and semantic prototype based on the joint optimization objective to obtain a robust multi-view common representation, and performs clustering based on the common representation to output the clustering label of each instance.
[0065] To verify the effectiveness of the multi-view clustering method for imperfect information provided in this invention, this embodiment conducted experimental evaluations using several publicly available multi-view clustering datasets. These datasets include Scene15, LandUse21, Reuters, CCV20, and HandWritten, covering various data types such as video, images, text, and multilingual text. The specific datasets are described below: Scene15 contains 4485 images from 15 scene categories, and the experiments use two extracted image features as two modalities (PHOG, GIST); LandUse21 contains 2100 satellite images from 21 categories, and the experiments use two extracted image features as two modalities (PHOG, LBP); Reuters is a multilingual text dataset consisting of 18758 samples in 6 languages, and the experiments use English and French modalities; CCV20 contains 6773 web videos from 20 categories, and the experiments use Scale Invariant Feature Transform (SIFT) and Mel Frequency Cepstral Coefficients (MFCC) as visual and audio features, respectively; HandWritten contains 2000 handwritten digit images from 0-9, and the experiments use pixel features and Fourier coefficient features as two modalities.
[0066] In the experiment, an incomplete view was first simulated by randomly occluding part of the view. Then, the cross-view correspondence of some complete instances was randomly shuffled to simulate noise association. By default, the proportion of incomplete views was equal to the proportion of noise association. This patent defines this proportion as the Imperfect Information Rate (IIR) and conducts experiments under several different proportions of IIR.
[0067] In this embodiment, to evaluate the robustness of the method of the present invention in scenarios with imperfect information, the method of the present invention is compared with three representative multi-view clustering methods, including the incomplete view processing method (DIVIDE), the noise association processing method (CorrGen), and the method for incomplete information (CAMERA). The evaluation metrics used are clustering accuracy (ACC), normalized mutual information (NMI), and adjusted RAND index (ARI). The values of these three metrics range from 0 to 1, with higher values indicating better clustering performance. The experimental results are shown in Table 1.
[0068] Table 1 shows the average clustering performance of different methods on five public multiview datasets.
[0069] As shown in Table 1, under different imperfect information rate settings, the method of this invention outperforms the comparison methods in terms of average ACC, NMI, and ARI in five classic multi-view clustering datasets. In particular, it can still maintain high clustering performance in the severely imperfect information scenario with IIR=0.8, indicating that the present invention can effectively alleviate the degradation problem caused by the coexistence of incomplete views and noise associations, and has good robustness and generalization ability.
[0070] This embodiment provides a computer device, which includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, it causes the processor to perform the steps of the robust multi-view clustering method for imperfect information described above.
[0071] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.
[0072] The memory includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory may also be an external storage device of the computer device, such as a plug-in hard disk, smart memory card (SMC), secure digital storage card (SD card), flash memory card, etc., equipped on the computer device. Of course, the memory may include both internal storage units and external storage devices of the computer device. In this embodiment, the memory is often used to store the operating system and various application software installed on the computer device, such as the program code of the robust multi-view clustering method for imperfect information. In addition, the memory can also be used to temporarily store clustering results that have been output or will be output.
[0073] In some embodiments, the processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor is typically used to control the overall operation of the computer device. In this embodiment, the processor is used to run program code stored in the memory or process data, for example, to run the program code for the robust multi-view clustering method for imperfect information.
[0074] This embodiment provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it causes the processor to perform the steps of the robust multi-view clustering method for imperfect information described above.
[0075] The computer-readable storage medium stores a clustering program that can be executed by at least one processor to perform the steps of the robust multi-view clustering method for imperfect information as described above.
[0076] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a terminal device to execute the robust multi-view clustering method for imperfect information described in the embodiments of this application.
[0077] This invention provides a multi-view clustering method, system, and storage medium for imperfect information, which alleviates the performance degradation problem of multi-view clustering caused by imperfect multi-view information.
Claims
1. A multi-view clustering method, system and storage medium for imperfect information, characterized in that, The method includes: S1: Determine the availability status of each instance in the multi-view data in each view, construct an imperfect information modeling method, and model the target view that is correctly matched with the anchor view as a latent variable; S2: Calculate cross-view consistency loss for complete instances, estimate instance-level pairing reliability based on loss distribution, and identify reliable cross-view associations and potential noisy associations; S3: Construct semantic prototypes for each view, calculate semantic attribution vectors from instances to prototypes, construct a reliability-weighted cross-view prototype co-occurrence matrix based on instance-level pairing reliability, and obtain a cross-view prototype transmission scheme through optimal transmission. S4: Combining information from the observable view with prototype-level semantic transmission results, perform posterior inference on potential paired views in the target view; S5: Construct a joint optimization objective based on the inferred potential paired views, perform robust learning on noisy associations and incomplete views, obtain a common representation of multiple views, and output clustering results; Step S1 specifically includes: S11: Given a multi-view dataset, determine the observation state of each instance in different views and construct a common representation space in the multi-view clustering task so that multiple instances can be divided into different semantic clusters; S12: Define imperfect information data as a set of instances with incomplete views or noise associations; S13: For the observed instances in the anchor view, model their expected paired views in the target view as latent variables and establish a posterior distribution model; Step S2 specifically includes: S21: For a complete instance that contains both anchor view and target view observations, extract features using a view-specific encoder and compute a cross-view consistency loss set. S22: A two-component Gaussian mixture model is used to fit the loss distribution of the complete instance; S23: Based on the memory effect of deep neural networks prioritizing the fitting of simple patterns, Gaussian components with smaller means are regarded as reliable pairing components, and instance-level pairing reliability is calculated based on the posterior probability that the loss belongs to the component with the lower loss mean. Step S3 specifically includes: S31: Maintain a set of view-specific semantic prototypes for each view, and assign observation instances to each semantic prototype to obtain the semantic attribution vector of the instance; S32: Introduce instance-level pairing reliability into cross-view prototype co-occurrence estimation to construct a reliability-weighted prototype co-occurrence matrix; S33: Construct the cost matrix for optimal transmission based on the prototype co-occurrence matrix; S34: Solve the optimal transmission problem of the above cost matrix to obtain the cross-view prototype transmission scheme, and transmit the semantic structure in the anchor view to the target view; Step S4 specifically includes: S41: Approximate the posterior mean of potential paired views as a weighted combination of instance-level candidate paired views and cross-view semantic transport paired views; S42: Calculate adaptive weights based on instance-level pairing reliability and prototype semantic allocation consistency; S43: For complete instances with reliable pairing relationships, directly use the target view observations as potential pairing views; S44: For complete instances with unreliable pairing relationships, correct the target view observations using the transmitted semantic pairing views; S45: For incomplete instances where the target view is missing, combine the candidate paired views generated by the cross-view generator with the transmitted semantic paired views; Step S5 specifically includes: S51: Construct a joint optimization objective based on the inferred potential pairing views; S52: For a complete instance, construct cross-view contrast constraints using the inferred potential paired views; S53: For incomplete instances, use the transport semantic structure to constrain the semantic consistency of the inferred potential paired views; S54: Based on the joint optimization objective, update the view-specific encoder, cross-view generator, and semantic prototype to obtain a robust multi-view common representation, and perform clustering based on the common representation to output the clustering label of each instance. 2.The multi-view clustering method, system and storage medium for facing imperfect information according to claim 1, wherein, In step S12, imperfect information data is defined as a set of instances with incomplete views or noise associations, and its definition formula is: in, For multi-view datasets, For a set of instances contaminated by imperfect information, For the first One example, For the number of views, For view indexing, Indicates the first The first instance in Available states in each view This indicates that the view is unavailable. This indicates that the view is available. For anchor point view indexing, Index the target view. This is an indicator function that takes the value 1 when two view observations belong to the same instance, and 0 otherwise. Representation of instances The A view.
3. The multi-view clustering method, system, and storage medium for imperfect information according to claim 1, characterized in that, In step S13, for each observation instance in the anchor view, its expected paired view in the target view is modeled as a latent variable, and a posterior distribution model is established, the formula of which is: wherein, represents a posterior distribution model, represents a pair view desired by the target view, represents an observed amount of information. 4.The multi-view clustering method, system and storage medium for imperfect information, according to claim 1, wherein, In steps S32 to S34, the calculation formulas for the reliability-weighted prototype co-occurrence matrix, the optimal transmission cost matrix, and the cross-view semantic structure transmission are as follows: in, For view With View The prototype co-occurrence matrix between them For view With View Between The reliability of pairing two views for an instance. and The first An instance in the view and view The prototype semantic attribution vector in the text. For view To view The transmission cost matrix, It is an infinite norm. For anchor point view Transfer to target view The semantic attribution vector, For the first An instance in the view The semantic attribution vector, This is a cross-view prototype transfer scheme. 5.The multi-view clustering method, system and storage medium for imperfect information, according to claim 1, wherein, The approximate formula for the posterior mean of the latent pair view in step S41 is: in, Represents the posterior expectation. This indicates the expected paired view of the anchor view. For adaptive weights, This represents the amount of information observed. For candidate matching views in the target view, This is a semantic pairing view obtained from prototype-level semantic transfer. For the first An instance in the view Passing semantics to the view Later assigned to the The probability of a semantic prototype. For view The first in A semantic prototype.
6. The multi-view clustering method, system, and storage medium for imperfect information according to claim 1, characterized in that, The calculation formulas for the joint optimization objective, cross-view comparison constraint, and semantic consistency constraint in steps S51 to S53 are as follows: in, For the total loss, For the loss term used to handle noise association, For the loss term used to handle incomplete views, To balance hyperparameters, The loss function used to measure the consistency between two views. For view The corresponding view-specific encoder, Representation of instances The One view, This indicates the expected paired view of the anchor view. Let KL divergence be the KL divergence. For anchor point view Transfer to target view The semantic attribution vector, This is the prototype semantic attribution vector of the inferred potential paired view.