Clustering and model training methods, devices, equipment, and storage media

By constructing a unified target distribution representation and utilizing graph neural network model, the problem of poor clustering of characters in videos caused by relying on face information in the existing technology is solved, and effective clustering of multimodal data and adaptation to complex scenes is achieved.

CN114387650BActive Publication Date: 2025-07-01ZHEJIANG SENSETIME TECH DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210028976.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-11
Publication Date
2025-07-01
Estimated Expiration
2042-01-11

AI Technical Summary

Technical Problem

The prior art relies solely on facial information when characters are clustered in videos, resulting in missing clips of characters appearing when faces are invisible, and artificial design strategies are difficult to adapt to complex scenes.

Method used

By obtaining multimodal data sets, including face, human body and sound data, a unified target distribution representation is constructed, and a graph neural network model is used for clustering to achieve end-to-end multimodal clustering.

Benefits of technology

The clustering effect is improved, so that multimodal data belonging to the same object can be effectively clustered, reduce labor costs, and is suitable for complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114387650B_ABST
    Figure CN114387650B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a clustering and model training method, apparatus, device, and storage medium. The clustering method includes: obtaining a dataset to be processed, where the dataset to be processed includes at least two modality data, and the at least two modality data belong to at least one object; determining a target distribution representation for each of the modality data, where the target distribution representation for each of the modality data is at least used to represent the probability that each of the modality data belongs to the same object as other modality data in the dataset to be processed; and clustering each of the modality data based on each of the target distribution representations to obtain at least one clustering cluster, where each clustering cluster includes at least one modality data belonging to the same object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to, but is not limited to, the field of computer technology, and in particular, to a clustering and model training method, apparatus, device, and storage medium. Background Art

[0002] Clustering refers to the process of dividing a set of physical or abstract objects into multiple classes composed of similar objects. For example, person clustering refers to clustering various information (including face, body, and voice) of a specific person appearing in a video together, which is of great significance for video plot understanding, video editing, etc. Summary of the Invention

[0003] Embodiments of the present disclosure provide a clustering and model training method, apparatus, device, and storage medium.

[0004] The technical solution of the embodiments of the present disclosure is implemented as follows:

[0005] Embodiments of the present disclosure provide a clustering method, the method including:

[0006] Obtain a dataset to be processed, where the dataset to be processed includes at least two modality data, and the at least two modality data belong to at least one object;

[0007] Determine the target distribution representation of each modality data, where the target distribution representation of each modality data is at least used to represent the probability that each modality data belongs to the same object as other modality data in the dataset to be processed;

[0008] Based on the target distribution representation of each modality data, cluster each modality data to obtain at least one clustering cluster, and each clustering cluster includes at least one modality data belonging to the same object.

[0009] Embodiments of the present disclosure provide a model training method, the method including:

[0010] Obtain a sample set, where the sample set includes at least one data subset, each data subset includes at least two modality data, the at least two modality data belong to at least one object, and each modality data has label information;

[0011] Use a graph neural network model to be trained to determine the target distribution representation of each modality data in each data subset, where the target distribution representation of each modality data is at least used to represent the probability that each modality data belongs to the same object as other modality data in the data subset;

[0012] Based on the target distribution representation of each modality data in each data subset and the label information of each modality data, determine a target loss value;

[0013] When the target loss value meets the preset conditions, update the parameters of the graph neural network model.

[0014] An embodiment of the present disclosure provides a clustering device, which includes:

[0015] A first acquisition module, configured to acquire a dataset to be processed, where the dataset to be processed includes at least two modality data, and the at least two modality data belong to at least one object;

[0016] A first determination module, configured to determine the target distribution representation of each modality data, and the target distribution representation of each modality data is at least used to represent the probability that each modality data and other modality data in the dataset to be processed belong to the same object;

[0017] A first clustering module, configured to cluster each modality data based on each target distribution representation to obtain at least one clustering cluster, and each clustering cluster includes at least one modality data belonging to the same object.

[0018] An embodiment of the present disclosure provides a model training device, which includes:

[0019] A second acquisition module, configured to acquire a sample set, where the sample set includes at least one data subset, each data subset includes at least two modality data, the at least two modality data belong to at least one object, and each modality data has label information;

[0020] A second determination module, configured to use a graph neural network model to be trained to determine the target distribution representation of each modality data in each data subset, and the target distribution representation of each modality data is at least used to represent the probability that each modality data and other modality data in the data subset belong to the same object;

[0021] A third determination module, configured to determine a target loss value based on the target distribution representation and label value of each modality data in each data subset;

[0022] A first update module, configured to update the parameters of the graph neural network model when the target loss value meets the preset conditions.

[0023] An embodiment of the present disclosure provides an electronic device, including a processor and a memory, where the memory stores a computer program that can run on the processor, and the processor implements the above method when executing the computer program.

[0024] An embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above method is implemented.

[0025] In the embodiments of the present disclosure, by obtaining a dataset to be processed, the dataset to be processed includes at least two modality data, and the at least two modality data belong to at least one object; determining the target distribution representation of each modality data, and the target distribution representation of each modality data is at least used to represent the probability that each modality data and other modality data in the dataset to be processed belong to the same object; based on each target distribution representation, clustering each modality data to obtain at least one clustering cluster, and each clustering cluster includes at least one modality data belonging to the same object. In this way, by constructing a unified distribution representation for each modality data and using this distribution representation to cluster each modality data, at least one modality data belonging to the same object can be clustered, thereby improving the clustering effect.

[0026] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification. These drawings illustrate embodiments consistent with the present disclosure and are used together with the specification to explain the technical solutions of the present disclosure.

[0028] Figure 1 Schematic diagram of the implementation process of a clustering method provided by an embodiment of the present disclosure;

[0029] Figure 2 Schematic diagram of the implementation process of a clustering method provided by an embodiment of the present disclosure;

[0030] Figure 3A Schematic diagram of the implementation process of a clustering method provided by an embodiment of the present disclosure;

[0031] Figure 3B Schematic diagram of constructing a feature map based on a feature map neural network provided by an embodiment of the present disclosure;

[0032] Figure 4A Schematic diagram of the implementation process of a clustering method provided by an embodiment of the present disclosure;

[0033] Figure 4B Schematic diagram of constructing a distribution map based on a distribution map neural network provided by an embodiment of the present disclosure;

[0034] Figure 5 Schematic diagram of the implementation process of a model training method provided by an embodiment of the present disclosure;

[0035] Figure 6A Schematic diagram of a clustering system provided by an embodiment of the present disclosure;

[0036] Figure 6B Schematic diagram of the implementation of an initialization module provided by an embodiment of the present disclosure;

[0037] Figure 6C Schematic diagram of the initialization of distributed representation provided by an embodiment of the present disclosure;

[0038] Figure 6D Schematic diagram of retrieval based on features and distributed representation provided by an embodiment of the present disclosure;

[0039] Figure 7 Schematic diagram of the composition structure of a clustering device provided by an embodiment of the present disclosure;

[0040] Figure 8 Schematic diagram of the composition structure of a model training device provided by an embodiment of the present disclosure;

[0041] Figure 9 Schematic diagram of a hardware entity of an electronic device in an embodiment of the present disclosure. Detailed implementation manners

[0042] In order to make the objectives, technical solutions, and advantages of the present disclosure clearer, the present disclosure will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limitations on the present disclosure. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present disclosure.

[0043] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0044] In the following description, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein.

[0045] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this disclosure belongs. The terms used herein are only for the purpose of describing the embodiments of the present disclosure and are not intended to limit the present disclosure.

[0046] Clustering refers to the process of dividing a set of physical or abstract objects into multiple classes composed of similar objects. For example, person clustering refers to clustering various information (including face, body, and voice) of a specific person that appears in a video. Compared with traditional face clustering methods, person clustering not only needs to cluster the occurrences of faces but also needs to cluster the human body, clustering the face and body information of the same person together. This way of person clustering that includes multi-modal information of face and body can obtain all the information of a specific person's appearances in the video, which is of great significance for video plot understanding, video question answering, and video editing based on specific persons.

[0047] In related technologies, most existing methods for clustering persons in videos only use face information for clustering. However, in video scenarios, many times the face is not visible when the back is photographed or blocked. If only face information is used for clustering, many segments of a person's appearance will be missed, which has a great impact on obtaining the information of a specific person in the video completely and comprehensively. Currently, there are also a small number of methods that can process multi-modal information of persons in videos for clustering, but they all adopt different manually designed strategies for different modalities. Such a method of clustering different modalities with different strategies by manually designed rules cannot be well applied to more complex scenarios.

[0048] Embodiments of the present disclosure provide a clustering method. By constructing a unified distribution representation for each modality data and using this distribution representation to cluster each modality data, at least one modality data belonging to the same object can be clustered, thereby improving the clustering effect. The clustering method and model training method provided by the embodiments of the present disclosure can both be executed by an electronic device. The electronic device can be various types of terminals such as a laptop computer, a tablet computer, a desktop computer, a set-top box, a mobile device (such as a mobile phone, a portable music player, a personal digital assistant, a dedicated messaging device, a portable game device), etc., or can also be implemented as a server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0049] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure.

[0050] Figure 1The following is a schematic flowchart of an implementation process of a clustering method provided by an embodiment of the present disclosure. As Figure 1 shown, the method includes steps S11 to S13:

[0051] Step S11: Obtain a dataset to be processed, where the dataset to be processed includes at least two modal data, and the at least two modal data belong to at least one object.

[0052] Here, the modal data may include, but is not limited to, modal data containing modal information. For example, modal data containing a face, modal data containing a human body, modal data containing an iris, modal data containing sound, modal data containing a fingerprint, etc. Different modal data may be modal data belonging to the same object or different objects. In some embodiments, the object may include, but is not limited to, a person, an animal, etc.

[0053] In some embodiments, the dataset to be processed may include, but is not limited to, at least one type of modal data of the same object or different objects, etc. For example, the dataset to be processed may contain the same type of modal data of the same object. For example, the dataset to be processed includes modal data of the face of object A. Another example is that the dataset to be processed may contain at least two types of modal data of the same object, that is, the dataset to be processed contains at least two cross-modal data belonging to the same object. For example, the dataset to be processed includes modal data of the face of object A and modal data of the human body of object A. Another example is that the dataset to be processed may include the same type of model data of different objects. For example, the dataset to be processed may include modal data of the face of object A and modal data of the face of object B. Still another example is that the dataset to be processed may include at least two types of modal data of different objects. For example, the dataset to be processed includes modal data of the face of object A, modal data of the human body of object A, modal data of the face of object B, modal data of the sound of object B, etc.

[0054] In some embodiments, the dataset to be processed can be uploaded or set by a user through an operation interface, where the operation interface includes an interactive interface for configuring operations and information display on the dataset to be processed. The operation interface can be displayed on any suitable electronic device with interface interaction capabilities. In implementation, the electronic device displaying the operation interface and the device executing the clustering method can be the same or different, which is not limited here. For example, the electronic device executing the clustering can be a laptop computer, and the electronic device displaying the operation interface can also be the laptop computer. The operation interface can be the interactive interface of the client running on the laptop computer or the web page displayed in the browser running on the laptop computer. Another example is that the computer device executing the clustering method can be a server, and the electronic device displaying the operation interface can be a laptop computer. The operation interface can be the interactive interface of the client running on the laptop computer or the web page displayed in the browser running on the laptop computer. The laptop computer can access the server through the client or the browser.

[0055] In some embodiments, the dataset to be processed is obtained by performing feature extraction processing on the source data.

[0056] Here, the source data can include, but is not limited to, image sets, videos, etc.

[0057] For example, in an application scenario of finding the modal data of a target object in a target video, the user can input the specified modal data of the face containing the target object and the target video in the operation interface. After processing the target video, a dataset to be processed is obtained, where the dataset to be processed contains various types of modal data of multiple objects. For example, it contains modal data of the face, modal data of the human body, and modal data of the voice. The electronic device performs clustering on the dataset to be processed and outputs the modal data of the face containing the target object, the modal data of the human body, and the modal data of the voice to the same clustering result.

[0058] Step S12: Determine the target distribution representation of each piece of modal data. The target distribution representation of each piece of modal data is at least used to represent the probability that each piece of modal data belongs to the same object as other modal data in the dataset to be processed.

[0059] Here, the target distribution representation of each piece of modal data is a unified distribution representation, which has nothing to do with the modal information in the modal data.

[0060] In some embodiments, the target distribution representation of each piece of modal data can be determined based on a trained graph neural network model. In implementation, by inputting the dataset to be processed into the trained graph neural network model, the target distribution representation of each piece of modal data in the dataset to be processed can be obtained.

[0061] In some embodiments, the distribution representation of each modality data can be initialized based on the correlation relationship between every two modality data, and the distribution representation can be updated based on the update parameters to determine the target distribution representation of each modality data.

[0062] Here, the correlation relationship can include but is not limited to belonging to the same object, etc. In implementation, the correlation relationship can be preset by the user or obtained during the process of acquiring the dataset to be processed.

[0063] In some embodiments, the correlation relationship can be set by means of manual annotation, etc. For example, the user sets the correlation relationship between some modality data in the dataset to be processed through manual annotation.

[0064] In some embodiments, the source data can be subjected to feature extraction processing to obtain the correlation relationship. For example, in the case where the source data includes images of at least one object, feature extraction of the face and body is performed on the image pair to obtain modality data including the face and modality data including the body of at least one object, and the modality data including the face and the modality data including the body belonging to the same object are set as the correlation relationship.

[0065] The update parameters can include but are not limited to the first feature similarity between every two modality data within different modalities, the second feature similarity between every two modality data within the same modality, the fusion distribution correlation degree of every two modality data, etc. The first feature similarity can include but is not limited to the similarity between the correlation representations of every two modality data within different modalities. Here, the correlation representation of each modality data is used to represent the correlation information of each modality data. The second feature similarity can include but is not limited to the similarity between the correlation representations of every two modality data within the same modality. The fusion distribution correlation degree of every two modality data can include but is not limited to the correlation relationship between the distribution representations of every two modality data.

[0066] In some embodiments, the distribution representation of each modality data is initialized based on the correlation relationship between every two modality data, and the initialized distribution representation is updated at least once based on the update parameters.

[0067] For example, if the dataset to be processed includes N modality data, where N is a positive integer, then The represents the probability that the i-th modality data and all modality data 1, 2,......, N belong to the same object, and can be initialized based on each correlation relationship. In implementation, the following formula (1-1) can be used to initialize the distribution representation of each modality data:

[0068] ​

[0069] Among them, t(i) represents the object to which the i-th modal data belongs, t(j) represents the object to which the j-th modal data belongs, η is a probability value, and in the case where the i-th modal data and the j-th modal data belong to the same object in the association relationship, then this can be set to η. For example, η = 0.7; in the case where the i-th modal data and the j-th modal data belong to different objects in the association relationship, then this can be set to (1 - η). For example, η = 0.3. In implementation, those skilled in the art can select the initialization method of the distribution representation according to actual needs, and the embodiments of the present disclosure do not make limitations.

[0070] Step S13: Cluster each piece of the modal data based on each of the target distribution representations to obtain at least one cluster, and each cluster includes at least one piece of modal data belonging to the same object.

[0071] In some embodiments, clustering can be performed based on the similarity, distance, etc. between every two target distribution representations.

[0072] For example, determine the distance between the target distribution representations of every two pieces of modal data. In the case where the distance is not less than the first threshold, it indicates that these two pieces of modal data belong to the same object, and cluster these two pieces of modal data into the same class; in the case where the distance is less than the first threshold, it indicates that these two pieces of modal data belong to different objects, and cluster these two pieces of modal data into different classes. Among them, the first threshold can be similarity.

[0073] In some embodiments, step S13 includes:

[0074] Step S131: Cluster each piece of the modal data based on the similarity between every two of the target distribution representations to obtain at least one cluster.

[0075] In some embodiments, clustering of at least one type of modal data of the same object can be achieved based on the similarity between every two target distribution representations and a second threshold. Among them, the second threshold can include but is not limited to similarity, proximity, irrelevance, etc.

[0076] For example, when the second threshold is similarity, if the similarity between two target distribution representations is not less than the second threshold, it indicates that the two modal data are very similar. Then, the two modal data belong to the same object, and the two modal data are clustered into the same cluster. For another example, when the second threshold is dissimilarity, if the similarity between two target distribution representations is not less than the second threshold, it indicates that the two modal data are not similar. Then, the two modal data belong to different objects, and the two modal data are clustered into different clusters.

[0077] In the embodiments of the present disclosure, by obtaining a dataset to be processed, the dataset to be processed includes at least two modal data, and the at least two modal data belong to at least one object; determining the target distribution representation of each modal data, and the target distribution representation of each modal data is at least used to represent the probability that each modal data belongs to the same object as other modal data in the dataset to be processed; based on each target distribution representation, clustering each modal data to obtain at least one cluster, and each cluster includes at least one modal data belonging to the same object. In this way, by constructing a unified distribution representation for each modal data and using this distribution representation to cluster each modal data, it is possible to cluster the modal data of at least one modality belonging to the same object together, thereby improving the clustering effect.

[0078] Figure 2 It is a schematic flowchart of the implementation process of a clustering method provided by the embodiments of the present disclosure. As Figure 2 shown, the method includes steps S21 to S23:

[0079] Step S21, obtain a dataset to be processed, the dataset to be processed includes at least two modal data, and the at least two modal data belong to at least one object.

[0080] The above step S21 corresponds to the foregoing step S11. When implemented, the specific implementation manner of the foregoing step S11 can be referred to.

[0081] Step S22, use the trained graph neural network model to update the distribution representation of each modal data at least once. When the number of updates reaches a preset value, respectively determine the updated distribution representation of each modal data as the target distribution representation of each modal data, and the distribution representation of each modal data is at least used to represent the probability that each modal data belongs to the same object as other modal data in the dataset to be processed.

[0082] Here, the graph neural network model can be a model constructed based on the graph neural network. Among them, the graph neural network can represent the vertices in the graph as low-dimensional vectors by retaining the network topology structure and node content information of the graph, so as to be processed using simple machine learning algorithms (such as support vector machine classification, etc.) or deep learning algorithms. The preset value can include but is not limited to set empirical values, values calculated based on the clustering effects of multiple clusterings, etc. In implementation, those skilled in the art can independently determine the preset value according to actual needs, and the embodiments of the present disclosure do not make limitations.

[0083] In some embodiments, the trained graph neural network model at least includes a distribution graph neural network.

[0084] In some embodiments, the distribution graph neural network is used to construct a distribution graph of each modality data. The distribution graph includes at least one distribution node and the distribution connection relationship between the distribution nodes. Each distribution node is used to represent the distribution representation of each modality data, and each distribution connection relationship is used to represent the probability that every two distribution nodes belong to the same object. Based on the update parameters, at least one update is performed on each distribution node and its distribution connection relationship in the distribution graph constructed by the distribution graph neural network, and each distribution node in the last distribution graph is used as the target distribution representation of the modality data. Among them, the update parameters can include but are not limited to the first feature similarity between every two modality data of different modalities, the second feature similarity between every two modality data of the same modality, the fusion distribution correlation degree between every two modality data, etc.

[0085] In some embodiments, the trained graph neural network model at least includes a feature graph neural network and a distribution graph neural network. In implementation, the output parameters of the feature graph neural network are used as update parameters to update the distribution representation of each modality data in the distribution graph, and the updated distribution representation in the distribution graph is used as the update parameter to update the association representation in the feature graph neural network, so that multiple modality data belonging to the same object can be clustered into the same cluster. Among them, the feature graph neural network is used to construct a feature graph of each modality data. The feature graph includes at least one feature node and the feature connection relationship between the feature nodes. Each feature node is used to represent the association representation of each modality data. The association representation is used to represent the association information of each modality data, and each feature connection relationship is used to represent the feature association relationship between every two modality data. Among them, the feature association relationship can include a first feature connection relationship and a second feature connection relationship. The first feature connection relationship is used to represent the probability that every two feature nodes within the same modality belong to the same object, and the second feature connection relationship is used to represent that every two feature nodes between different modalities belong to the same object. In some embodiments, the first feature connection relationship can be represented by a first feature connection line, and the second feature connection relationship can be represented by the second feature connection line.

[0086] For example, the dataset to be processed includes N modal data, and the N modal data includes O types of modal data of M objects. At this time, the feature map can include O feature submaps, where each feature node in each feature submap is used to represent multiple modal data belonging to the same modal category. In implementation, according to the association relationship between every two modal data, multiple modal data belonging to the same modality can be connected through the first feature connection line, and modal data of different modalities belonging to the same object can be connected through the second feature connection line.

[0087] For instance, 200 modal data are the three types of modal data of face, body, and voice of three objects A, B. Among them, there are 50 face data of object A, 30 body data of object A, 25 voice data of object A, 30 face data of object B, 40 body data of object A, and 25 voice data of object A. At this time, according to the association relationship between every two modal data, 80 face data belonging to the face type are connected through the first feature connection line, 70 body data belonging to the body type are connected, and 50 voice data belonging to the voice type are connected; through the second feature connection line, the face data, body data, and voice data of object A are respectively connected, and the face data, body data, and voice data of object B are respectively connected.

[0088] Step S23: Based on each of the target distribution representations, perform clustering on each of the modal data to obtain at least one clustering cluster, and each clustering cluster includes at least one modal data belonging to the same object.

[0089] The above step S23 corresponds to the foregoing step S13. In implementation, the specific implementation manner of the foregoing step S13 can be referred to.

[0090] In an embodiment of the present disclosure, by obtaining a dataset to be processed, the dataset to be processed includes at least two modal data, and the at least two modal data belong to at least one object; using a trained graph neural network model, the distribution representation of each modal data is updated at least once, and when the number of updates reaches a preset value, the updated distribution representation of each modal data is respectively determined as the target distribution representation of each modal data, and the distribution representation of each modal data is at least used to represent the probability that each modal data and other modal data in the dataset to be processed belong to the same object; based on each target distribution representation, each modal data is clustered to obtain at least one cluster, and each cluster includes at least one modal data belonging to the same object. In this way, a unified target distribution representation can be constructed for each modal data through the graph neural network model, and each modal data is clustered using the target distribution representation, realizing end-to-end clustering of at least one modal data belonging to the same object, without the need to manually set and maintain a large number of modal fusion rules, which can not only greatly reduce the labor cost, but also improve the clustering effect.

[0091] Figure 3A FIG. is a schematic flowchart of an implementation of a clustering method provided by an embodiment of the present disclosure, as Figure 3A shown, the method includes steps S31 to S33:

[0092] Step S31, obtain a dataset to be processed, the dataset to be processed includes at least two modal data, and the at least two modal data belong to at least one object.

[0093] The above step S31 corresponds to the foregoing step S11, and in implementation, the specific implementation manner of the foregoing step S11 may be referred to.

[0094] Step S32, using the feature graph neural network in the trained graph neural network model, update the distribution representation of each modal data based on the association representation of each modal data, and when the number of updates reaches a preset value, respectively determine the updated distribution representation of each modal data as the target distribution representation of each modal data, and the target distribution representation of each modal data is at least used to represent the probability that each modal data and other modal data in the dataset to be processed belong to the same object, and each association representation is used to represent the association information of each modal data.

[0095] Here, the graph neural network model at least includes a feature graph neural network, which is used to construct a feature graph of each modal data, and the feature graph includes at least one feature node and a feature connection relationship between feature nodes, each feature node is used to characterize the association representation of each modal data, and each feature connection relationship is used to characterize the feature association relationship between each two modal data. The feature association relationship may include a first connection relationship and a second connection relationship, wherein the first connection relationship is used to characterize the probability that each two modal data in the same modality belong to the same object, and the second connection relationship is used to connect two modal data belonging to the same object in different modalities.

[0096] In some embodiments, the first connection relationship may be represented by a first characteristic connection line, and the second connection relationship may be represented by a second characteristic connection line.

[0097] In the case where the data set to be processed includes 6 modal data, the 6 modal data respectively represent two face data and one body data of object A, and one face data and two body data of object B. Figure 3B A schematic diagram of constructing a feature map based on a feature map neural network provided in an embodiment of the present disclosure, such as Figure 3B As shown, at this time, the feature graph 300 includes 6 feature nodes, the first feature node is the face data NP1 of object A, the second feature node is the face data NP2 of object A, the third feature node is the face data NP3 of object B, the fourth feature node is the body data NP4 of object A, the fifth feature node is the body data NP5 of object B, and the sixth feature node is the body data NP6 of object B, wherein the first feature node NP1 to the third feature node NP3 constitute a plurality of modal data belonging to the face category, the fourth feature node NP4 to the sixth feature node NP6 constitute a plurality of modal data belonging to the body category, and every two feature nodes in the face category and every two feature nodes in the body category are connected by a first feature connecting line SP1, and each first feature connecting line SP1 is used to characterize the probability that the two connected feature nodes belong to the same object, and two feature nodes belonging to the same object in the face category and the body category are connected by a second feature connecting line SP2, and each second feature connecting line SP2 is used to characterize that the two connected feature nodes belong to the same object.

[0098] Step S33: clustering each of the modal data based on each of the target distribution representations to obtain at least one cluster, each of the clusters including at least one modal data belonging to the same object.

[0099] The above step S33 corresponds to the above step S13. When implementing the step S33, reference may be made to the specific implementation of the above step S13.

[0100] In some embodiments, using the feature graph neural network in the trained graph neural network model to update the distribution representation of each modality data based on the correlation representation of each modality data includes steps S321 to S322:

[0101] Step S321: Using the trained feature graph neural network, based on the correlation representation of each modality data, determine the current fusion distribution correlation degree between every two modality data.

[0102] Here, the correlation representation is used to represent the correlation information of each modality data, and the current fusion distribution correlation degree is used to represent the correlation information between the distribution representations of every two modality data in the current time.

[0103] In some embodiments, the current distribution correlation degree between every two modality data can be determined based on the correlation representation, the second feature similarity, between every two modality data within the same modality, and the correlation relationship, the feature correlation degree, the first feature similarity, etc. between every two modality data within different modalities. Here, the correlation relationship can be used to represent that every two modality data belong to the same object, or every two modality data belong to different objects, and this correlation relationship can be preset by the user or obtained during the process of acquiring the dataset to be processed. The feature correlation degree can include but is not limited to the correlation relationship between every two modality data within different modalities.

[0104] Step S322: Based on the current fusion distribution correlation degree between every two modality data, update the distribution representation of each modality data.

[0105] In some embodiments, based on the current fusion distribution correlation degree between every two modality data, the previous distribution representation of each modality data can be updated to the current distribution representation of each modality data. During implementation, the distribution representation of each modality data can be updated through the following formula (3-1):

[0106]

[0107] Wherein, is used to represent the distribution representation of the i-th modality data in the l-th update, is a feature transformation block with a fully connected layer and a non-linear activation layer, and || represents the concatenation operation, is used to represent the distribution representation of the i-th modality data in the (l-1)-th update, is used to represent the current fusion distribution correlation degree between the i-th modality data and the j-th modality data in the l-th update.

[0108] In some embodiments, the dataset to be processed includes at least two sub - datasets, and each of the sub - datasets corresponds to a modality respectively. The step S321 includes steps S331 to S333:

[0109] Step S331: Determine the first feature similarity between the associated representations of every two modalities in the sub - datasets of different modalities.

[0110] Here, the first feature similarity is used to characterize the similarity between the associated representations of every two modalities in the sub - datasets of different modalities.

[0111] In some embodiments, the first feature similarity between every two modalities in the sub - datasets of different modalities can be determined based on the association relationship, feature correlation degree between every two modalities in the sub - datasets of different modalities, the second feature similarity between every two modalities in the sub - datasets of the same modality, etc.

[0112] Step S332: Determine the second feature similarity between the associated representations of every two modalities in the sub - datasets of the same modality.

[0113] Here, the second feature similarity is used to characterize the similarity between the associated representations of every two modalities in the sub - datasets of the same modality.

[0114] In some embodiments, the second feature similarity between every two modalities in the sub - datasets of the same modality can be determined based on the associated representations of every two modalities in the sub - datasets of the same modality.

[0115] In some embodiments, the step S332 includes step S3321:

[0116] Step S3321: Based on the previous associated representations of each modality data in the sub - datasets of the same modality, determine the second feature similarity between the associated representations of every two of the modality data.

[0117] Here, the second feature similarity is used to characterize the similarity between the previous associated representations of every two modalities in the sub - datasets of the same modality.

[0118] In some embodiments, for the second feature similarity in the l - th update can be expressed as In implementation, the second feature similarity between every two of the modality data can be determined by the following formula (3 - 2):

[0119]

[0120] Wherein, It is used to represent the associated representation of the i-th modal data after the (l - 1)-th update. o(i) is used to represent the modal type of the i-th modal data, and RELU(·) is the activation function.

[0121] Step S333: Based on each of the first feature similarities and each of the second feature similarities, determine the current fusion distribution correlation degree between every two modal data.

[0122] Here, the current fusion distribution correlation degree is used to characterize the correlation information between the distribution representations of every two modal data in the current time. In implementation, the current fusion distribution correlation degree between every two modal data can be determined by the following formula (3-3):

[0123]

[0124] Among them, is the normalization within the modal similarity matrix, is used to represent the similarity between the associated representations of every two modal data within the same modal in the l-th time, M is used to represent the association relationship between every two modal data, and α l is a hyperparameter used to balance the second similarity from the same modal and the first similarity from other modalities

[0125] In some embodiments, the step S331 includes steps S341 to S343:

[0126] Step S341: Determine the feature correlation degree between every two modal data in the subsets of different modalities.

[0127] Here, the feature correlation degree is used to characterize the association relationship between every two modal data in the subsets of different modalities.

[0128] In some embodiments, the feature correlation degree between every two modal data in the subsets of different modalities can be determined based on the association relationship between every two modal data in the subsets of different modalities.

[0129] Step S342: Perform normalization processing on each of the second feature similarities.

[0130] Here, the second feature similarity is used to represent the similarity between the associated representations of every two modal data in the subset of the same modal. In implementation, the normalization processing on each of the second feature similarities can be performed by the following formula (3-4):

[0131]

[0132] Among them, D lA diagonal matrix, used to represent the aggregated second feature similarity at the l-th time. Each element in this diagonal matrix can be obtained through the following formula (3-5):

[0133]

[0134] where, is used to represent the similarity between the associated representations of the i-th modal data and the j-th modal data within the same modality at the l-th time.

[0135] Step S343, based on each of the normalized second feature similarities and each of the feature correlation degrees, determine the first feature similarity between the associated representations of every two modal data.

[0136] Here, the first feature similarity is used to characterize the similarity between the associated representations of every two modal data in the sub-datasets of different modalities.

[0137] In some embodiments, the step S341 includes steps S351 to S352:

[0138] Step S351, determine the association relationship between every two modal data in the sub-datasets of different modalities.

[0139] Here, the association relationship can be set by the user in advance, or obtained during the process of acquiring the dataset to be processed.

[0140] Step S352, based on each of the association relationships, determine each of the feature correlation degrees.

[0141] Here, the feature correlation degree is used to characterize the association relationship between every two modal data in the sub-datasets of different modalities. During implementation, the first feature similarity between every two modal data can be calculated through the following formula (3-6):

[0142]

[0143] where t(i) represents the object to which the i-th modal data belongs, t(j) represents the object to which the j-th modal data belongs, and o(i) represents the modal type of the i-th modal data. In the case where the i-th modal data and the j-th modal data represent different modal types of the same object, this M i,j is 1, otherwise, M i,j is 0.

[0144] In an embodiment of the present disclosure, by obtaining a dataset to be processed, the dataset to be processed includes at least two modal data, and the at least two modal data belong to at least one object; using a feature graph neural network in a trained graph neural network model, based on the associated representation of each modal data, update the distribution representation of each modal data. When the number of updates reaches a preset value, determine the updated distribution representation of each modal data as the target distribution representation of each modal data respectively. The target distribution representation of each modal data is at least used to represent the probability that each modal data and other modal data in the dataset to be processed belong to the same object, and each associated representation is used to represent the associated information of each modal data; based on each target distribution representation, perform clustering on each modal data to obtain at least one clustering cluster, and each clustering cluster includes at least one modal data belonging to the same object. In this way, by updating the distribution representation of each modality based on the associated representation of each modal data through the feature graph neural network of the model, a more accurate distribution representation of each modal data can be obtained, thereby improving the clustering effect.

[0145] Figure 4A It is a schematic flowchart of the implementation process of a clustering method provided by an embodiment of the present disclosure, as Figure 4A shown, the method includes steps S41 to S43:

[0146] Step S41, obtain a dataset to be processed, the dataset to be processed includes at least two modal data, and the at least two modal data belong to at least one object.

[0147] The above step S41 corresponds to the foregoing step S11. When implemented, the specific implementation manner of the foregoing step S11 can be referred to.

[0148] Step S42, use a trained graph neural network model to update the distribution representation and the associated representation of each modal data at least once. When the number of updates reaches a preset value, determine the updated distribution representation of each modal data as the target distribution representation of each modal data respectively. Each associated representation is used to represent the associated information of each modal data, and the target distribution representation of each modal data is at least used to represent the probability that each modal data and other modal data in the dataset to be processed belong to the same object.

[0149] In some embodiments, the graph neural network model includes a feature graph neural network and a distribution graph neural network. The use of a trained graph neural network model to update the distribution representation and the associated representation of each modal data at least once includes steps S421 to S422:

[0150] Step S421: Using the trained feature map neural network, update the distribution representation of each modality data based on the associated representation of each modality data.

[0151] The above step S421 corresponds to the foregoing steps S321 to S322. When implemented, the specific implementation manners of the foregoing steps S321 to S322 may be referred to.

[0152] Step S422: Using the trained distribution map neural network, update the associated representation of each modality data based on the updated distribution representation of each modality data.

[0153] Here, the trained graph neural network model at least includes a distribution map neural network, which is used to construct a distribution map of each modality data. The distribution map includes at least one distribution node and the distribution connection relationship between the distribution nodes. Each distribution node is used to represent the distribution representation of each modality data, and each distribution connection relationship is used to represent the probability that every two distribution nodes belong to the same object.

[0154] In some embodiments, the distribution connection relationship may be represented by a first distribution connection line.

[0155] In the case where the dataset to be processed includes 6 modality data, which respectively represent two face data and one body data of object A, and one face data and two body data of object B, Figure 4B FIG. is a schematic diagram of constructing a distribution map based on a distribution map neural network provided by an embodiment of the present disclosure. As Figure 4B shown, at this time, the distribution map 410 includes 6 distribution nodes, namely distribution nodes ND1 to ND6. Each distribution node is used to represent the distribution representation of each modality data. Every two distribution nodes are connected by a first distribution connection line SD1, and each first distribution connection line SD1 is used to represent the probability that the two connected distribution nodes belong to the same object.

[0156] In some embodiments, the step S422 includes steps S431 to S432:

[0157] Step S431: Using the trained distribution map neural network, determine the current fusion feature correlation degree between every two modality data based on the updated distribution representation of each modality data.

[0158] Here, the current fusion feature correlation degree is used to represent the correlation relationship between the associated representations of every two modality data in this current time.

[0159] In some embodiments, for the aggregation feature correlation degree at the l-th time It can be expressed as In implementation, the current fusion feature correlation between each two modal data can be determined by the following formula (4-1):

[0160]

[0161] in, It is used to represent the distribution representation of the i-th modal data in the th time. The similarity block of the distribution representation of two fully connected layers and one activation layer is obtained.

[0162] Step S432: based on the correlation degree of each current fusion feature, update the correlation representation of each modality data.

[0163] In some embodiments, for the first aggregation feature correlation It can be expressed as In implementation, the associated representation of each modal data can be determined by the following formula (4-2):

[0164]

[0165] Among them, o(i) represents the mode type of the i-th node, represents the correlation degree of the fusion features between the i-th modal data and the j-th modal data in the l-th time, represents the association representation of the jth node in the l-1th order, is a learnable gated residual block, and m is the total number of modal data.

[0166] Step S43: clustering each of the modal data based on each of the target distribution representations to obtain at least one cluster, each of the clusters including at least one modal data belonging to the same object.

[0167] The above step S43 corresponds to the above step S13. When implementing the step S43, reference may be made to the specific implementation of the above step S13.

[0168] In the embodiments of the present disclosure, by obtaining a dataset to be processed, where the dataset to be processed includes at least two modal data, and the at least two modal data belong to at least one object; using a trained graph neural network model, at least one update is performed on the distribution representation and the association representation of each modal data. When the number of updates reaches a preset value, the updated distribution representation of each modal data is respectively determined as the target distribution representation of each modal data. Each association representation is used to represent the association information of each modal data, and the target distribution representation of each modal data is at least used to represent the probability that each modal data and other modal data in the dataset to be processed belong to the same object; based on each target distribution representation, each modal data is clustered to obtain at least one cluster, and each cluster includes at least one modal data belonging to the same object. In this way, through the feature graph neural network of the model, the distribution representation of each modality is updated based on the association representation of each modal data, and through the distribution graph neural network of the model, the association representation of each modal data is updated based on the updated distribution representation of each modal data. In this way of cyclic update and mutual enhancement, a more accurate distribution representation of each modal data can be obtained, thereby improving the clustering effect.

[0169] Figure 5 It is a schematic flowchart of the implementation process of a model training method provided by the embodiments of the present disclosure, as Figure 5 shown, the method includes steps S51 to S54:

[0170] Step S51, obtain a sample set, where the sample set includes at least one data subset, each data subset includes at least two modal data, the at least two modal data belong to at least one object, and each modal data has label information.

[0171] Here, the modal data may include, but is not limited to, modal data containing modal information. For example, modal data containing a face, modal data containing a human body, modal data containing an iris, modal data containing a voice, modal data containing a fingerprint, etc. Different modal data may be modal data belonging to the same object or different objects. In some embodiments, the object may include, but is not limited to, a person, an animal, etc.

[0172] In some embodiments, each data subset may include, but is not limited to, at least one type of modal data of the same object or different objects, etc.

[0173] The label information is used to indicate the object to which each modal data belongs. In some embodiments, it is determined whether two modal data come from the same object by comparing their label information.

[0174] In some embodiments, the sample set can be uploaded or set by the user through the operation interface, or can be obtained by performing feature extraction processing on the source data. Here, the source data can include, but is not limited to, image sets, videos, etc.

[0175] Step S52: Use the graph neural network model to be trained to determine the target distribution representation of each modality data in each data subset, and the target distribution representation of each modality data is at least used to represent the probability that each modality data belongs to the same object as other modality data in the data subset.

[0176] Here, the graph neural network model can be a model constructed based on the graph neural network. During implementation, input each data subset into the graph neural network model to be trained, and the target distribution representation of each modality data in each data subset can be obtained.

[0177] Step S53: Based on the target distribution representation of each modality data in each data subset and the label information of each modality data, determine the target loss value.

[0178] Here, the target loss value is used to represent the difference between the label information of each modality data and the target distribution representation.

[0179] Step S54: Update the parameters of the graph neural network model when the target loss value meets the preset conditions.

[0180] Here, the preset conditions can include, but are not limited to, meeting the convergence conditions, etc. Among them, the convergence conditions can include, but are not limited to, the target loss value being greater than a threshold, etc. During implementation, those skilled in the art can independently determine the preset conditions according to actual needs, and the embodiments of the present disclosure do not make limitations.

[0181] The parameters of the graph neural network model can include, but are not limited to, learnable gated residual blocks The similarity block σ of the distribution representation, the feature transformation block ψ, etc.

[0182] In some embodiments, when the target loss value does not meet the preset conditions, use the current graph neural network model as the trained graph neural network model.

[0183] For example, when the target loss value is less than the threshold, use the current graph neural network model as the trained graph neural network model.

[0184] In some embodiments, step S52 includes:

[0185] Step S521: Use the graph neural network model to be trained to update the distribution representation of each modality data at least once.

[0186] In some embodiments, the graph neural network model to be trained includes at least a distribution graph neural network, which is used to construct a distribution graph of each modality data. The distribution graph includes at least one distribution node and distribution connection relationships between the distribution nodes. Each distribution node is used to represent the distribution representation of each modality data, and each distribution connection relationship is used to represent the probability that every two distribution nodes belong to the same object. The target distribution representation of each modality data is determined through the distribution graph neural network. Based on the update parameters, each distribution node and its distribution connection relationships in the distribution graph constructed by the distribution graph neural network are updated at least once, and each distribution node in the last distribution graph is used as the target distribution representation of the modality data. Among them, the update parameters may include, but are not limited to, the first feature similarity between every two modality data within different modalities, the second feature similarity between every two modality data within the same modality, the fusion distribution correlation degree of every two modality data, etc. The first feature similarity may include, but is not limited to, the similarity between the correlation representations of every two modality data within different modalities. Here, the correlation representation of each modality data is used to represent the correlation information of each modality data. The second feature similarity may include, but is not limited to, the similarity between the correlation representations of every two modality data within the same modality. The fusion distribution correlation degree of every two modality data may include, but is not limited to, the correlation relationship between the distribution representations of every two modality data.

[0187] In some embodiments, the distribution representation of each modality data can be initialized based on the correlation relationship between every two modality data, and the distribution representation is updated based on the update parameters to determine the target distribution representation of each modality data. Among them, the correlation relationship may include, but is not limited to, belonging to the same object, etc. In implementation, the correlation relationship can be preset by the user or obtained during the process of obtaining the sample set.

[0188] In implementation, the distribution representation of each modality data is initialized through the correlation relationship between every two modality data, and based on the update parameters, the initialized distribution representation is updated at least once.

[0189] For example, the dataset to be processed includes N modality data, where N is a positive integer. Then The represents the probability that the i-th modality data and all modality data 1, 2,......, N belong to the same object. Based on the correlation relationship between every two modality data, the above formula (1-1) can be used to be initialized. In implementation, those skilled in the art can select the initialization method of the distribution representation according to actual needs, and the embodiments of the present disclosure do not make limitations.

[0190] In some embodiments, the output parameters of the feature map neural network can be used as update parameters to update the distribution representation of each modal data in the distribution map, and the updated distribution representation in the distribution map can be used as update parameters to update the association representation in the feature map neural network. Through the way of cyclic update, a more accurate distribution representation of each modal data can be obtained. Among them, the feature map neural network is used to construct the feature map of each modal data, and the feature map includes at least one feature node and the feature connection relationship between the feature nodes. Each feature node is used to represent the association representation of each modal data, the association representation is used to represent the association information of each modal data, and each feature connection relationship is used to represent the feature association relationship between every two modal data. Among them, the feature association relationship can include a first feature connection relationship and a second feature connection relationship. The first feature connection relationship is used to represent the probability that every two feature nodes within the same modality belong to the same object, and the second feature connection relationship is used to represent that every two feature nodes within different modalities belong to the same object. In some embodiments, the first feature connection relationship can be represented by a first feature connection line, and the second feature connection relationship can be represented by the second feature connection line.

[0191] Step S522: When the number of updates reaches the preset value, respectively determine the updated distribution representation of each modal data as the target distribution representation of each modal data.

[0192] Here, the number of updates can include but is not limited to set empirical values, calculated based on the clustering effects of multiple clusterings, etc. In implementation, those skilled in the art can independently determine the number of updates according to actual needs, and the embodiments of the present disclosure do not make limitations.

[0193] In some embodiments, the graph neural network model includes a feature map neural network, and step S521 includes step S5211:

[0194] Step S5211: Use the feature map neural network to be trained to update the distribution representation of each modal data based on the association representation of each modal data, and each association representation is used to represent the association information of each modal data.

[0195] Here, the graph neural network model to be trained at least includes a feature graph neural network to be trained, and the feature graph neural network is used to construct a feature graph of each modality data. The feature graph includes at least one feature node and the feature connection relationships between the feature nodes. Each feature node is used to represent the associated representation of each modality data, and each feature connection relationship is used to represent the feature association relationship between every two modality data. Among them, the feature association relationship may include a first connection relationship and a second connection relationship. The first connection relationship is used to represent the probability that every two modality data within the same modality belong to the same object, and the second connection relationship is used to connect two modality data belonging to the same object within different modalities. In some embodiments, the first connection relationship can be represented by a first feature connection line, and the second connection relationship can be represented by the second feature connection line.

[0196] In some embodiments, the graph neural network model further includes a distribution graph neural network, and the method further includes:

[0197] Step S5212: Using the distribution graph neural network to be trained, based on the updated distribution representation of each modality data, update the associated representation of each modality data.

[0198] Here, the graph neural network model to be trained at least includes a distribution graph neural network, and the distribution graph neural network is used to construct a distribution graph of each modality data. The distribution graph includes at least one distribution node and the distribution connection relationships between the distribution nodes. Each distribution node is used to represent the distribution representation of each modality data, and each distribution connection relationship is used to represent the probability that every two distribution nodes belong to the same object. In some embodiments, the distribution connection relationship can be represented by a first distribution connection line.

[0199] During implementation, the associated representation of each modality data can be initialized in advance, and based on the updated distribution representation of each modality data, the initialized associated representation is updated at least once.

[0200] For example, the dataset to be processed includes N modality data of three types: face, body, and voice, where N is a positive integer. Then can be represented as (f1,......, f pq , b1,......, b pq , u1,......, u p ), where f i , b j , u krespectively indicate that the i-th modal data is modal data including a face, the i-th modal data is modal data including a body, and the k-th modal data is modal data including sound. In implementation, those skilled in the art can select the initialization method of the associated representation according to actual needs, which is not limited in the embodiments of the present disclosure.

[0201] In some embodiments, each of the data subsets includes at least two sub-datasets, and each of the sub-datasets corresponds to a modality respectively. The step S53 includes:

[0202] Step S531: Determine a feature similarity loss value based on the second feature similarity between the associated representations of every two modal data in the sub-datasets of the same modality in each update and the label information of each modal data.

[0203] Here, the second feature similarity is used to characterize the similarity between the associated representations of every two modal data in the sub-datasets of the same modality. The label information is used to indicate the object to which each modal data belongs.

[0204] In implementation, the feature similarity loss value can be determined by the following formula (5-1):

[0205]

[0206] where L represents the number of updates; N represents the total number of modal data; is used to characterize the second feature similarity between the i-th modal data and the j-th modal data within the same modality in the l-th time; BCE represents binary cross-entropy loss; o(i) represents the modality type of the i-th modal data; is a weight value; II(x) is an indicator function, when x is 1, II(x) is 1, and when x is 0, II(x) is 0; y i,j is a joint label. When the i-th modal data and the j-th modal data belong to the same object, y i,j is 1, otherwise y i,j is 0.

[0207] In some embodiments, the joint label can be determined based on the label information of each modal data.

[0208] For example, the label information i indicates that the i-th modal data belongs to object A, and the label information j indicates that the j-th modal data belongs to object A. At this time, the joint label y i,i is 1. For another example, the label information i indicates that the i-th modal data belongs to object A, and the label information j indicates that the j-th modal data belongs to object B. At this time, the joint label y i,j is 0.

[0209] Step S532: Determine the distribution similarity loss value based on the distribution similarity between the distribution representations of every two modal data in each historical update, the distribution similarity between the target distribution representations of every two modal data, and the label information of each modal data.

[0210] Here, the distribution similarity is used to represent the similarity between the distribution representations of every two modal data. The label information is used to indicate the object to which each modal data belongs.

[0211] In implementation, the feature similarity loss value can be determined by the following formula (5-2):

[0212]

[0213] where L represents the number of updates; N represents the total number of modal data; represents the similarity between the distribution representations of the i-th modal data and the j-th modal data in the l-th time; BCE represents the binary cross-entropy loss; is the weight value; y i,j is a joint label. When the i-th modal data and the j-th modal data belong to the same object, y i,j is 1, otherwise y i,j is 0.

[0214] Step S533: Determine the target loss value based on the feature similarity loss value and the distribution similarity loss value.

[0215] In some embodiments, the feature similarity loss value can be determined by the following formula (5-3):

[0216]

[0217] where λ f and λ d are both hyperparameters, is the feature similarity loss value, is the distribution similarity loss value.

[0218] In the embodiments of the present disclosure, based on a preset sample set with label information, the graph neural network model is trained to achieve end-to-end learning. In this way, the trained graph neural network can construct a unified target distribution representation for each modal data, and use this target distribution representation to cluster each modal data, realizing end-to-end clustering of at least one type of modal data belonging to the same object, without the need to manually set and maintain a large number of modal fusion rules, which can not only greatly reduce the labor cost, but also improve the clustering effect.

[0219] Figure 6ASchematic diagram of a clustering system 60 provided by an embodiment of the present disclosure, as Figure 6A shown. The clustering system 60 includes an initialization module 61, an update module 62, and a clustering module 63, where:

[0220] The initialization module 61 is configured to initialize the graph neural network model based on the dataset to be processed, where the dataset to be processed includes at least two modality data, and the at least two modality data belong to at least one object.

[0221] The update module 62 is configured to update the distribution representation of each modality data at least once based on the graph neural network. When the number of updates reaches a preset value, the updated distribution representation of each modality data is respectively determined as the target distribution representation of each modality data. The target distribution representation of each modality data is at least used to represent the probability that each modality data belongs to the same object as other modality data in the dataset to be processed.

[0222] The clustering module 63 is configured to cluster each modality data based on the target distribution representation of each modality data to obtain at least one clustering cluster, and each clustering cluster includes at least one modality data belonging to the same object.

[0223] In some embodiments, the graph neural network model includes a feature graph neural network and a distribution graph neural network, and the initialization module 61 includes a feature initialization module and a distribution initialization module.

[0224] The feature initialization module is configured to construct a feature graph of each modality data based on the trained feature graph neural network. The feature graph includes at least one feature node and the feature connection relationship between the feature nodes. Each feature node is used to represent the associated representation of each modality data, and each feature connection relationship is used to represent the feature association relationship between every two modality data. Among them, the feature association relationship may include a first feature connection relationship and a second feature connection relationship. The first feature connection relationship is used to represent the probability that every two feature nodes within the same modality belong to the same object, and the second feature connection relationship is used to represent that every two feature nodes in different modalities belong to the same object.

[0225] In some embodiments, the first feature connection relationship can be represented by a first feature connection line, and the second feature connection relationship can be represented by the second feature connection line.

[0226] The distribution initialization module is used to construct a distribution map of each modality data based on the trained distribution map neural network. The distribution map includes at least one distribution node and distribution connection relationships between the distribution nodes. Each distribution node is used to represent the distribution representation of each modality data, and each distribution connection relationship is used to represent the probability that every two distribution nodes belong to the same object.

[0227] In some embodiments, the distribution connection relationship can be represented by a first distribution connection line.

[0228] For example, the dataset to be processed includes 6 modality data, which respectively represent two face data and one body data of object A, and one face data and two body data of object B. Then, through the feature initialization module and the distribution initialization module, the initialized feature map and distribution map can be obtained. Figure 6B It is a schematic diagram of the implementation of an initialization module 61 provided by an embodiment of the present disclosure, as Figure 6B shown. At this time, the feature map 610 includes 6 feature nodes. The first feature node is the face data NP1 of object A, the second feature node is the face data NP2 of object A, the third feature node is the face data NP3 of object B, the fourth feature node is the body data NP4 of object A, the fifth feature node is the body data NP5 of object B, and the sixth feature node is the body data NP6 of object B. Among them, the first feature node NP1 to the third feature node NP3 constitute multiple modality data belonging to the face category, and the fourth feature node NP4 to the sixth feature node NP6 constitute multiple modality data belonging to the body category. Each two feature nodes in the face category and each two feature nodes in the body category are connected by a first feature connection line SP1, and two feature nodes belonging to the same object in the face category and the body category are connected by a second feature connection line SP2. Each of the first feature connection lines SP1 is used to represent the probability that the two connected feature nodes belong to the same object, and each of the second feature connection lines SP2 is used to represent the probability that the two connected feature nodes belong to the same object; the distribution map 620 includes 6 distribution nodes, namely distribution nodes ND1 to ND6. Each distribution node is used to represent the distribution representation of each modality data, and each two distribution nodes are connected by a first distribution connection line SD1. Each of the first distribution connection lines SD1 is used to represent the probability that the two connected distribution nodes belong to the same object.

[0229] In some embodiments, the update module 62 may include a distribution update module and a feature update module.

[0230] The distribution update module is configured to use the trained feature map neural network to update the distribution representation of each modality data based on the correlation representation of each modality data, where each correlation representation is used to represent the correlation information of each modality data.

[0231] The feature update module is configured to use the trained distribution map neural network to update the correlation representation of each modality data based on the updated distribution representation of each modality data.

[0232] In some embodiments, the distribution update module may include a fusion distribution correlation degree module and a distribution representation update module. Among them, the fusion distribution correlation degree module is configured to use the trained feature map neural network to determine the current fusion distribution correlation degree between every two modality data based on the correlation representation of each modality data. The distribution representation update module is configured to update the distribution representation of each modality data based on the current fusion distribution correlation degree between every two modality data.

[0233] In some embodiments, the dataset to be processed includes at least two sub-datasets, and each sub-dataset corresponds to one modality respectively; the fusion distribution correlation degree module may include a first feature similarity module, a second feature similarity module, and a current fusion distribution correlation degree module. Among them, the first feature similarity module is configured to determine the first feature similarity between the correlation representations of every two modality data in the sub-datasets of different modalities; the second feature similarity module is configured to determine the second feature similarity between the correlation representations of every two modality data in the sub-datasets of the same modality; the current fusion distribution correlation degree module is configured to determine the current fusion distribution correlation degree between every two modality data based on each first feature similarity and each second feature similarity.

[0234] In some embodiments, the first feature similarity module may include a feature correlation degree module, a normalization module, and a current first feature similarity module. Among them, the feature correlation degree module is configured to determine the feature correlation degree between every two modality data in the sub-datasets of different modalities; the normalization module is configured to perform normalization processing on each second feature similarity; the current first feature similarity module is configured to determine the first feature similarity between the correlation representations of every two modality data based on each normalized second feature similarity and each feature correlation degree.

[0235] In some embodiments, the feature correlation degree module may include a correlation relationship module and a current first feature similarity module. Among them, the correlation relationship module is configured to determine the correlation relationship between every two modality data in the sub-datasets of different modalities; the current first feature similarity module is configured to determine each feature correlation degree based on each correlation relationship.

[0236] In some embodiments, the second feature similarity module may include a current second feature similarity module. The current second feature similarity module is configured to determine a second feature similarity between the correlation representations of every two pieces of the modality data based on the previous correlation representation of each piece of modality data in a sub-dataset of the same modality.

[0237] In some embodiments, the feature update module may include a fused feature correlation degree module and a correlation representation module. The fused feature correlation degree module is configured to use the trained distribution graph neural network to determine a current fused feature correlation degree between every two modality data based on the updated distribution representation of each piece of modality data. The correlation representation module is configured to update the correlation representation of each piece of modality data based on each current fused feature correlation degree.

[0238] In some embodiments, the clustering module 63 includes a current clustering module. The current clustering module is configured to cluster each piece of modality data based on the similarity between every two target distribution representations to obtain at least one clustering cluster.

[0239] Compared with the clustering method in the related art, the clustering method provided by the embodiments of the present disclosure has at least the following improvements:

[0240] 1) In the related art, the method for clustering people only uses face information for clustering. There are also a few clustering methods for processing multi-modal information, but they all only perform clustering in the feature space. Since the features of different modalities are implemented by different feature extraction networks, these features are modality-related and cannot be directly compared in terms of similarity, so clustering cannot be directly performed. In the embodiments of the present disclosure, by constructing a unified distribution representation for each piece of modality data, each piece of modality data is independent of the modality type, so that multi-modal clustering can be achieved.

[0241] 2) In the related art, the supervised clustering method through network learning usually only performs enhancement and update in the feature space. In the embodiments of the present disclosure, by constructing a distribution representation for each piece of modality data in the distribution space and using the mutual enhancement and cyclic update of the feature graph neural network and the distribution graph neural network, a better distribution representation can be obtained.

[0242] 3) In the related art, there are a few clustering methods for processing multi-modal information that design different strategies to process information of different modalities. Because they rely on artificially designed rules, it is difficult to process complex data distributions in actual scenarios. In the embodiments of the present disclosure, by constructing a unified distribution representation for each piece of modality data, each piece of modality data is independent of the modality type. By calculating the similarity between the distribution representations of every two modality data, clustering of multiple modality data belonging to the same object can be achieved.

[0243] The clustering method provided by the embodiments of the present disclosure has at least the following beneficial effects: 1) By constructing the distribution representation of each modality data, the modality data is made independent of the modality type, so that multi-modal clustering can be achieved; 2) The feature map neural network of the model updates the distribution representation of each modality based on the correlation representation of each modality data, and the distribution map neural network of the model updates the correlation representation of each modality based on the updated distribution representation of each modality data. In this way of cyclic updating and mutual enhancement, a more accurate target distribution representation of each modality data can be obtained, thereby improving the clustering effect; 3) Using the target distribution representation of each modality data to cluster each modality data realizes end-to-end clustering of the modality data of at least one modality belonging to the same object, without the need to manually set and maintain a large number of modality fusion rules, which can not only greatly reduce the labor cost, but also improve the clustering effect.

[0244] To better illustrate the beneficial effects of the embodiments of the present disclosure, the experimental data of the clustering method provided by the embodiments of the present application and the clustering methods in the related art are compared and described below.

[0245] (1) Dataset to be processed

[0246] The dataset VPCD is used, which includes 5 datasets: Movie 1, TV Series 2, Movie 3, TV Series 4, and TV Series 5. Among them, a TV series can include one or several episodes, and each episode contains at least one object. These 5 datasets include a total of 32,999 face modality data, 36,724 body modality data, and 9,863 voice modality data.

[0247] (2) Evaluation metrics

[0248] Normalized Mutual Information (NMI), Weighted Cluster Purity (WCP), Character Precision (CP), Recall (CR), and F-score (CF) are used respectively. The higher the values of these metrics, the more accurate the clustering result.

[0249] (3) Verification method

[0250] Cross-validation is used to evaluate the clustering performance of the graph neural network model. Specifically, four of the five subsets, namely Movie 1, TV Series 2, Movie 3, and TV Series 5, are selected as the training set, and the other TV Series 4 is used as the test set.

[0251] (4) Initialization of some parameters in the graph neural network model

[0252] All experiments used the Adam optimizer with an initial learning rate of 10 -3 , and the learning rate decayed by 0.1. The number of updates was set to 2. The loss weights λ f and λ d represent the hyperparameters of the feature similarity loss value and the distribution similarity loss value . For all datasets except Drama 4, λ f was fixed at 1, and λ d was set to 0.2. For Drama 4, λ d was set to 0.3.

[0253] (5) Experimental results

[0254] 1) Compared with existing clustering methods

[0255] Table 1 shows the clustering results of the clustering method provided in the embodiments of the present disclosure and the clustering methods in related technologies. The MuHPC clustering method uses face, body, and voice information for person clustering, so its performance is better than that of B-ReID (only including body information) and B-C1C (including face and body information). Compared with the MuHPC, a clustering method in the prior art, the clustering method provided in the embodiments of the present disclosure, namely the Modality-Agnostic distribution Graph NETwork (MAGNET), is superior by 6.1% in WCP, 2.5% in NMI, and 6% in CF. MuHPC manually designed three different rules to utilize all modal information, which may not be able to capture different people in complex scenarios. In contrast, the method provided in the embodiments of the present disclosure can learn the similarity between people on the distribution graph by using the cyclic update strategy of the Multi-modality ClueFeature Graph (MMFG) and the Modality-Agnostic Distribution Graph (MADG), so as to capture people in complex scenarios. To obtain a better distribution, MAGNET uses MMFG to capture adjacent information within the modality and uses MADG to refine the distribution representation of different modalities. In addition, MAGNET improves the most in datasets with frequent scene switches and person distributions (for example, in TV drama 2: NMI +1.84%, in Movie 1: +8.39% NMI, and in TV drama 4: +1.97 NMI), which further illustrates the superiority of MAGNET when dealing with complex scenarios. In addition, MAGNET is compared with three multi-view clustering methods, namely the Autoencoder in Autoencoder Networks (AE 2- Nets), COMIC, and the Incomplete Multi-view Clustering via Contrastive Prediction (COMPLETER). In these methods, different modalities can be regarded as different views. By using the correlation between different views, the multi-view features are projected into a unified feature space for clustering. However, in the person clustering task, it is difficult to project face features, body features, and voice features into a unified space because the correlation between these features is very weak. The MAGNET provided in the embodiments of the present disclosure can obtain better results compared with the multi-view clustering methods because the multi-modal data is clustered using modality-agnostic distribution representations provided in the embodiments of the present disclosure.

[0256] Table 1 Clustering results of different clustering methods

[0257]

[0258] 2) Person clustering results with noise correlations

[0259] Person clustering relies on the correlation information given by different modalities. The correlation between the face and the body is usually determined by the intersection over union (IOU) between the face bounding box and the body bounding box, which may be invalid in crowded situations. In this case, when two people stand too close, the face of one person may be wrongly associated with the body of another person, which brings noise to the feature correlation degree between the two modal data of cross-modalities.

[0260] To demonstrate the robustness of MAGNET to noise correlations, misassociated bodies are simulated by randomly swapping features between tracks with a given probability ρ. ρ is denoted as the noise ratio because it can control the ratio of noise correlations after random swapping. Experimentally, ρ is set from 0.1 to 0.5. The results of NMI are shown in Table 2. Even when the noise ratio is 0.5, the NMI in Drama 5 only decreases by 4.1%. Since the distribution features in MADG aggregate information from all modalities, the clustering results on the distribution map are robust to data with noise correlations.

[0261] Table 2 Person clustering results with noise correlations

[0262] ρ Movie 1 TV Drama 2 Movie 3 TV Drama 4 TV Drama 5 0 78.30 61.84 76.56 85.07 93.11 0.1 78.09 61.62 75.62 84.09 91.97 0.2 77.55 61.39 74.89 83.51 91.67 0.3 77.42 61.58 74.76 82.78 90.31 0.4 77.02 60.82 73.99 81.76 89.40 0.5 76.49 60.61 73.95 81.14 88.82

[0263] 3) Effectiveness of MAGNET

[0264] To study the effectiveness of MMFG and MADG, MMFG and MADG are respectively removed from the original model MAGNET. MMFGonly is the model without the distribution map, and feature aggregation is performed through feature similarity. Modal data with the same modality are clustered separately in MMFG, and then modal data with different modalities are grouped according to their co-occurrence in the same object. For MADG only, MMFG is removed. Therefore, the correlation degree in MADG is fixed by the intra-modal feature similarity of the original features. At the same time, to verify the effectiveness of the multi-modal fusion module, by setting α in Equation (3-3) l= 0 to remove the module, denoted as MAGNET wo Fusion. The results are shown in Table 3. By comparing MMFG only and MAGNET, MADG improves MAGNET in NMI by 4% because it can obtain identity-based information from all modalities for clustering; however, MMFG can only obtain unimodal and pairwise similarities for clustering. Compared with MADG only, MAGNET's NMI is improved by 3.8%, which indicates that directly adopting the original features cannot well define the distribution correlation in MADG because MMFG can aggregate valuable modality-specific information in the feature space. Compared with MAGNET wo Fusion, the full model MAGNET improves WCP by 3.2% and increases NMI by 1.4%, demonstrating the effectiveness of the multimodal fusion module.

[0265] Table 3 Effectiveness of MAGNET

[0266] Method WCP CP CR CF NMI MMFG only 92.60 89.07 60.32 71.93 75.07 MADG only 92.33 91.90 62.80 74.61 75.28 MAGNET without Fusion 90.07 88.39 65.08 74.96 77.62 MAGNET 93.27 87.32 66.31 75.38 79.06

[0267] 4) Clustering effect of multimodality

[0268] As shown in Table 4, clustering using the modality data of all modalities has a better clustering effect than using the modality data of partial modalities. MAGNET_f, MAGNET_fb, and MAGNET_fv respectively represent person clustering using only the modality data of the face, the modality data of the face and body, and the modality data of the face and voice. By comparing MAGNET_fb, MAGNET_fv, and MAGNET_f, it can be seen that using the body or voice for clustering can bring improvements of 1.65% and 0.43% in NMI. Compared with MAGNET_fb and MAGNET_fv, MAGNET using all three types of modality data can improve person clustering by 2.07% and 3.29% respectively.

[0269] Table 4 Clustering effect of multimodality

[0270]

[0271]

[0272] 5) Setting of the initial value of η

[0273] Figure 6C It is a schematic diagram for initializing the distribution representation provided by an embodiment of the present disclosure. In formula (1-1), η is the value for initialization. Adjust η in VPCD and show the average NMI in Figure 6C When η is 0.5, the distribution representations of modality data from the same modality or different modalities will be initialized to the same value. In this case, due to the ignorance of the modal information in the dataset, the NMI will decrease. When η is high, the NMI will decrease because MAGNET loses the ability to tolerate noisy data. Therefore, η = 0.7 is the best choice for the experiment.

[0274] 6) Example of retrieval results using feature and distribution representations

[0275] Figure 6D FIG. is a schematic diagram of retrieval according to feature and distribution representations provided by an embodiment of the present disclosure. As Figure 6D shown, given a face image 630 as a query, the top 5 most similar person modal data with feature and distribution representations are obtained respectively. It shows that using distribution representation for retrieval is better than using feature retrieval in two aspects. First, person modal data with different modalities can be retrieved through distribution representation rather than feature retrieval. For example, in Figure 6D the second row, a body image of a given face image is retrieved using distribution representation. However, when using feature retrieval, the body image cannot be retrieved because the feature similarity of two modal data with different modalities cannot reflect the similarity that these two modal data belong to the same object. Second, using distribution representation for retrieval is more robust than using features for retrieval. For example, in Figure 6D the first row, there are two incorrect samples when using feature retrieval because the features are not obvious enough in the case of insufficient image light. However, using distribution representation for retrieval can avoid this problem because it fuses information from different modalities, making the retrieval of images with multiple modalities based on identity more robust.

[0276] 7) Comparison with multi-view clustering

[0277] Multi-view clustering is to cluster instances with multiple features from different views. The VoxCeleb2 dataset is used to evaluate the performance of multi-view clustering. The VoxCeleb2 dataset is split into a test set with 2048 objects and a disjoint training set with 3434 objects. In addition, 512 identities are sampled from 2048 identities to obtain a smaller test set. Several methods are used with the test protocol, including Kmeans, Spectral, AHC, ARO, and LGCN. The results are shown in Table 5. LGCN simply concatenates face and audio features as joint features and uses GCN to aggregate features for clustering. MAGNET provided by the embodiment of the present disclosure treats different modalities as one instance and adaptively uses intra-modal and cross-modal information to cluster multi-modal person cues in a modality-independent distribution space. The F-score of MAGENT provided by the embodiment of the present disclosure is improved by 6.6% and 8.2% on test sets with 512 and 2048 objects respectively.

[0278] Table 5 Multi-view clustering results

[0279]

[0280]

[0281] Based on the above embodiments, an embodiment of the present disclosure provides a clustering device. Figure 7 It is a schematic structural diagram of a clustering device provided by an embodiment of the present disclosure, as Figure 7 shown. The clustering device 70 includes a first acquisition module 71, a first determination module 72, and a first clustering module 73.

[0282] The first acquisition module 71 is configured to acquire a dataset to be processed, where the dataset to be processed includes at least two modal data, and the at least two modal data belong to at least one object.

[0283] The first determination module 72 is configured to determine a target distribution representation of each modal data, and the target distribution representation of each modal data is at least used to represent the probability that each modal data belongs to the same object as other modal data in the dataset to be processed.

[0284] The first clustering module 73 is configured to cluster each modal data based on each target distribution representation to obtain at least one clustering cluster, and each clustering cluster includes at least one modal data belonging to the same object.

[0285] In some embodiments, the first determination module 72 is further configured to: use a trained graph neural network model to update the distribution representation of each modal data at least once, and when the number of updates reaches a preset value, determine the updated distribution representation of each modal data as the target distribution representation of each modal data respectively.

[0286] In some embodiments, the graph neural network model includes a feature graph neural network, and the first determination module 72 is further configured to: use the trained feature graph neural network to update the distribution representation of each modal data based on the association representation of each modal data, and each association representation is used to represent the association information of each modal data.

[0287] In some embodiments, the first determination module 72 is further configured to: use a trained graph neural network model to update the distribution representation and the association representation of each modal data at least once, and each association representation is used to represent the association information of each modal data.

[0288] In some embodiments, the graph neural network model includes a feature graph neural network and a distribution graph neural network. The first determination module 72 is further configured to update the distribution representation of each modality data based on the correlation representation of each modality data by using the trained feature graph neural network; and update the correlation representation of each modality data based on the updated distribution representation of each modality data by using the trained distribution graph neural network.

[0289] In some embodiments, the first determination module 72 is further configured to: determine the current fusion distribution correlation degree between every two modality data based on the correlation representation of each modality data by using the trained feature graph neural network; and update the distribution representation of each modality data based on the current fusion distribution correlation degree between every two modality data.

[0290] In some embodiments, the dataset to be processed includes at least two sub-datasets, and each sub-dataset corresponds to one modality respectively. The first determination module 72 is further configured to: determine the first feature similarity between the correlation representations of every two modality data in the sub-datasets of different modalities; determine the second feature similarity between the correlation representations of every two modality data in the sub-datasets of the same modality; and determine the current fusion distribution correlation degree between every two modality data based on each first feature similarity and each second feature similarity.

[0291] In some embodiments, the first determination module 72 is further configured to: determine the feature correlation degree between every two modality data in the sub-datasets of different modalities; perform normalization processing on each second feature similarity; and determine the first feature similarity between the correlation representations of every two modality data based on each normalized second feature similarity and each feature correlation degree.

[0292] In some embodiments, the first determination module 72 is further configured to: determine the correlation relationship between every two modality data in the sub-datasets of different modalities; and determine each feature correlation degree based on each correlation relationship.

[0293] In some embodiments, the first determination module 72 is further configured to: determine the second feature similarity between the correlation representations of every two modality data based on the previous correlation representation of each modality data in the sub-datasets of the same modality.

[0294] In some embodiments, the first determination module 72 is further configured to: determine the current fusion feature correlation degree between every two modality data based on the updated distribution representation of each modality data by using the trained distribution graph neural network; and update the correlation representation of each modality data based on each current fusion feature correlation degree.

[0295] In some embodiments, the first clustering module 73 is further configured to: cluster each piece of modality data based on the similarity between every two of the target distribution representations, to obtain at least one clustering cluster.

[0296] Based on the above embodiments, an embodiment of the present disclosure provides a model training apparatus, Figure 8 which is a schematic structural diagram of the composition of a model training apparatus provided by an embodiment of the present disclosure. As Figure 8 shown, the model training apparatus 80 includes a second acquisition module 81, a second determination module 82, a third determination module 83, and a first update module 84.

[0297] The second acquisition module 81 is configured to acquire a sample set, where the sample set includes at least one data subset, each data subset includes at least two pieces of modality data, the at least two pieces of modality data belong to at least one object, and each piece of modality data has label information;

[0298] The second determination module 82 is configured to use a graph neural network model to be trained to determine the target distribution representation of each piece of modality data in each data subset, and the target distribution representation of each piece of modality data is at least used to represent the probability that each piece of modality data belongs to the same object as other modality data in the data subset;

[0299] The third determination module 83 is configured to determine a target loss value based on the target distribution representation and label value of each piece of modality data in each data subset;

[0300] The first update module 84 is configured to update the parameters of the graph neural network model when the target loss value meets a preset condition.

[0301] In some embodiments, the second determination module 82 is further configured to: use the graph neural network model to be trained to update the distribution representation of each piece of modality data at least once; when the number of updates reaches a preset value, respectively determine the updated distribution representation of each piece of modality data as the target distribution representation of each piece of modality data.

[0302] In some embodiments, the graph neural network model includes a feature graph neural network, and the second determination module 82 is further configured to: use the feature graph neural network to be trained to update the distribution representation of each piece of modality data based on the association representation of each piece of modality data, and each association representation is used to represent the association information of each piece of modality data.

[0303] In some embodiments, the graph neural network model further includes a distribution graph neural network, and the second determination module 82 is further configured to: use the to-be-trained distribution graph neural network to update the association representation of each modality data based on the updated distribution representation of each modality data.

[0304] In some embodiments, each of the data subsets includes at least two sub-datasets, and each of the sub-datasets corresponds to one modality respectively; the third determination module 83 is further configured to: determine a feature similarity loss value based on the second feature similarity between the association representations of every two modality data in the sub-datasets of the same modality in each update and the label information of each modality data; determine a distribution similarity loss value based on the distribution similarity between the distribution representations of every two modality data in each historical update, the distribution similarity between the target distribution representations of every two modality data, and the label information of each modality data; determine a loss value based on the feature similarity loss value and the distribution similarity loss value.

[0305] The description of the above device embodiments is similar to the description of the above method embodiments and has similar beneficial effects to those of the method embodiments. For the technical details not disclosed in the device embodiments of the present disclosure, please refer to the description of the method embodiments of the present disclosure for understanding.

[0306] It should be noted that in the embodiments of the present disclosure, if the above method is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present disclosure, in essence or the part that contributes to the related art, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing an electronic device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of the present disclosure. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes. In this way, the embodiments of the present disclosure are not limited to any specific combination of hardware and software.

[0307] The embodiments of the present disclosure provide an electronic device, including a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the computer program, the above method is implemented.

[0308] The embodiments of the present disclosure provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above method is implemented. The computer-readable storage medium can be transient or non-transient.

[0309] Embodiments of the present disclosure provide a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. The computer program product can be specifically implemented in a manner of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.

[0310] It should be noted that Figure 9 is a schematic diagram of a hardware entity of an electronic device in an embodiment of the present disclosure. As Figure 9 shown, the hardware entity of the electronic device 900 includes: a processor 901, a communication interface 902, and a memory 903, where:

[0311] The processor 901 generally controls the overall operation of the electronic device 900.

[0312] The communication interface 902 can enable the electronic device to communicate with other terminals or servers through a network.

[0313] The memory 903 is configured to store instructions and applications executable by the processor 901, and can also cache data to be processed or already processed by the processor 901 and each module in the electronic device 900 (for example, image data, audio data, voice communication data, and video communication data), and can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM). Data transmission can be performed between the processor 901, the communication interface 902, and the memory 903 through a bus 904.

[0314] It should be pointed out here that: The descriptions of the above storage medium and device embodiments are similar to those of the above method embodiments, and have beneficial effects similar to those of the method embodiments. For technical details not disclosed in the storage medium and device embodiments of the present disclosure, please refer to the descriptions of the method embodiments of the present disclosure for understanding.

[0315] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present disclosure. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics may be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present disclosure, the magnitudes of the serial numbers of the above processes do not mean the sequence of execution, and the execution sequence of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure. The serial numbers of the embodiments of the present disclosure above are only for description and do not represent the advantages or disadvantages of the embodiments.

[0316] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.

[0317] In several embodiments provided by the present disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling or communication connection between the components shown or discussed with each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0318] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units; they may be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0319] In addition, each functional unit in the embodiments of the present disclosure can be all integrated in a processing unit, or each unit can be separately a unit alone, or two or more units can be integrated in one unit; the above integrated unit can be implemented in the form of hardware, or in the form of a hardware plus a software functional unit.

[0320] Those of ordinary skill in the art will understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments; and the foregoing storage medium includes: removable storage devices, read-only memory (ROM), magnetic disks, or optical discs and other various media that can store program codes.

[0321] Alternatively, if the above-mentioned integrated unit is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present disclosure, in essence or the part that contributes to the related art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present disclosure. And the foregoing storage medium includes: removable storage devices, ROM, magnetic disks, or optical discs and other various media that can store program codes.

[0322] As described above, the above are only the implementation manners of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should all be covered by the protection scope of the present disclosure.

Claims

1. A clustering method, characterized in that, Including: Obtain a dataset to be processed, where the dataset to be processed includes at least two types of modal data, the at least two types of modal data belong to at least one object, and the modal data includes one of the following: modal data containing a face, modal data containing a human body, modal data containing an iris, modal data containing a voice, modal data containing a fingerprint; Determine the target distribution representation of each piece of modal data, and the target distribution representation of each piece of modal data is at least used to represent the probability that each piece of modal data belongs to the same object as other modal data in the dataset to be processed; Based on each target distribution representation, perform clustering on each piece of modal data to obtain at least one clustering cluster, and each clustering cluster includes at least one piece of modal data belonging to the same object.

2. The method according to claim 1, characterized in that, The determining the target distribution representation of each piece of modal data includes: Using a trained graph neural network model, update the distribution representation of each piece of modal data at least once, and when the number of updates reaches a preset value, respectively determine the updated distribution representation of each piece of modal data as the target distribution representation of each piece of modal data; wherein, the graph neural network model includes at least one of the following: a feature graph neural network, a distribution graph neural network; The feature graph neural network is used to construct a feature graph of each piece of modal data, the feature graph includes at least one feature node and a feature connection relationship between the at least one feature node, each feature node is used to represent the associated representation of each piece of modal data, the associated representation is used to represent the associated information of each piece of modal data, each feature connection relationship is used to represent the feature association relationship between every two pieces of modal data, and the feature association relationship is used to represent the probability that every two feature nodes within the same modality belong to the same object or the probability that every two feature nodes in different modalities belong to the same object; The distribution graph neural network is used to construct a distribution graph of each piece of modal data, the distribution graph includes at least one distribution node and a distribution connection relationship between the at least one distribution node, each distribution node is used to represent the distribution representation of each piece of modal data, and each distribution connection relationship is used to represent the probability that every two distribution nodes belong to the same object.

3. The method according to claim 2, wherein The graph neural network model includes a feature graph neural network; The using the trained graph neural network model to update the distribution representation of each piece of modal data at least once includes: Using the trained feature graph neural network, based on the associated representation of each piece of modal data, update the distribution representation of each piece of modal data, and each associated representation is used to represent the associated information of each piece of modal data.

4. The method according to claim 2, wherein The using the trained graph neural network model to update the distribution representation of each piece of modal data at least once includes: Using the trained graph neural network model, update the distribution representation and the associated representation of each piece of modal data at least once, and each associated representation is used to represent the associated information of each piece of modal data.

5. The method according to claim 4, wherein The graph neural network model includes a feature graph neural network and a distribution graph neural network. Using the trained graph neural network model to update the distribution representation and the correlation representation of each modality data at least once includes: Using the trained feature graph neural network to update the distribution representation of each modality data based on the correlation representation of each modality data; Using the trained distribution graph neural network to update the correlation representation of each modality data based on the updated distribution representation of each modality data.

6. The method according to claim 3, characterized in that, The step of using the trained feature graph neural network to update the distribution representation of each modality data based on the correlation representation of each modality data includes: Using the trained feature graph neural network to determine the current fusion distribution correlation degree between every two modality data based on the correlation representation of each modality data; Updating the distribution representation of each modality data based on the current fusion distribution correlation degree between every two modality data.

7. The method according to claim 6, wherein The dataset to be processed includes at least two sub-datasets, and each sub-dataset corresponds to one modality respectively; The step of using the trained feature graph neural network to determine the current fusion distribution correlation degree between every two modality data based on the correlation representation of each modality data includes: Determining the first feature similarity between the correlation representations of every two modality data in the sub-datasets of different modalities; Determining the second feature similarity between the correlation representations of every two modality data in the sub-datasets of the same modality; Determining the current fusion distribution correlation degree between every two modality data based on each first feature similarity and each second feature similarity.

8. The method according to claim 7, characterized in that, The step of determining the first feature similarity between the correlation representations of every two modality data in the sub-datasets of different modalities includes: Determining the feature correlation degree between every two modality data in the sub-datasets of different modalities; Performing normalization processing on each second feature similarity; Determining the first feature similarity between the correlation representations of every two modality data based on each normalized second feature similarity and each feature correlation degree.

9. The method according to claim 8, characterized in that The step of determining the feature correlation degree between every two modality data in the sub-datasets of different modalities includes: Determining the correlation relationship between every two modality data in the sub-datasets of different modalities; Determining each feature correlation degree based on each correlation relationship.

10. The method according to claim 7, characterized in that, The step of determining the second feature similarity between the correlation representations of every two modality data in the sub-datasets of the same modality includes: Determining the second feature similarity between the correlation representations of every two modality data based on the previous correlation representation of each modality data in the sub-datasets of the same modality.

11. The method according to claim 5, wherein The step of using the trained distribution graph neural network to update the correlation representation of each modality data based on the updated distribution representation of each modality data includes: Using the trained distribution graph neural network to determine the current fusion feature correlation degree between every two modality data based on the updated distribution representation of each modality data; Updating the correlation representation of each modality data based on each current fusion feature correlation degree.

12. The method according to any one of claims 1 to 11, characterized in that, Performing clustering on each piece of the modality data based on each of the target distribution representations to obtain at least one clustering cluster, including: Performing clustering on each piece of the modality data based on the similarity between every two of the target distribution representations to obtain at least one clustering cluster.

13. A model training method, characterized in that, The method includes: Obtaining a sample set, where the sample set includes at least one data subset, each data subset includes at least two types of modality data, the at least two types of modality data belong to at least one object, each piece of the modality data has label information, the sample set includes at least two types of modality data, and the modality data includes one of the following: modality data containing a face, modality data containing a human body, modality data containing an iris, modality data containing a voice, modality data containing a fingerprint; Using a graph neural network model to be trained to determine the target distribution representation of each piece of modality data in each data subset, where the target distribution representation of each piece of modality data is at least used to represent the probability that each piece of modality data belongs to the same object as other modality data in the data subset, and the graph neural network model includes at least one of the following: a feature graph neural network, a distribution graph neural network; the feature graph neural network is used to construct a feature graph of each piece of modality data, the feature graph includes at least one feature node and the feature connection relationship between the at least one feature node, each feature node is used to represent the associated representation of each piece of modality data, the associated representation is used to represent the associated information of each piece of modality data, each feature connection relationship is used to represent the feature association relationship between every two pieces of modality data, and the feature association relationship is used to represent the probability that every two feature nodes within the same modality belong to the same object or the probability that every two feature nodes in different modalities belong to the same object; the distribution graph neural network is used to construct a distribution graph of each piece of modality data, the distribution graph includes at least one distribution node and the distribution connection relationship between the at least one distribution node, each distribution node is used to represent the distribution representation of each piece of modality data, and each distribution connection relationship is used to represent the probability that every two distribution nodes belong to the same object; Determining a target loss value based on the target distribution representation of each piece of modality data in each data subset and the label information of each piece of modality data; Updating the parameters of the graph neural network model when the target loss value meets a preset condition.

14. The method according to claim 13, wherein The using the graph neural network model to be trained to determine the target distribution representation of each piece of modality data in each data subset includes: Using the graph neural network model to be trained to update the distribution representation of each piece of modality data at least once, and when the number of updates reaches a preset value, respectively determining the updated distribution representation of each piece of modality data as the target distribution representation of each piece of modality data.

15. The method according to claim 14, wherein The graph neural network model includes a feature graph neural network; The using the graph neural network model to be trained to update the distribution representation of each piece of modality data at least once includes: Using the feature map neural network to be trained, based on the associated representation of each modality data, update the distribution representation of each modality data, where each associated representation is used to represent the association information of each modality data.

16. The method according to claim 15, wherein The graph neural network model further includes a distribution graph neural network, and the method further includes: Using the distribution graph neural network to be trained, based on the updated distribution representation of each modality data, update the associated representation of each modality data.

17. The method according to claim 15 or 16, characterized in that, Each data subset includes at least two sub-datasets, and each sub-dataset corresponds to one modality respectively; Determining the target loss value based on the target distribution representation of each modality data and the label information of each modality data in each data subset includes: Based on the second feature similarity between the associated representations of every two modality data in the sub-datasets of the same modality in each update and the label information of each modality data, determining the feature similarity loss value; Based on the distribution similarity between the distribution representations of every two modality data in each historical update, the distribution similarity between the target distribution representations of every two modality data, and the label information of each modality data, determining the distribution similarity loss value; Based on the feature similarity loss value and the distribution similarity loss value, determining the target loss value.

18. A clustering device, characterized in that, The device includes: A first acquisition module, configured to acquire a dataset to be processed, where the dataset to be processed includes at least two types of modality data, the at least two types of modality data belong to at least one object, and the modality data includes one of the following: modality data including a face, modality data including a human body, modality data including an iris, modality data including sound, modality data including a fingerprint; A first determination module, configured to determine the target distribution representation of each modality data, where the target distribution representation of each modality data is at least used to represent the probability that each modality data belongs to the same object as other modality data in the dataset to be processed; A first clustering module, configured to cluster each modality data based on each target distribution representation to obtain at least one clustering cluster, and each clustering cluster includes at least one modality data belonging to the same object.

19. A model training device, characterized in that, The device includes: A second acquisition module, configured to acquire a sample set, the sample set includes at least one data subset, each data subset includes at least two types of modality data, the at least two types of modality data belong to at least one object, each modality data has label information, and the modality data includes one of the following: modality data including a face, modality data including a human body, modality data including an iris, modality data including sound, modality data including a fingerprint; A second determination module, configured to use a graph neural network model to be trained to determine a target distribution representation of each modality data in each of the data subsets. The target distribution representation of each modality data is at least used to represent the probability that each modality data and other modality data in the data subset belong to the same object. The graph neural network model includes at least one of the following: a feature graph neural network, a distribution graph neural network; the feature graph neural network is configured to construct a feature graph of each modality data, the feature graph includes at least one feature node and a feature connection relationship between the at least one feature node, each feature node is configured to represent an association representation of each modality data, the association representation is used to represent the association information of each modality data, each feature connection relationship is used to represent a feature association relationship between every two modality data, and the feature association relationship is used to represent the probability that every two feature nodes within the same modality belong to the same object or the probability that every two feature nodes in different modalities belong to the same object; the distribution graph neural network is configured to construct a distribution graph of each modality data, the distribution graph includes at least one distribution node and a distribution connection relationship between the at least one distribution node, each distribution node is configured to represent a distribution representation of each modality data, and each distribution connection relationship is used to represent the probability that every two distribution nodes belong to the same object; A third determination module, configured to determine a target loss value based on the target distribution representation and the label value of each modality data in each of the data subsets; A first update module, configured to update the parameters of the graph neural network model when the loss value meets a preset condition.

20. An electronic device, comprising a processor and a memory, the memory storing a computer program that can run on the processor, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 17 is implemented.

21. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by a processor, the method according to any one of claims 1 to 17 is implemented.