Clustering method, clustering model training method and related device
Through the hierarchical clustering method of cascading codebook comparison module and difference feature operation, the multimodal clustering algorithm's calculation resource occupation under massive data and poor clustering effect of long-tail data is solved, and efficient and accurate multimodal representation information aggregation is achieved.
Patent Information
- Application Number
- CN202411060484.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-02
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-08-02
AI Technical Summary
The existing multimodal clustering algorithm occupies a large amount of computing resources in scenarios with diversified massive data and sample distribution, and the clustering effect of long-tail data is poor, which affects the generalization ability.
The codebook comparison module of N cascades is used to cluster the multimodal representation information output from the multimodal pre-trained model in layered clustering, and the difference characteristics are used to perform cascade comparison to determine the invisible classification label.
It improves clustering accuracy and generalization ability, can efficiently aggregate sparse samples, and improves the purity and distinction of the clustering method.
Smart Images

Figure CN118964996B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a clustering method, a clustering model training method, and related devices. Background Art
[0002] In the scenario of general content understanding, the evolution of multi-modal pre-training technology is getting faster and faster. The multi-modal representation information output by mainstream training paradigms such as the multi-modal text-image pre-training model (Contrastive Language-Image Pre-training, CLIP) is becoming increasingly dense. The cosine distance between multi-modal features is usually used to characterize the similarity between multiple samples. Therefore, due to the needs of business construction, the industry has performed content clustering on a large amount of business data based on feature cosine similarity. Typical methods include traditional clustering algorithms such as K-means and Density-Based Spatial Clustering of Applications with Noise (DBSCAN). These methods can achieve very generalizable effects in scenarios with a small amount of data and uniform sample distribution. However, with the evolution of a large amount of data and the diversification of sample distribution, the disadvantages of such clustering algorithms are exposed. The one-time clustering solution for all data occupies a very large amount of CPU computing resources and requires pre-layered parallelization operations. At the same time, a considerable part of the long-tail data will be clustered into the same category with a coarser granularity, and the clustering effect of long-tail samples has low discrimination, which is not conducive to the generalization of the clustering algorithm. Therefore, how to perform efficient clustering under the premise of a large amount of data and diversified sample distribution has become a current research hotspot. Summary of the Invention
[0003] In view of this, embodiments of the present disclosure provide a clustering method, a clustering model training method, and related devices, which can achieve very good high-purity aggregation even for a very sparse class of samples.
[0004] The clustering method described in the embodiments of the present disclosure is applied to a clustering model. The clustering model includes N cascaded codebook comparison modules, and each level of codebook comparison module corresponds to a codebook in the N-level codebook space, where N is an integer greater than or equal to 1. The clustering method includes: inputting the multi-modal representation information output by the multi-modal pre-training model into the clustering model; in each codebook comparison module of the clustering model, respectively determining the codeword in its corresponding codebook that is closest to the input feature, using it as the codeword corresponding to the codebook comparison module, and using the difference feature between the input feature and the codeword as the input feature of the next-level codebook comparison module; and determining the hidden classification label corresponding to the multi-modal representation information based on the codewords corresponding to the N cascaded codebook comparison modules.
[0005] In an embodiment of the present disclosure, in each codebook comparison module of the clustering model, the codeword in the corresponding codebook that is closest to the input feature is respectively determined as the codeword corresponding to the codebook comparison module, and the difference feature between the input feature and the codeword is used as the input feature of the next-level codebook comparison module, including: in the first-level codebook comparison module, the input multi-modal representation information is compared with multiple codewords in the first-level codebook corresponding to the first-level codebook comparison module, the codeword closest to the multi-modal representation information is determined as the codeword corresponding to the first-level codebook comparison module, a difference operation is performed on the multi-modal representation information and the first-level codeword to obtain a difference feature, and the difference feature is input into the second-level codebook comparison module; in the i-th codebook comparison module, where i is greater than 1 and less than N, the input difference feature is compared with multiple codewords in the i-th codebook corresponding to the i-th codebook comparison module, the codeword closest to the difference feature is determined as the codeword corresponding to the i-th codebook comparison module, a difference operation is performed on the difference feature and the i-th codeword to obtain a difference feature, and the difference feature is input into the next-level codebook comparison module; and in the N-th codebook comparison module, the input difference feature is compared with multiple codewords in the N-th codebook corresponding to the N-th codebook comparison module, and the codeword closest to the difference feature is determined as the codeword corresponding to the N-th codebook comparison module.
[0006] In an embodiment of the present disclosure, the above difference operation includes: determining the difference feature based on the following expression: D j =C j -CW j ; where j represents the j-th codebook comparison module; D j represents the difference feature determined by the j-th codebook comparison module; C j represents the input feature of the j-th codebook comparison module; and CW j represents the codeword corresponding to the j-th codebook comparison module.
[0007] In an embodiment of the present disclosure, determining the invisible classification label corresponding to the multi-modal representation information based on the codewords corresponding to the N cascaded codebook comparison modules includes: determining the invisible classification label corresponding to the multi-modal representation information based on the following expression: where T represents the invisible classification label corresponding to the multi-modal representation information; j represents the j-th codebook comparison module; represents the serial number of the codeword CW j corresponding to the j-th codebook comparison module in the j-th codebook; and M represents the number of codewords in each codebook.
[0008] In an embodiment of the present disclosure, each of the above N cascaded codebooks includes M codewords, where M is an integer greater than 1; determining the invisible classification label corresponding to the multimodal representation information based on the codewords corresponding to the N cascaded codebook comparison modules includes: respectively selecting one codeword from each of the N cascaded codebooks for combination, obtaining a total of M*N codeword combinations; encoding the M*N codeword combinations to obtain the identifiers of the above M*N codeword combinations; determining the codeword combination corresponding to the multimodal representation information based on the codewords corresponding to the N cascaded codebook comparison modules; and determining the identifier of the codeword combination corresponding to the multimodal representation information as the invisible classification label corresponding to the multimodal representation information.
[0009] An embodiment of the present disclosure also discloses a clustering model training method, which is applied to a clustering model, where the clustering model includes N cascaded codebook comparison modules, and N is an integer greater than or equal to 1; wherein, the clustering model training method includes: pre-constructing an N-layer codebook space; wherein, the N-layer codebook space corresponds one-to-one with the N cascaded codebook comparison modules; for each layer of the codebook space, initializing M cluster centers; where M is an integer greater than 1; during the training process, for each codebook comparison module of the clustering model, respectively execute: determining the cluster center closest to the input sample feature among the M cluster centers, and using it as the candidate cluster center corresponding to the codebook comparison module, and recording the distance between the input sample feature and the candidate cluster center as the candidate distance corresponding to the codebook comparison module; and using the difference feature between the input sample feature and the candidate cluster center as the input feature of the next-level codebook comparison module; summing the candidate distances corresponding to the N cascaded codebook comparison modules to obtain a similarity loss; and optimizing the M cluster centers of the N cascaded codebook comparison modules with the minimum similarity loss as the first optimization goal to obtain the optimized M cluster centers; performing codebook restoration and feature reconstruction based on the candidate cluster centers corresponding to the N cascaded codebook comparison modules to obtain reconstructed features; using the difference between the input sample feature and the reconstructed feature as a reconstruction loss; and optimizing the M cluster centers of the N cascaded codebook comparison modules with the minimum reconstruction loss as the second optimization goal to obtain the optimized M cluster centers; and when the training end condition is satisfied, respectively using the M cluster centers corresponding to the N cascaded codebook comparison modules as the M codewords of the codebooks corresponding to the N cascaded codebook comparison modules.
[0010] Embodiments of the present disclosure also disclose a clustering model, including: N cascaded codebook comparison modules; where N is an integer greater than or equal to 1; and each level of the N cascaded codebook comparison modules corresponds to a codebook in the N-layer codebook space, and is used to compare the input features with each codeword in its corresponding codebook, determine the codeword with the closest distance to the input features, use it as the codeword corresponding to itself, and use the difference feature between the input features and the codeword as the input feature of the next-level codebook comparison module.
[0011] Embodiments of the present disclosure also disclose a clustering device, including: an input interface, the clustering model as claimed in claim 6, and a classification label determination module; where the input interface is used to receive the multimodal representation information output by the multimodal pre-training model, and input the received multimodal representation information into the clustering model; each codebook comparison module of the clustering model respectively determines the codeword with the closest distance to the input features in its corresponding codebook, uses it as the codeword corresponding to the codebook comparison module, and uses the difference feature between the input features and the codeword as the input feature of the next-level codebook comparison module; and the classification label determination module is used to determine the latent classification label corresponding to the multimodal representation information based on the codewords corresponding to the N cascaded codebook comparison modules.
[0012] In addition, embodiments of the present disclosure also provide an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the above clustering method is implemented.
[0013] Embodiments of the present disclosure also provide a non-transitory computer-readable storage medium, where the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the above clustering method.
[0014] Embodiments of the present disclosure also provide a computer program product, including computer program instructions, and when the computer program instructions run on a computer, the computer is caused to execute the above clustering method.
[0015] The clustering method, clustering model training method, and related devices described in the embodiments of the present disclosure propose a discretized clustering scheme based on deep learning, and through introducing a residual model structure, hierarchical clustering is performed on the attributes of different dimensions of the input information-intensive multimodal representation information, so as to achieve the goal of fully mapping the dense multimodal representation information into a feature space with high discrimination and distinct content levels. The above scheme can achieve very good high-purity aggregation even for a class with a very sparse sample size, thereby greatly improving the accuracy and generalization ability of the clustering method. Description of the Drawings
[0016] To more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the embodiments or the description of related technologies. Obviously, the drawings in the following description are only embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0017] Figure 1 It shows the structure of the clustering model according to some embodiments of the present disclosure.
[0018] Figure 2 It shows the implementation process of the clustering method according to some embodiments of the present disclosure.
[0019] Figure 3 It shows the implementation process of the clustering model training method according to some embodiments of the present disclosure.
[0020] Figure 4 It shows the internal structure of the clustering device according to some embodiments of the present disclosure.
[0021] Figure 5 It shows a more specific schematic diagram of the hardware structure of an electronic device according to some embodiments of the present disclosure. Detailed Embodiments
[0022] To make the purpose, technical solutions, and advantages of the present disclosure clearer and more understandable, the following further details the present disclosure in combination with specific embodiments and with reference to the drawings.
[0023] It should be noted that unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the ordinary meaning understood by those of ordinary skill in the art to which the present disclosure pertains. The "first", "second", and similar terms used in the embodiments of the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. The terms such as "include" or "comprise" mean that the elements or objects appearing before this term cover the elements or objects listed after this term and their equivalents, without excluding other elements or objects. The terms such as "connect" or "be connected" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0024] It can be understood that before using the technical solutions of the various embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner and the user's authorization will be obtained.
[0025] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operation of the disclosed technical solution based on the prompt message.
[0026] As an optional but non-limiting implementation, in response to a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0027] It is understandable that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0028] As mentioned earlier, current clustering methods based on multimodal pre-training can achieve very generalized results in scenarios with relatively small amounts of data and uniform sample distribution. However, with the evolution of massive data and the diversification of sample distribution, one-time clustering of all data consumes a significant amount of CPU resources and requires parallel operations such as pre-stratification. Furthermore, a considerable portion of long-tail data is clustered into the same category at a coarse granularity, resulting in a less differentiated clustering effect for long-tail samples.
[0029] Specifically, in the task of content understanding, various modal carrier information that can be collected is usually utilized in the modeling stage, including video, voice, optical recognition (OCR), comments, etc., and a very rich and dense feature space can be represented in the pre-training stage. However, the output high-dimensional features are not sufficient to directly characterize the aggregation of massive samples at the content level. The subsequent optimization process usually introduces a clustering algorithm, performs intensive similarity calculations on high-dimensional features, and constructs circle relationships between data. For massive samples with highly similar content, clustering algorithms can usually achieve very good generalization effects, but for very sparse samples, they are usually clustered together with data with large content differences, affecting the downstream data confidence.
[0030] In order to solve the above technical problems, the embodiments of the present disclosure provide a clustering model that can perform clustering based on the information-dense multimodal representation information output by the multimodal pre-training model, and output invisible classification labels corresponding to the input multimodal representation information.
[0031] Figure 1 The structure of the clustering model 100 described in some embodiments of the present disclosure is shown.Figure 1 As shown, the clustering model 100 described in the embodiments of the present disclosure includes: N cascaded codebook comparison modules; where N is an integer greater than or equal to 1.
[0032] In the embodiments of the present disclosure, the above-mentioned N cascaded codebook comparison modules have substantially the same structure. Each level of the codebook comparison module among the above-mentioned N cascaded codebook comparison modules corresponds to one codebook in the N-layer codebook space. The N codebooks corresponding to the above-mentioned N cascaded codebook comparison modules constitute the N-layer codebook space. Among them, each layer of the codebook space depicts the attributes of the samples in one dimension. In this way, the N-layer codebook space can perform hierarchical clustering on the input multi-modal representation information from the attributes of N dimensions, so as to achieve the goal of fully mapping the dense multi-modal representation information to a feature space with high discrimination and distinct content levels, thereby improving the accuracy and generalization ability of the clustering method.
[0033] In the embodiments of the present disclosure, each level of the codebook comparison module among the above-mentioned N cascaded codebook comparison modules can be used to compare its own input feature Z with each codeword in its corresponding codebook, determine the codeword with the closest distance to the input feature, output it as the codeword corresponding to itself, and use the difference feature between the input feature and the codeword as the input feature of the next-level codebook comparison module.
[0034] Specifically, for the first-level codebook comparison module, the first-level codebook comparison module compares the input multi-modal representation information with multiple codewords in the first-level codebook corresponding to the first-level codebook comparison module, determines the codeword with the closest distance to the multi-modal representation information as the codeword corresponding to the first-level codebook comparison module, performs a difference operation on the multi-modal representation information and the first-level codeword to obtain a difference feature, and inputs the difference feature into the second-level codebook comparison module.
[0035] For the i-th level codebook comparison module, where i is greater than 1 and less than N, the i-th level codebook comparison module compares the input difference feature with multiple codewords in the i-th level codebook corresponding to the i-th level codebook comparison module, determines the codeword with the closest distance to the difference feature as the codeword corresponding to the i-th level codebook comparison module, performs a difference operation on the difference feature and the i-th level codeword to obtain a difference feature, and inputs the difference feature into the next-level codebook comparison module.
[0036] For the N-th level codebook comparison module, the N-th level codebook comparison module compares the input difference feature with multiple codewords in the N-th level codebook corresponding to the N-th level codebook comparison module, and determines the codeword with the closest distance to the difference feature as the codeword corresponding to the N-th level codebook comparison module.
[0037] For example, Figure 1The first-level codebook comparison module, the second-level codebook comparison module, ……, and the N-level codebook comparison module are shown. Among them, in the first-level codebook comparison module, the input feature Z is closest to codeword 1. Therefore, the first-level codebook comparison module will output codeword 1, and calculate the difference feature between the input feature Z and codeword 1 as the input of the second-level codebook comparison module. In the second-level codebook comparison module, the input feature Z is closest to codeword 2. Therefore, the second-level codebook comparison module will output codeword 2, and calculate the difference feature between the input feature Z and codeword 2 as the input of the third-level codebook comparison module. ……. In the N-level codebook comparison module, the input feature Z is closest to codeword 5. Therefore, the N-level codebook comparison module will output codeword 5. It should be specifically noted that the N-level codebook comparison module is the last-level module. Therefore, there is no need to output the difference feature to the next-level module.
[0038] It should be noted that, in the embodiments of the present disclosure, the above distance may refer to any one of various similarity distances, for example, Euclidean distance, cosine similarity distance, etc. The embodiments of the present disclosure do not limit the specific meaning and algorithm of the above distance.
[0039] In the embodiments of the present disclosure, the above difference operation may include: determining the difference feature based on the following expression:
[0040] D j =C j -CW j
[0041] where j represents the j-level codebook comparison module; D j represents the difference feature determined by the j-level codebook comparison module; C j represents the input feature of the j-level codebook comparison module; and CW j represents the codeword corresponding to the j-level codebook comparison module.
[0042] Based on the clustering model 100 shown above Figure 1 the embodiments of the present disclosure also propose a clustering method. Figure 2 shows the implementation process of the above clustering method. As Figure 2 shown, the above clustering method may include the following steps:
[0043] In step 210, input the multi-modal representation information output by the multi-modal pre-training model into the above clustering model.
[0044] In an embodiment of the present disclosure, the above-mentioned multimodal pre-training model may be a multimodal pre-training model such as the CLIP model. The embodiment of the present disclosure does not limit the specific type of the multimodal pre-training model used. As mentioned above, in the scenario of general content understanding, with the evolution of the multimodal pre-training model, the multimodal representation information output by it is becoming more and more dense. The clustering method proposed in the embodiment of the present disclosure is based on the dense multimodal representation information output by the above-mentioned multimodal pre-training model for clustering.
[0045] In addition, the above-mentioned multimodal representation information may be the fusion feature of multiple models. Further, if the back-end service also needs to rely on other statistical features, the statistical values required can also be encoded through the encoding algorithm of numerical features, and the features obtained after encoding are further fused into the above-mentioned multimodal representation information as the input of the above-mentioned clustering model.
[0046] In step 220, in each codebook comparison module of the above-mentioned clustering model, the codeword in its corresponding codebook that is closest to the input feature is respectively determined as the codeword corresponding to the codebook comparison module, and the difference feature between the above-mentioned input feature and the above-mentioned codeword is used as the input feature of the next-level codebook comparison module.
[0047] In an embodiment of the present disclosure, the above-mentioned step 220 may specifically include the following steps:
[0048] First, in the first-level codebook comparison module, the input multimodal representation information is compared with multiple codewords in the first-level codebook corresponding to the first-level codebook comparison module, the codeword closest to the multimodal representation information is determined as the codeword corresponding to the first-level codebook comparison module, a difference operation is performed between the multimodal representation information and the first-level codeword to obtain a difference feature, and the difference feature is input into the second-level codebook comparison module;
[0049] Secondly, in the i-th level codebook comparison module, where i is greater than 1 and less than N, the input difference feature is compared with multiple codewords in the i-th level codebook corresponding to the i-th level codebook comparison module, the codeword closest to the difference feature is determined as the codeword corresponding to the i-th level codebook comparison module, a difference operation is performed between the difference feature and the i-th level codeword to obtain a difference feature, and the difference feature is input into the next-level codebook comparison module; and
[0050] Finally, in the N-th level codebook comparison module, the input difference feature is compared with multiple codewords in the N-th level codebook corresponding to the N-th level codebook comparison module, and the codeword closest to the difference feature is determined as the codeword corresponding to the N-th level codebook comparison module.
[0051] In an embodiment of the present disclosure, the above-mentioned difference operation may include: determining the difference feature based on the following expression:
[0052] D j = C j - CW j
[0053] where j represents the j-th codebook comparison module; D j represents the difference feature determined by the j-th codebook comparison module; C j represents the input feature of the j-th codebook comparison module; and CW j represents the codeword corresponding to the j-th codebook comparison module.
[0054] In step 230, determine the invisible classification label corresponding to the above-mentioned multimodal representation information based on the codewords corresponding to the above-mentioned N cascaded codebook comparison modules.
[0055] In an embodiment of the present disclosure, each of the above-mentioned N cascaded codebooks contains M codewords, where M is an integer greater than 1. In this case, the above step 230 may include the following multiple steps:
[0056] First, select one codeword from each of the above-mentioned N cascaded codebooks for combination, and a total of M * N different codeword combinations are obtained;
[0057] Second, encode the above M * N codeword combinations to obtain the identifiers (IDs) of the above M * N codeword combinations;
[0058] Then, determine the codeword combination corresponding to the above-mentioned multimodal representation information based on the codewords corresponding to the above-mentioned N cascaded codebook comparison modules; and
[0059] Finally, determine the identifier of the codeword combination corresponding to the above-mentioned multimodal representation information as the invisible classification label corresponding to the multimodal representation information.
[0060] As described above, the N codebooks corresponding to the above-mentioned N cascaded codebook comparison modules constitute an N-layer codebook space. Among them, each layer of the codebook space characterizes the attributes of the samples in one dimension. That is to say, by constructing the above-mentioned N-layer codebook space, the above-mentioned multimodal representation information can be classified separately in N dimensions. In addition, each of the above-mentioned N cascaded codebooks contains M codewords. Among them, the M codewords in one of the above-mentioned codebooks represent that in its corresponding dimension, the above-mentioned multimodal representation information can be specifically divided into one of M categories. Through such a setting, the final classification result will be to specifically divide the multimodal representation information into one of M*N categories from N dimensions. Therefore, after encoding the combination of the above-mentioned M*N codewords and obtaining the ID of the combination of the above-mentioned M*N codewords, once the ID of the codeword combination corresponding to the above-mentioned multimodal representation information is determined, it can be determined which specific category among the M*N categories the above-mentioned multimodal representation information is divided into. Since what is output in the above method is the ID of the codeword combination, rather than an explicit classification label, it is called an implicit classification label. The above-mentioned implicit classification label can provide a reference for the degree of divisibility of the samples at the content level for various downstream tasks.
[0061] In some embodiments of the present disclosure, a specific encoding method is given. In these embodiments, the implicit classification label corresponding to the above-mentioned multimodal representation information can be directly determined based on the following expression:
[0062]
[0063] where T represents the implicit classification label corresponding to the above-mentioned multimodal representation information; j represents the j-th level codebook comparison module; represents the serial number of the codeword corresponding to the j-th level codebook comparison module in the j-th level codebook; and M represents the number of codewords in each codebook.
[0064] It can be seen that through the above method, the implicit classification label can be directly determined by using the serial numbers of the codewords corresponding to each level of codebook comparison module in their codebooks.
[0065] The above clustering model and clustering method make full use of the quantization mapping of the multi-dimensional feature space to fully map the dense multimodal representation information into a feature space with high discrimination and distinct content levels, and output implicit classification labels. It can be understood that by constructing a multi-layer codebook space for similarity quantization of feature vectors, it can be ensured that the samples aggregated in each layer of the codebook space have as high a similarity as possible in different content attributes, thereby effectively reducing the impact of the sample distribution on the clustering effect. Moreover, as the amount of feature information extracted by pre-training becomes more and more dense, this discrete clustering method can be unaffected by the sample size, and the feature aggregation between different samples becomes finer and finer.
[0066] Based on the above clustering model, an embodiment of the present disclosure also provides a method for training a clustering model. Figure 3 shows the implementation process of the clustering model training method described in the embodiment of the present disclosure. As Figure 3 shown, the method includes the following steps:
[0067] In step 310, an N-layer codebook space is pre-constructed.
[0068] In the embodiment of the present disclosure, the above N-layer codebook space corresponds one-to-one with the N cascaded codebook comparison modules. Among them, each layer of the codebook space can correspond to the attributes of the samples in one dimension.
[0069] In step 320, for each layer of the codebook space, M cluster centers are initialized; where M is an integer greater than 1.
[0070] It can be understood that in the embodiment of the present disclosure, the meaning of the cluster center of each codebook above is similar to the center of each classification. However, before the model training, the above initialized cluster centers are inaccurate, and the goal of the model training is to optimize and obtain suitable cluster centers, that is, to achieve the optimal clustering result of each layer of the codebook.
[0071] In step 330, during the training process, for each codebook comparison module of the above clustering model, the following steps are respectively executed:
[0072] Determine the cluster center closest to the input sample feature among the M cluster centers;
[0073] Take it as the candidate cluster center corresponding to the above codebook comparison module;
[0074] Record the distance between the input sample feature and the candidate cluster center as the candidate distance corresponding to the above codebook comparison module; and
[0075] Take the difference feature between the input sample feature and the candidate cluster center as the input feature of the next-level codebook comparison module.
[0076] In the embodiment of the present disclosure, the above distance can refer to any one of various similarity distances, for example, Euclidean distance, cosine similarity distance, etc.
[0077] In step 340, sum the candidate distances corresponding to the above N cascaded codebook comparison modules to obtain a similarity loss. Then, optimize the M cluster centers of the N cascaded codebook comparison modules with the minimum similarity loss as the first optimization goal to obtain the optimized M cluster centers.
[0078] In an embodiment of the present disclosure, by summing the candidate distances corresponding to the above-mentioned N cascaded codebook comparison modules, this is used as an optimization item to be added to the training, and it is expected that this optimization item will be constrained to become smaller during the subsequent training process, so as to achieve the effect that the distance between similar samples is smaller and the distance between dissimilar samples is larger.
[0079] In step 350, based on the candidate cluster centers corresponding to the above-mentioned N cascaded codebook comparison modules, codebook restoration and feature reconstruction are performed to obtain reconstructed features. The difference between the input sample features and the reconstructed features is used as the reconstruction loss, and the M cluster centers of the N cascaded codebook comparison modules are optimized with the minimum of the reconstruction loss as the second optimization target to obtain the optimized M cluster centers.
[0080] It should be noted that the above-mentioned codebook restoration and feature reconstruction can be implemented by using existing codebook restoration and feature reconstruction methods, and the embodiments of the present disclosure do not limit the specific implementation manners of the above steps.
[0081] It can be seen that the above-mentioned clustering model training method further introduces a feature mapping module after the above-mentioned N cascaded codebook comparison modules to highly restore the input features, and the distance metric between the input and output features is used as an optimization item to be added to the training, with the expectation of retaining as much information as possible during the feature compression process.
[0082] In step 360, when the training end condition is satisfied, the M cluster centers corresponding to the above-mentioned N cascaded codebook comparison modules are respectively used as the M codewords of the codebook corresponding to the above-mentioned N cascaded codebook comparison modules.
[0083] In an embodiment of the present disclosure, the above-mentioned training end condition can be one or more pre-set conditions. For example, the number of iterations exceeds a pre-set threshold, the similarity loss and / or the reconstruction loss meet the set conditions, and so on.
[0084] Through the above-mentioned clustering model training method, suitable M cluster centers can be determined in the above-mentioned N-layer codebook spaces respectively, that is, the M codewords of each layer of codebook space are determined respectively. Thus, during the process of model inference, by sequentially determining the codewords with the closest distance to the input features in each level of codebook comparison module, the classification of multi-modal representation information in N dimensions is completed in sequence, so as to achieve the goal of fully mapping the dense multi-modal representation information into a feature space with higher discrimination and distinct content levels.
[0085] Embodiments of the present disclosure also disclose a clustering device, the internal structure of which is as Figure 4 shown, including: an input interface 410, the above-mentioned clustering model 420, and a classification label determination module 430.
[0086] Among them, the above input interface is used to receive the multi-modal representation information output by the multi-modal pre-training model and input the received multi-modal representation information into the above clustering model 420.
[0087] For each level of codebook comparison module of the above clustering model 420, it respectively determines the codeword in its corresponding codebook that is closest to the input feature, takes it as the codeword corresponding to the codebook comparison module, and uses the difference feature between the input feature and the codeword as the input feature of the next-level codebook comparison module.
[0088] The above classification label determination module 430 is used to determine the latent classification label corresponding to the above multi-modal representation information based on the codewords corresponding to each level of codebook comparison module.
[0089] The clustering method, clustering model training method, and related devices described in the embodiments of the present disclosure propose a deep learning-based discretized clustering scheme. By introducing a residual model structure, hierarchical clustering is performed on the attributes of different dimensions of the input information-dense multi-modal representation information, so as to achieve the goal of fully mapping the dense multi-modal representation information into a feature space with high discrimination and distinct content levels. Even for a category with a very sparse sample size, the above scheme can achieve very good high-purity aggregation, thereby greatly improving the accuracy and generalization ability of the clustering method.
[0090] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the clustering method described in any one of the above embodiments.
[0091] Figure 5 FIG. shows a more specific hardware structure diagram of the electronic device provided in this embodiment. The device may include: a processor 2010, a memory 2020, an input / output interface 2030, a communication interface 2040, and a bus 2050. Among them, the processor 2010, the memory 2020, the input / output interface 2030, and the communication interface 2040 are communicatively connected to each other inside the device through the bus 2050.
[0092] The processor 2010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present specification.
[0093] The memory 2020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 2020 can store the operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 2020 and called and executed by the processor 2010.
[0094] The input / output interface 2030 is used to connect to input / output devices to achieve information input and output. Among them, the input / output devices can be configured as components in the device or externally connected to the device to provide corresponding functions. The input devices can include microphones, various sensors, etc., and the output devices can include displays, speakers, vibrators, indicator lights, etc.
[0095] The communication interface 2040 is used to connect to a communication module (not shown in the figure) to achieve communication interaction between this device and other devices. The communication module can communicate through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0096] The bus 2050 includes a path for transmitting information between various components of the device (such as the processor 2010, the memory 2020, the input / output interface 2030, and the communication interface 2040).
[0097] It should be noted that although the above device only shows the processor 2010, the memory 2020, the input / output interface 2030, the communication interface 2040, and the bus 2050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solutions of the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0098] The electronic device in the above embodiment is used to implement the corresponding clustering method in any of the foregoing embodiments and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0099] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the clustering method as described in any of the foregoing embodiments.
[0100] The computer-readable medium of this embodiment includes both permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0101] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the task processing method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0102] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the concept of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of brevity.
[0103] In addition, for simplicity of explanation and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the devices may be shown in block diagram form to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be fully within the understanding of those skilled in the art). In the case where specific details (such as circuits) are set forth to describe the exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0104] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0105] Embodiments of the present disclosure are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Accordingly, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the embodiments of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A clustering method, applied to a clustering model, wherein, The clustering model includes N cascaded codebook comparison modules, and each level of the codebook comparison module corresponds to a codebook in the N-level codebook space, where N is an integer greater than or equal to 1; the clustering method includes: Input the multi-modal representation information output by the multi-modal pre-training model based on the modal carrier information into the clustering model; wherein, the modal carrier information includes: one or a combination of image, text, audio, and video; In each codebook comparison module of the clustering model, respectively determine the codeword in its corresponding codebook that is closest to the input feature, use it as the codeword corresponding to the codebook comparison module, and use the difference feature between the input feature and the codeword as the input feature of the next-level codebook comparison module; and Determine the latent classification label corresponding to the multi-modal representation information based on the codewords corresponding to the N cascaded codebook comparison modules; wherein, Determining the latent classification label corresponding to the multi-modal representation information based on the codewords corresponding to the N cascaded codebook comparison modules includes: Determine the latent classification label corresponding to the multi-modal representation information based on the following expression: Wherein, T represents the latent classification label corresponding to the multimodal representation information; j represents the j-th codebook comparison module; represents the codeword CW corresponding to the j-th codebook comparison module j the serial number in the j-th codebook; and M represents the number of codewords in each codebook.
2. The clustering method according to claim 1, wherein In each codebook comparison module of the clustering model, respectively determine the codeword in its corresponding codebook that is closest to the input feature, use it as the codeword corresponding to the codebook comparison module, and use the difference feature between the input feature and the codeword as the input feature of the next-level codebook comparison module includes: In the first-level codebook comparison module, compare the input multi-modal representation information with multiple codewords in the first-level codebook corresponding to the first-level codebook comparison module, determine the codeword closest to the multi-modal representation information as the first-level codeword corresponding to the first-level codebook comparison module, perform a difference operation on the multi-modal representation information and the first-level codeword to obtain a difference feature, and input the difference feature into the second-level codebook comparison module; In the i-th level codebook comparison module, where i is greater than 1 and less than N, compare the input difference feature with multiple codewords in the i-th level codebook corresponding to the i-th level codebook comparison module, determine the codeword closest to the difference feature as the i-th level codeword corresponding to the i-th level codebook comparison module, perform a difference operation on the difference feature and the i-th level codeword to obtain a difference feature, and input the difference feature into the next-level codebook comparison module; and In the N-th level codebook comparison module, compare the input difference feature with multiple codewords in the N-th level codebook corresponding to the N-th level codebook comparison module, determine the codeword closest to the difference feature as the codeword corresponding to the N-th level codebook comparison module.
3. The clustering method according to claim 2, wherein, The difference operation includes: Determine the difference feature based on the following expression: D j = C j - CW j Among them, j represents the j-th level codebook comparison module; D j represents the difference feature determined by the j-th level codebook comparison module; C j represents the input feature of the j-th level codebook comparison module; and CW j represents the j-th level codeword corresponding to the j-th level codebook comparison module.
4. A clustering model training method, applied to a clustering model, wherein, The clustering model includes N cascaded codebook comparison modules, where N is an integer greater than or equal to 1; wherein, the training method of the clustering model includes: Pre-construct an N-level codebook space; wherein, the N-level codebook space corresponds one-to-one with the N cascaded codebook comparison modules; For each level of the codebook space, initialize M cluster centers; where M is an integer greater than 1; During the training process, the following steps are respectively executed for each codebook comparison module of the clustering model: determining the cluster center among the M cluster centers that is closest to the input sample feature, taking it as the candidate cluster center corresponding to the codebook comparison module, and recording the distance between the input sample feature and the candidate cluster center as the candidate distance corresponding to the codebook comparison module; and taking the difference feature between the input sample feature and the candidate cluster center as the input feature of the next-level codebook comparison module. Summing up the candidate distances corresponding to the N cascaded codebook comparison modules to obtain a similarity loss; and optimizing the M cluster centers of the N cascaded codebook comparison modules respectively with the minimum of the similarity loss as the first optimization objective to obtain the optimized M cluster centers. Based on the candidate cluster centers corresponding to the N cascaded codebook comparison modules, perform codebook restoration and feature reconstruction to obtain reconstructed features; taking the difference between the input sample feature and the reconstructed features as a reconstruction loss; and optimizing the M cluster centers of the N cascaded codebook comparison modules respectively with the minimum of the reconstruction loss as the second optimization objective to obtain the optimized M cluster centers; and When the training end condition is satisfied, taking the M cluster centers corresponding to the N cascaded codebook comparison modules respectively as the M codewords of the codebooks corresponding to the N cascaded codebook comparison modules.
5. A clustering device, comprising: An input interface, a clustering model, and a classification label determination module; where The input interface is used to receive the multimodal representation information output by the multimodal pre-training model and input the received multimodal representation information into the clustering model. The clustering model includes: N cascaded codebook comparison modules; where N is an integer greater than or equal to 1; and each level of the N cascaded codebook comparison modules respectively corresponds to a codebook in the N-layer codebook space, and is used to compare the input feature with each codeword in its corresponding codebook, determine the codeword that is closest to the input feature, take it as its corresponding codeword, and take the difference feature between the input feature and the codeword as the input feature of the next-level codebook comparison module; and The classification label determination module is used to determine the latent classification label corresponding to the multimodal representation information based on the codewords corresponding to the N cascaded codebook comparison modules; where Determining the latent classification label corresponding to the multimodal representation information based on the codewords corresponding to the N cascaded codebook comparison modules includes: Determining the latent classification label corresponding to the multimodal representation information based on the following expression: Wherein, T represents the stealth classification label corresponding to the multi-modal representation information; j represents the j-th level codebook comparison module; represents the serial number of the codeword CWj corresponding to the j-th level codebook comparison module in the j-th level codebook; and M represents the number of codewords in each codebook.
6. An electronic device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the clustering method according to any one of claims 1-3.
7. A non-transitory computer-readable storage medium, the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the clustering method according to any one of claims 1-3.
8. A computer program product comprising computer program instructions which, when run on a computer, cause the computer to execute the clustering method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Fast database search systems and methods
CN110168525A