Incremental Clustering Method, Device, Computer Equipment and Storage Medium

By following group attribute constraints and manual discrimination in the process of incremental clustering, the problem of insufficient confirmation of suspected merged results in incremental clustering is solved, and the accuracy and efficiency of image classification and recognition are improved.

CN113936162BActive Publication Date: 2025-06-10HENGRUI (CHONGQING) ARTIFICIAL INTELLIGENCE TECH RES INST CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111248666.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-26
Publication Date
2025-06-10
Estimated Expiration
2041-10-26

AI Technical Summary

Technical Problem

During the incremental clustering process, the confirmation timeliness of suspected merge results is short, the impact is weak and it cannot be persisted, resulting in low accuracy and efficiency of image classification and recognition. Each manual intervention requires pausing the incremental clustering process, affecting work efficiency.

Method used

By obtaining the pending data, incremental clustering is performed according to the group attribute constraints, clustering results are generated and historical data are updated. The candidate merge queue is managed for suspected merge data, the processing priority is determined, and the merge or split results are determined through manual judgment, the discrimination results are temporarily stored and processed when the preset conditions are met, and the incremental clustering process is restored.

Benefits of technology

It effectively solves the timeliness and influence of confirming suspected merge results in incremental clustering, improves the accuracy and efficiency of image classification and recognition, and reduces the impact of manual intervention on the incremental clustering process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113936162B_ABST
    Figure CN113936162B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of big data processing, and specifically provides an incremental clustering method, device, computer device and storage medium, aiming to solve the problem of how to process the suspected merged result pairs in the incremental clustering result. For this purpose, the method of the present invention includes: obtaining new data and historical data, where the data sample contains group attributes, and the group attributes include group names and team names; performing incremental clustering and / or manual intervention to process suspected merged data pairs in accordance with the group attribute constraints, and adding a cache and batch processing mechanism for the suspected merge discrimination results in the manual intervention processing flow, so that incremental clustering and manual discrimination can be performed simultaneously. By applying the method of the present invention, the incorrect merges in clustering are suppressed, and the correct merges are maintained, improving the clustering accuracy; at the same time, the timeliness and influence of the manual confirmation information are strengthened, enabling it to be persistent; and due to the introduction of the parallel mechanism of incremental clustering and manual discrimination, the clustering efficiency is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of big data processing, and specifically provides an incremental clustering method, apparatus, computer device, and storage medium. Background Art

[0002] Clustering is the process of classifying data into different classes or clusters. Cluster analysis is an exploratory analysis. During the classification process, people do not need to give a classification standard in advance. Cluster analysis can start from sample data and automatically classify. Cluster analysis has been widely used in fields such as business data analysis, pattern recognition, and image processing.

[0003] How to efficiently obtain information from massive data has become the focus of research on clustering algorithms. In current clustering algorithms, as the data scale increases and the difficulty of data distribution increases, the probability of high-score negative examples also increases, seriously affecting the accuracy of clustering and resulting in merger errors in clustering results. When confirming suspected merger results in the clustering results of incremental clustering, if simply processing by merging or not merging according to the confirmation results, although the accuracy of clustering can be increased, since the information not merged is simply discarded and not inherited and continued, it cannot generate more value, and thus has a greater impact on the accuracy of image classification and recognition in practical applications such as image recognition and pattern recognition; moreover, each manual intervention requires pausing the incremental clustering process, and the mutual waiting between the incremental clustering process and the human-computer interaction process will seriously affect work efficiency and reduce the clustering efficiency of the entire incremental clustering system. Therefore, how to effectively solve the problems of short timeliness, weak influence, and inability to be persistent when confirming suspected merger results generated during the incremental clustering process, improve the accuracy and efficiency of image classification and recognition, and at the same time reduce the impact of manual intervention on the incremental clustering process has become an urgent problem in this field.

[0004] Correspondingly, a new solution is needed in this field to solve the above problems. Summary of the Invention

[0005] The present invention aims to solve the above technical problems, that is, to solve the problems of short timeliness, weak influence, and inability to be persistent when confirming suspected merger results generated during the incremental clustering process, and each manual intervention requires pausing the incremental clustering process, resulting in low accuracy and efficiency of image classification and recognition.

[0006] In a first aspect, the present invention provides an incremental clustering method, and the method includes:

[0007] The method includes:

[0008] S1. Obtain the data to be processed, where the data to be processed includes new data and first clustered data. The first clustered data is historical data that has been clustered. The data type of the data to be processed includes image data or text data, and each sample in the data to be processed contains group attributes, where the group attributes include group numbers and team numbers;

[0009] S2. Follow the group attribute constraints to perform incremental clustering on the data to be processed, obtaining a clustering result. The clustering result includes second clustered data that has completed incremental clustering and / or suspected merge data. The suspected merge data contains one or more suspected merge data pairs. Each suspected merge data pair includes two merge objects and the similarity score between the two merge objects. Each merge object contains a label and at least one sample;

[0010] S3. Update the first clustered data according to the second clustered data;

[0011] S4. Manage the candidate merge queue for the suspected merge data to obtain the processing priority queue of the suspected merge data pairs;

[0012] S5. Obtain the suspected merge data pair with the highest priority in the processing priority queue, and discriminate the suspected merge data pair with the highest priority to obtain a category discrimination result. The category discrimination result includes one of impure category, non - merge category, or merge category;

[0013] S6. Temporarily store the category discrimination result;

[0014] S7. In response to a preset execution condition, continue to execute step S1 or enter the category discrimination result processing flow.

[0015] In an embodiment of the above incremental clustering method, the steps of the "category discrimination result processing flow" specifically include:

[0016] S8. Pause the incremental clustering operation;

[0017] S9. According to the category discrimination result temporarily stored in step S6 and following the group attribute constraints, obtain third - type clustered data;

[0018] S10. Update the first clustered data according to the third clustered data;

[0019] S11. Resume the incremental clustering operation and execute step S1 to start a new incremental clustering.

[0020] In an embodiment of the above incremental clustering method, the step of "according to the category discrimination result temporarily stored in step S6 and following the group attribute constraints, obtain third - type clustered data" specifically includes:

[0021] S31. Obtain the temporarily stored class discrimination result;

[0022] S32. When the class discrimination result is an impure class, perform group attribute splitting on the samples of the suspected merged data pair to obtain the third type of clustering data, and execute step S35;

[0023] S33. When the class discrimination result is a non-merging class, perform a first group attribute assignment on the group attributes of the samples of the suspected merged data pair to obtain the third type of clustering data, and execute step S35;

[0024] S34. When the discrimination result is a merging class, follow the group attribute constraints and execute the processing flow of the merging class to obtain the third type of clustering data;

[0025] S35. Check whether all the temporarily stored class discrimination results have been processed;

[0026] If not, return to step S31;

[0027] If so, return to step S10.

[0028] In an embodiment of the above incremental clustering method, the step of "when the class discrimination result is an impure class, perform group attribute splitting on the samples of the suspected merged data pair" specifically includes:

[0029] Divide the samples in the suspected merged data pair with the same group number and team number into the first split class;

[0030] Divide the samples in the remaining suspected merged data pairs into the second split class;

[0031] Set new labels for the first split class and the second split class respectively.

[0032] In an embodiment of the above incremental clustering method, the step of "when the class discrimination result is a non-merging class, perform a first group attribute assignment on the group attributes of the samples of the suspected merged data pair" specifically includes:

[0033] Assign a brand-new and identical group number to the samples in the suspected merged data pair;

[0034] Assign team numbers with different values to the samples corresponding to the two types of labels in the suspected merged data pair respectively.

[0035] In an embodiment of the above incremental clustering method, the step of "when the discrimination result is a merging class, follow the group attribute constraints and execute the processing flow of the merging class" specifically includes:

[0036] Assign a second group attribute value to the group attribute of the sample in the suspected merged data pair;

[0037] Obtain the latest label of the sample and the sample corresponding to the latest label from the suspected merged data pair;

[0038] Check whether there is a conflict in the group attribute constraints of the sample corresponding to the latest label;

[0039] If there is a conflict, do not merge;

[0040] If there is no conflict, merge;

[0041] Among them, the step of "assigning a second group attribute value to the group attribute of the sample in the suspected merged data pair" specifically includes:

[0042] Assign the same new group number to the samples in the suspected merged data pair;

[0043] Assign the same team number to the samples in the suspected merged data pair.

[0044] In an embodiment of the above incremental clustering method, the step of "checking whether there is a conflict in the group attribute constraints of the sample corresponding to the latest label" specifically includes:

[0045] Select all samples with a group number not equal to 0;

[0046] When the group numbers are the same and the team numbers are the same, the group attributes do not conflict;

[0047] When the group numbers are the same but the team numbers are different, the group attributes conflict.

[0048] In an embodiment of the above incremental clustering method, the method of group attribute constraints includes:

[0049] The sample with a group number of 0 has an invalid group attribute and has no constraint relationship with any other clustering data;

[0050] The sample with a group number not equal to 0 has a valid group attribute;

[0051] Samples with the same group number and the same team number have the same label;

[0052] Samples with the same group number and different team numbers have different labels;

[0053] There is no constraint relationship between samples with different group numbers.

[0054] In an embodiment of the above incremental clustering method, the step of "responding to a preset execution condition and continuing to execute step S1 or entering the category discrimination result processing flow" specifically includes:

[0055] When the number of suspected merge data pairs obtained in step S2 reaches a preset number threshold, or when the cumulative statistical incremental clustering completion time in step S2 reaches a preset time threshold, the process enters the category discrimination result processing flow.

[0056] In a second aspect, the present invention provides an incremental clustering device, the device comprising: The device comprises:

[0057] A data loading module, the data acquisition module is configured to acquire data to be processed, the data to be processed includes new data and first clustered data, the first clustered data is historical data that has been clustered, the data type of the data to be processed includes image data or text data, and each sample in the data to be processed includes group attributes, the group attributes include group numbers and team numbers;

[0058] An incremental clustering module, the incremental clustering module is configured to perform incremental clustering on the data to be processed following group attribute constraints, to obtain a clustering result, the clustering result includes second clustered data that has completed incremental clustering and / or suspected merge data, the suspected merge data includes one or more suspected merge data pairs, the suspected merge data pair includes two merge objects and the similarity score between the two merge objects, and each merge object includes a label and at least one sample;

[0059] A historical data maintenance module, the historical data maintenance module is configured to update the first clustered data according to the second clustered data;

[0060] A suspected merge data management module, the suspected merge data management module is configured to perform candidate merge queue management on the suspected merge data to obtain a processing priority queue of the suspected merge data pairs;

[0061] A human-computer interaction discrimination module, the human-computer interaction discrimination module is configured to perform the following operations:

[0062] Obtain the suspected merge data pair with the highest priority in the processing priority queue, discriminate the suspected merge data pair with the highest priority to obtain a category discrimination result, the category discrimination result includes one of an impure category, a non-merging category or a merging category;

[0063] Temporarily store the category discrimination result;

[0064] In response to a preset execution condition, continue a new incremental clustering process or enter the category discrimination result processing flow.

[0065] In an embodiment of the above incremental clustering device, the human-computer interaction discrimination module is further configured to perform the following operations:

[0066] Suspend the incremental clustering operation;

[0067] Obtain the third type of clustering data according to the temporarily stored class discrimination result and following the group attribute constraint;

[0068] Update the first clustering data according to the third clustering data;

[0069] Resume the incremental clustering process and start a new incremental clustering process.

[0070] In an embodiment of the above incremental clustering device, the human-computer interaction discrimination module is further configured to perform the following operations:

[0071] Obtain the temporarily stored class discrimination result;

[0072] When the class discrimination result is an impure class, perform group attribute splitting on the samples of the suspected merging data pair to obtain the third type of clustering data, and go to "Check whether all the temporarily stored class discrimination results have been processed";

[0073] When the class discrimination result is a non-merging class, perform first group attribute assignment on the group attributes of the samples of the suspected merging data pair to obtain the third type of clustering data, and go to "Check whether all the temporarily stored class discrimination results have been processed";

[0074] When the discrimination result is a merging class, follow the group attribute constraint and execute the process for handling the merging class to obtain the third type of clustering data;

[0075] Check whether all the temporarily stored class discrimination results have been processed;

[0076] If not, obtain and process the unprocessed temporarily stored class discrimination result;

[0077] If so, go to "Update the first clustering data according to the third clustering data".

[0078] In an embodiment of the above incremental clustering device, the human-computer interaction discrimination module specifically performs the following operations:

[0079] Divide the samples in the suspected merging data pair with the same group number and team number into the first split class;

[0080] Divide the samples in the remaining suspected merging data pairs into the second split class;

[0081] Set new labels for the first split class and the second split class respectively.

[0082] In one embodiment of the above incremental clustering device, the human-computer interaction discrimination module specifically performs the following operations:

[0083] Assign a new and numerically identical group number to the samples in the suspected merged data pair;

[0084] Assign numerically different team numbers to the samples corresponding to the two types of labels in the suspected merged data pair respectively.

[0085] In one embodiment of the above incremental clustering device, the human-computer interaction discrimination module specifically performs the following operations:

[0086] Perform a second group attribute assignment on the group attributes of the samples in the suspected merged data pair;

[0087] Obtain the latest label of the sample and the sample corresponding to the latest label according to the samples in the suspected merged data pair;

[0088] Check whether there is a conflict in the group attribute constraints of the samples corresponding to the latest label;

[0089] If there is a conflict, do not merge;

[0090] If there is no conflict, merge;

[0091] Among them, the step of "performing a second group attribute assignment on the group attributes of the samples in the suspected merged data pair" specifically includes:

[0092] Assign a new and numerically identical group number to the samples in the suspected merged data pair;

[0093] Assign numerically identical team numbers to the samples in the suspected merged data pair.

[0094] In one embodiment of the above incremental clustering device, the human-computer interaction discrimination module specifically performs the following operations:

[0095] Select all samples with a group number not equal to 0;

[0096] When the group numbers are the same and the team numbers are the same, the group attributes do not conflict;

[0097] When the group numbers are the same but the team numbers are different, the group attributes conflict.

[0098] In one embodiment of the above incremental clustering device, the method for group attribute constraints includes:

[0099] The samples with a group number of 0 have invalid group attributes and have no constraint relationship with any other clustering data;

[0100] The samples with a group number not equal to 0 have valid group attributes;

[0101] Samples with the same group number and the same team number have the same label;

[0102] Samples with the same group number but different team numbers have different labels;

[0103] There is no constraint relationship between samples with different group numbers.

[0104] In an embodiment of the above incremental clustering device, the suspected merge data management module specifically performs the following operations:

[0105] When the number of pairs of suspected merge data obtained during the incremental clustering process reaches a preset quantity threshold, or when the cumulative statistical incremental clustering completion time during the incremental clustering process reaches a preset time threshold, enter the category discrimination result processing flow.

[0106] In a third aspect, the present invention proposes a computer device, including a processor and a storage device, the storage device being adapted to store multiple program codes, and the program codes being adapted to be loaded and run by the processor to execute the incremental clustering method described in any one of the above solutions.

[0107] In a fourth aspect, the present invention proposes a storage medium, the storage medium being adapted to store multiple program codes, and the program codes being adapted to be loaded and run by a processor to execute the incremental clustering method described in any one of the above solutions.

[0108] In the case of adopting the above technical solutions, the present invention determines whether the pairs of suspected merge data generated during the incremental clustering process need to be merged or split manually, so as to suppress incorrect merges and maintain correct merges; group attribute information and related checks are added to the clustering data during the incremental clustering process and when manually confirming pairs of suspected merge data, strengthening the timeliness and influence of the manually intervened confirmation information, enabling it to be persistent, achieving the purpose of reinforcement learning, solving the problem of how to effectively process pairs of suspected merge data in incremental clustering, and further improving the accuracy and efficiency of image classification and recognition; and a manual discrimination result cache module is added during the process of processing the manually intervened confirmation information, enabling the incremental clustering and manual discrimination to be executed simultaneously, thereby further improving the efficiency of the incremental clustering. BRIEF DESCRIPTION OF THE DRAWINGS

[0109] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings, in which:

[0110] Figure 1 is the main step flowchart of the general incremental clustering method.

[0111] Figure 2 is the main step flowchart of the incremental clustering method of the embodiment of the present invention.

[0112] Figure 3 is Figure 2 The specific implementation flowchart of step S208 in

[0113] Figure 4 is Figure 3 The main step flowchart of obtaining the third clustering result according to the category discrimination result in step S2082 in

[0114] Figure 5 It is a schematic diagram of the composition structure of the incremental clustering device of the present invention. Specific implementation manners

[0115] First, read Figure 1 , Figure 1 is the main step flowchart of the general incremental clustering method. As Figure 1 shown, the general incremental clustering method includes:

[0116] Step S101: Obtain the data to be processed, where the data to be processed includes new data and historical data;

[0117] Step S102: Perform incremental clustering on the data to be processed to obtain an incremental clustering result;

[0118] Step S103: Update the historical data according to the incremental clustering result.

[0119] In the general incremental clustering method, usually only steps S101, S102, and S103 are included. After completing step S103, the incremental clustering process will return to step S101, use the updated historical data and new data, and repeat steps S101, S102, and S103 to complete the incremental clustering. Moreover, in the general incremental clustering method, the processing of suspected merged data between determined merge and determined non-merge is often not paid much attention to, but the suspected merged data is simply classified into the determined merge result, which sometimes causes clustering merge errors. Especially as the data scale increases and the data distribution difficulty increases, merge errors are prone to occur and accumulate in the incremental clustering result, as well as the situation of more and / or larger-scale misclassified clusters, thus affecting the overall clustering performance.

[0120] The incremental clustering method of the present invention is precisely proposed in view of the deficiency that the general incremental clustering method always fails to effectively process suspected merged data. It is hoped that by introducing a method of manually discriminating suspected merged data, the deficiency of the incremental clustering algorithm can be made up, the situation of clustering result merge errors in incremental clustering can be solved, thereby improving the accuracy of clustering, and enabling the information confirmed manually to be persistent, achieving the purpose of reinforcement learning.

[0121] Continue to read Figure 2 ,Figure 2 This is the main step flowchart of the incremental clustering method according to an embodiment of the present invention. As Figure 2 shown, the incremental clustering method of the present invention includes:

[0122] Step S201: Obtain data to be processed, where the data to be processed includes newly added data and first clustering data;

[0123] Step S202: Follow the group attribute constraint, perform incremental clustering on the data to be processed, and obtain an incremental clustering result, where the incremental clustering result includes second clustering data and / or suspected merge data;

[0124] Step S203: Update the first clustering data according to the second clustering data;

[0125] Step S204: Manage the candidate merge queue for the suspected merge data to obtain a processing priority queue for the suspected merge data pairs;

[0126] Step S205: Select the suspected merge data pair with the highest priority in the processing priority queue, perform manual discrimination, and obtain a category discrimination result;

[0127] Step S206: Temporarily store the category discrimination result;

[0128] Step S207: Select to execute Step S201 to start a new incremental clustering, or enter the category discrimination result processing flow in Step 208;

[0129] Step S208: Category discrimination result processing flow.

[0130] In Step S201, the data to be processed can be image data, for example, using the images of surveillance videos for face clustering; it can also be text data, for example, using the consumption data of users for customer clustering analysis, etc.

[0131] Each sample of the data to be processed contains label information and group attribute information. The label information is used to identify the classification result of the incremental clustering, and the group attribute information is additional incremental clustering information.

[0132] The group attribute consists of a group number (group) and a team number (team), and the incremental clustering process of the present invention follows the group attribute constraint. The method of the group attribute constraint specifically includes:

[0133] Samples with a group number of 0 have invalid group attributes and have no constraint relationship with any other clustering data;

[0134] Samples with a group number not equal to 0 have valid group attributes;

[0135] Samples with the same group number and the same team number have the same label;

[0136] Samples with the same group number but different team numbers have different labels;

[0137] There is no constraint relationship between samples with different group numbers.

[0138] The first clustering data is historical data that has been clustered, and the samples of the first clustering data include the labels added after the incremental clustering is completed. For new data, the default value of the label can be set, such as setting the default value of the label of the new data to 0 or other values, and the group number and team number in the initial group attributes of all samples are 0.

[0139] In addition, for incremental data, it is necessary to perform Batch (batch processing data) splitting according to business requirements and / or resource conditions, etc., to obtain the new data in step S201. The method of Batch splitting is not limited in this invention. As an example, it can be split according to time periods, such as setting 5 minutes as a time period. Those skilled in the art can select the method of Batch splitting according to the actual situation.

[0140] In step S202, the specific algorithm of incremental clustering is not limited in this invention. As an example, the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) density clustering algorithm, hierarchical clustering algorithm, etc. can be used. Those skilled in the art can select the incremental clustering algorithm according to the actual situation.

[0141] The incremental clustering result includes the second clustering data and / or suspected merge data. The second clustering data is the clustering data that has completed incremental clustering and has had labels assigned.

[0142] The suspected merge data contains one or more suspected merge data pairs. In this embodiment, preferably, a suspected merge data pair includes two merge objects and the similarity score between the two merge objects. Each merge object includes its own label and at least one sample. As an example, a suspected merge data pair can be expressed as: {(labelA, sampleA1), (labelB, sampleB1), Score}, and the specific meaning is that merge object 1 is the label labelA and the sample sampleA1 in labelA, merge object 2 is the label labelB and the sample sampleB1 in labelB, and the similarity score between merge object 1 and merge object 2 is Score. Also, the merge objects in the suspected merge data pair can include multiple samples. For example, the merge object is (labelA, sampleA1, sampleA2, sampleA3), and at this time, the label labelA of this merge object contains 3 samples, sampleA1, sampleA2, and sampleA3.

[0143] At this time, the clustering labels labelA and labelB in the suspected merged data pair can be labels newly generated by incremental clustering or labels that already exist in the first clustering data. Therefore, the combination forms of the labels in the suspected merged data pair may be one of the combinations of: new label + new label, new label + existing label, or existing label + existing label.

[0144] The present invention does not limit the calculation method of the similarity score. As an example, the inner product score of two merging objects can be used for calculation. Those skilled in the art can select the calculation method of the similarity score according to the actual situation.

[0145] It should be noted that the scheme of selecting two merging objects as a group for the suspected merged data pair is considered because the number 2 is the smallest unit of the suspected merged data pair, and the logical judgment is the simplest. At this time, manual discrimination only needs to give the result of merging or not merging, which is very convenient for manual processing. When there are multiple merging objects in the incremental clustering result, multiple data can be paired two by two so that each suspected merged data pair only contains two merging objects.

[0146] It should be noted that the group attributes of all samples are given by the result confirmed by manual intervention. That is, only in the relevant operations of step S205 can non-zero values be assigned to the group number and team number in the sample group attributes, while the group attributes remain unchanged in the incremental clustering in step S202. It is precisely through the group attributes in the samples that the manually intervened confirmation information can be persistent and strengthen the learning of the incremental clustering process.

[0147] In the incremental clustering of step S202, the group attribute constraints need to be followed. As an example, if the group numbers of sample sampleA1 (groupC, teamA) and sample sampleB1 (groupC, teamB) are the same, both being groupC, but the team numbers are different, being teamA and teamB respectively, when performing incremental clustering, the same label will not be assigned to sampleA1 and sampleB1.

[0148] In step S203, according to the second clustering data obtained in step S202, the first clustering data is updated based on the second clustering data.

[0149] In step S204, each time after the incremental clustering is completed, when the incremental clustering result contains suspected merged data, the suspected merged data pair is updated in real time; and according to the similarity scores included in the suspected merged data pair, the suspected merged data pair is sorted to form a processing priority queue. Usually, the higher the similarity score of the suspected merged data pair, the higher its processing priority.

[0150] In this embodiment, a human-computer interaction process for suspected merged data is introduced. That is, whether to merge the suspected merged data in step S205 and how to perform the merge are determined manually, and the other steps are preferentially executed by a computer program.

[0151] In step S205, the pair of suspected merged data with the highest priority in the processing priority queue is selected for manual discrimination to obtain a category discrimination result, which includes an impure category, a non-merging category, or a merging category.

[0152] In step S206, the category discrimination result obtained through manual discrimination in step S205 is temporarily stored. According to the preset execution conditions set in step S207, it is processed in a centralized batch manner.

[0153] In step S207, according to actual needs, it is possible to choose to execute step S201 to start a new incremental clustering or enter the category discrimination result processing flow in step 208. As an example, the execution conditions for the category discrimination result processing flow can be preset. These execution conditions can be randomly preset by the user or set according to some metrics. For example, a quantity threshold for the pair of suspected merged data can be set. When the set threshold is reached, a reminder is given that the category discrimination result processing flow needs to be executed; or an execution time threshold can be set. Starting from the start of a certain incremental clustering as the timing starting point, the time for the completion of each subsequent incremental clustering operation is cumulatively timed. If the cumulative timing reaches the set execution time threshold, the category discrimination result processing flow is entered.

[0154] Continue reading Figure 3 , Figure 3 is the specific implementation method of step S208, and its method includes:

[0155] Step S2081: Pause the incremental clustering operation;

[0156] Step S2082: According to the category discrimination result temporarily stored in step S206 and following the group attribute constraints, obtain the third clustering data;

[0157] Step S2084: Update the first clustering data according to the third clustering data;

[0158] Step S2085: Resume the incremental clustering operation and return to step S201 to start a new incremental clustering.

[0159] Considering that if the category discrimination result processing process and incremental clustering are carried out simultaneously, after the incremental clustering result is updated, the labels and samples in the historical data may change, resulting in the inability to ensure the accuracy of the data such as the latest labels and samples that need to be obtained again in the category discrimination result processing process. Therefore, in order to avoid the mutual influence between the incremental clustering result and the category discrimination result processing process, usually after suspending the incremental clustering operation in step S2081, the category discrimination result processing process is started.

[0160] Before suspending the incremental clustering, it is necessary to check whether there is an ongoing incremental clustering; if not, directly execute step S2082; if so, wait for the ongoing incremental clustering to complete, and after executing steps S203 and S204, then suspend the incremental clustering process.

[0161] Continue reading Figure 4 , Figure 4 For Figure 3 the flowchart of step S2083 in obtaining the third clustering result according to the category discrimination result in

[0162] In step S401, obtain the number of category discrimination results temporarily stored in step S206, and according to the priority order of manual processing, obtain the category discrimination result with the highest current priority.

[0163] Execute step S402 to check which of the three results the category discrimination result is. First, in step S403, judge whether the category discrimination result is an impure category.

[0164] When the category discrimination result is an impure category, execute step S404 for group attribute splitting, transfer to step S414 to obtain the third type of clustering result, and execute step S415 to check whether all the temporarily stored category discrimination results have been processed. Among them, the steps of group attribute splitting are:

[0165] Divide the samples in the suspected merged data pairs with the same group number and team number into the first split category;

[0166] Divide the samples in the remaining suspected merged data pairs into the second split category;

[0167] Set new labels for the first split category and the second split category respectively.

[0168] As an example, the suspected merged data pairs are labelA{sampleA1(groupA, teamA), sampleA2(groupA, teamA), sampleA3(groupA3, teamA3)}, labelB{sampleB1(groupB, teamB)}, the label labelA contains 3 samples sampleA1, sampleA2, and sampleA3, the label labelB contains 1 sample sampleB1, the group numbers and team numbers of samples sampleA1 and sampleA2 are the same, but different from those of sampleA3; after splitting by group attributes, the first split category is sampleA1 and sampleA2, and the second split category is sampleA3 and sampleB1; after setting new labels for the first split category and the second split category respectively, we get labelC{sampleA1(groupA, teamA), sampleA2(groupA, teamA)} and labelD{sampleB1(groupB, teamB, sampleA3(groupA3, teamA3)}.

[0169] When the category discrimination result is not an impure category, steps S405 and S406 are executed to obtain the label of the merging object in the currently highest-priority suspected merged data pair corresponding to the category discrimination result in the temporary storage and the samples corresponding to the label. It is judged in step S407 whether the category discrimination result is a merged category.

[0170] When the category discrimination result is a non-merging category, step S408 is executed to perform the first group attribute assignment, transfer to step S414 to obtain the third type of clustering result, and step S415 is executed to check whether all the category discrimination results in the temporary storage have been processed. Among them, the steps of the first group attribute assignment are as follows:

[0171] Assign brand-new and identical group numbers to the samples in the suspected merged data pair;

[0172] Assign different team numbers to the samples corresponding to the two types of labels in the suspected merged data pair.

[0173] As an example, the suspected merged data pairs are labelA{sampleA1(groupA, teamA),}, labelB{sampleB1(groupB, teamB)}; the data after the first group of attribute assignments are labelA{sampleA1(groupC, teamA1),}, labelB{sampleB1(groupC, teamB1)}. Among them, groupC is a new and numerically identical group number, and teamA1 and teamB1 are team numbers with different numerical values.

[0174] When the category discrimination result is the merged category, step S409 is executed to perform the second group of attribute assignments on the group attributes of the samples in the suspected merged data pairs; among them, the steps of the second group of attribute assignments are as follows:

[0175] Assign a new and numerically identical group number to the samples of the suspected merged data pairs;

[0176] Assign team numbers with the same numerical value to the samples of the suspected merged data pairs.

[0177] As an example, the suspected merged data pairs are labelA{sampleA1(groupA, teamA),}, labelB{sampleB1(groupB, teamB)}; the data after the second group of attribute assignments are labelA{sampleA1(groupC, teamD),}, labelB{sampleB1(groupC, teamD)}, where groupC is a new and numerically identical group number, and teamD is a team number with the same numerical value.

[0178] In step S410, according to the samples in the suspected merged data pairs, from all historical data, including the first type of clustering data, the second type of clustering data, and the third type of clustering data, obtain the latest label of the sample and all samples corresponding to the latest label. As an example, the label of the sample sampleA1 in the suspected merged data pair is labelA, but due to the incremental clustering continuously updating the clustering result, the label of the sample sampleA1 has changed in the first type of clustering data label, and the latest label is labelF. At this time, step S410 modifies the label labelA to labelF, and takes out all the data with the label labelF in the first type of clustering data and executes step S411. If the label of the sample sampleA1 has not changed in the first type of clustering data, only all the samples with the label labelA need to be taken out from the historical data.

[0179] In step S411, for the newly added samples, perform a group attribute constraint check. The steps of the group attribute constraint check are as follows:

[0180] Select all samples whose group numbers are not 0;

[0181] When the group numbers are the same and the team numbers are the same, the group attributes do not conflict;

[0182] When the group numbers are the same but the team numbers are different, the group attributes conflict.

[0183] In step S412, when it is determined that there is a conflict in the group attribute constraint check, there is no need to process the group attributes of the samples. Instead, go to step S414 to obtain the third type of clustering result, and execute step S415 to check whether all the temporarily stored class discrimination results have been processed.

[0184] In step S412, when it is determined that there is no conflict in the group attribute constraint check, execute step S413. In step S413, merge the suspected merge data pairs to obtain a merged label and multiple samples corresponding to the label, and each sample has the same group number and team number.

[0185] As an example, after step S410, the suspected merge data pairs are labelA{sampleA1(groupC, teamD),}, labelB{sampleB1(groupC, teamD)}. If the manual discrimination is to merge into labelA, the merged result is labelA{sampleA1(groupC, teamD), sampleB1(groupC, teamD)}.

[0186] In step S415, after each processing of the discrimination result of a group of suspected merge data pairs, the number of temporarily stored class discrimination results is reduced by 1. When step S415 checks that there are still temporarily stored discrimination results, return to step S401 to obtain the highest-priority remaining suspected merge data pairs for processing; when step S415 checks that all the temporarily stored discrimination results have been processed, execute step S2083 to update the first clustering data.

[0187] In steps S203 and S2083, the methods for updating the first clustering data include:

[0188] Add a label to the newly added data and add it to the first clustering data. For example, in the normal incremental clustering process, the second clustering data obtained after the newly added data is clustered;

[0189] Modify the labels of the first clustering data. For example, after the suspected merge data pairs are processed manually and the labels are merged, it may be necessary to modify the labels of some samples in the first clustering data according to the third clustering data.

[0190] In step S2084, resume the incremental clustering operation and execute step S201 to start a new incremental clustering.

[0191] It should be noted that, in this embodiment, after all the temporarily stored discrimination results are processed, the third type of clustering data is used to update the first type of clustering data. In other embodiments, it can be set to update the first type of clustering data immediately after obtaining the third type of clustering data for each pair of suspected merged data pairs whose discrimination results are processed; it can also be set to update the first type of clustering data immediately after obtaining the third type of clustering data for multiple pairs of suspected merged data pairs whose discrimination results are processed. Those skilled in the art can use different methods to implement the update method of the first type of clustering data for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0192] It should be noted that, in the present invention, due to the introduction of methods such as group attributes, group attribute constraint conditions, and temporarily storing category discrimination results in step S206, the process of manually discriminating pairs of suspected merged data in step S205 does not need to pause the incremental clustering process, but can be executed in parallel with the clustering process. The user can manually process pairs of suspected merged data at any time without interrupting the incremental clustering, and temporarily store the discrimination results. When the conditions preset in step S207 are met, it enters the process of processing category discrimination results in step S208 to update the historical data. Although step S208 needs to pause the incremental clustering, step S208 can be completed by a computer program, and the time consumed by computer processing can be almost ignored. Therefore, through the method of the present invention, the influence of the manual interaction process on the incremental clustering process is greatly reduced, and the clustering efficiency is improved.

[0193] Furthermore, the present invention also provides an incremental clustering device. As Figure 5 shown, the incremental clustering device 5 of the present invention mainly includes: a data loading module 51, an incremental clustering module 52, a suspected merged data management module 53, a human-computer interaction discrimination module 54, and a historical data maintenance module 55.

[0194] As an example, the data loading module 51 is configured to execute step S201 to obtain the data to be processed. The incremental clustering module 52 is configured to execute the operations in step S202. The suspected merged data management module 53 is configured to execute the operations in step S204. The human-computer interaction discrimination module 54 is configured to execute the operations in steps S205, S206, S207, S208, S2081, S2082, S2084, and steps S401 to S415. The historical data maintenance module 55 is configured to execute the operations in steps S203 and S2083.

[0195] Furthermore, the present invention also provides a computer device, which includes a processor and a storage device. The storage device can be configured to store and execute the program of the incremental clustering method in the above method embodiments. The processor can be configured to execute the program in the storage device, and the program includes, but is not limited to, the program of executing the incremental clustering method in the above method embodiments. For the sake of convenience of description, only the parts related to the embodiments of the present invention are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present invention. The incremental clustering device can be a control device formed by various electronic devices.

[0196] Furthermore, the present invention also provides a storage medium, which can be configured to store the program of executing the incremental clustering method in the above method embodiments. The program can be loaded and run by a processor to implement the above incremental clustering method. For the sake of convenience of description, only the parts related to the embodiments of the present invention are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present invention. The storage medium can be a storage device formed by various electronic devices. Optionally, the storage medium in the embodiments of the present invention is a non-transitory computer-readable storage medium.

[0197] Those skilled in the art should be able to realize that the method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of electronic hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in the form of electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0198] It should be noted that the ordinal numbers such as "first", "second", "third", etc. in the description, claims, and drawings of the present invention are only used to distinguish similar objects, rather than to describe or represent a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein.

[0199] It should be noted that in the description of the present application, the term "A and / or B" represents all possible combinations of A and B, such as only A, only B, or A and B.

[0200] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.

Claims

1. An incremental clustering method, characterized in that, the method includes: S1. Obtain the data to be processed, where the data to be processed includes new data and first clustering data. The first clustering data is historical data that has completed clustering. The data type of the data to be processed includes image data or text data, and each sample in the data to be processed contains group attributes, where the group attributes include group numbers and team numbers; S2. Follow the group attribute constraints to perform incremental clustering on the data to be processed, and obtain a clustering result. The clustering result includes second clustering data that has completed incremental clustering and / or suspected merge data. The suspected merge data contains one or more suspected merge data pairs. The suspected merge data pair includes two merge objects and the similarity score between the two merge objects. Each merge object contains a label and at least one sample; S3. Update the first clustering data according to the second clustering data; S4. Perform candidate merge queue management on the suspected merge data to obtain a processing priority queue for the suspected merge data pairs; S5. Obtain the suspected merge data pair with the highest priority in the processing priority queue; S6. Obtain and temporarily store the category discrimination result, where the category discrimination result is obtained by manually discriminating the suspected merge data pair with the highest priority. The category discrimination result includes one of impure category, non-merging category, or merging category; S7. In response to a preset execution condition, continue to execute step S1 or enter the category discrimination result processing flow; The steps of the "category discrimination result processing flow" specifically include: S8. Pause the incremental clustering operation; S9. According to the category discrimination result temporarily stored in step S6 and following the group attribute constraints, obtain the third type of clustering data; Among them, the method of the group attribute constraints includes: The sample with a group number of 0 has an invalid group attribute and has no constraint relationship with any other clustering data; The sample with a group number not equal to 0 has a valid group attribute; The samples with the same group number and the same team number have the same label; The samples with the same group number and different team numbers have different labels; There is no constraint relationship between the samples with different group numbers.

2. The incremental clustering method according to claim 1, characterized in that, the steps of the "category discrimination result processing flow" specifically further include: S10. Update the first clustering data according to the third type of clustering data; S11. Resume the incremental clustering operation and execute step S1 to start a new incremental clustering.

3. The incremental clustering method according to claim 2, characterized in that, the step of "according to the category discrimination result temporarily stored in step S6 and following the group attribute constraints, obtain the third type of clustering data" specifically includes: S31. Obtain the temporarily stored category discrimination result; S32. When the category discrimination result is an impure category, perform group attribute splitting on the samples of the suspected merge data pair to obtain the third type of clustering data, and execute step S35; S33. When the class discrimination result is non - merging classes, perform a first group - attribute assignment on the group attributes of the samples of the suspected merging data pairs to obtain the third - type clustering data, and execute step S35; S34. When the discrimination result is merging classes, follow the group - attribute constraints and execute the processing flow for merging classes to obtain the third - type clustering data; S35. Check whether all the temporarily stored class discrimination results have been processed; If not, return to step S31; If so, return to step S10.

4. The incremental clustering method according to claim 3, wherein, the step of "when the class discrimination result is impure classes, perform group - attribute splitting on the samples of the suspected merging data pairs" specifically includes: Dividing the samples in the suspected merging data pairs with the same group number and the same team number into a first split class; Dividing the samples in the remaining suspected merging data pairs into a second split class; Setting new labels for the first split class and the second split class respectively.

5. The incremental clustering method according to claim 3, wherein, the step of "when the class discrimination result is non - merging classes, perform a first group - attribute assignment on the group attributes of the samples of the suspected merging data pairs" specifically includes: Assigning a new and identical group number to the samples in the suspected merging data pairs; Assigning different team numbers to the samples corresponding to the two types of labels in the suspected merging data pairs respectively.

6. The incremental clustering method according to claim 3, wherein, the step of "when the discrimination result is merging classes, follow the group - attribute constraints and execute the processing flow for merging classes" specifically includes: Performing a second group - attribute assignment on the group attributes of the samples in the suspected merging data pairs; Obtaining the latest label of the sample and the sample corresponding to the latest label according to the samples in the suspected merging data pairs; Checking whether there is a conflict in the group - attribute constraints of the samples corresponding to the latest label; If there is a conflict, do not merge; If there is no conflict, merge; wherein, the step of "performing a second group - attribute assignment on the group attributes of the samples in the suspected merging data pairs" specifically includes: Assigning a new and identical group number to the samples of the suspected merging data pairs; Assigning an identical team number to the samples of the suspected merging data pairs.

7. The incremental clustering method according to claim 6, wherein, the step of "checking whether there is a conflict in the group - attribute constraints of the samples corresponding to the latest label" specifically includes: Selecting all samples with a group number not equal to 0; When the group numbers are the same and the team numbers are the same, the group attributes do not conflict; When the group numbers are the same but the team numbers are different, the group attributes conflict.

8. The incremental clustering method according to claim 1, wherein, the step of "in response to a preset execution condition, continue to execute step S1 or enter the class discrimination result processing flow" specifically includes: When the number of suspected merging data pairs obtained in step S2 reaches a preset quantity threshold, or when the cumulative statistical incremental clustering completion time in step S2 reaches a preset time threshold, enter the class discrimination result processing flow.

9. An incremental clustering device, characterized in that, the device includes: A data loading module, the data acquisition module is configured to acquire data to be processed, the data to be processed includes new data and first clustering data, the first clustering data is historical data that has completed clustering, the data type of the data to be processed includes image data or text data, and each sample in the data to be processed contains group attributes, and the group attributes include group numbers and team numbers; An incremental clustering module, the incremental clustering module is configured to perform incremental clustering on the data to be processed following group attribute constraints to obtain a clustering result, the clustering result includes second clustering data that has completed incremental clustering and / or suspected merger data, the suspected merger data contains one or more suspected merger data pairs, the suspected merger data pair includes two merger objects and the similarity score between the two merger objects, and each merger object contains a label and at least one sample; A historical data maintenance module, the historical data maintenance module is configured to update the first clustering data according to the second clustering data; A suspected merger data management module, the suspected merger data management module is configured to perform candidate merger queue management on the suspected merger data to obtain a processing priority queue of the suspected merger data pairs; A human-computer interaction discrimination module, the human-computer interaction discrimination module is configured to perform the following operations: Obtain the suspected merger data pair with the highest priority in the processing priority queue; Obtain and temporarily store a category discrimination result, where the category discrimination result is obtained by manually discriminating the suspected merger data pair with the highest priority, and the category discrimination result includes one of impure category, non-merger category or merger category; In response to a preset execution condition, continue a new incremental clustering process or enter a category discrimination result processing flow; The "category discrimination result processing flow" specifically performs the following operations: Pause the incremental clustering operation; According to the temporarily stored category discrimination result and following the group attribute constraints, obtain third-class clustering data; Among them, the method of the group attribute constraints includes: The sample with a group number of 0 has an invalid group attribute and has no constraint relationship with any other clustering data; The sample with a group number not equal to 0 has a valid group attribute; Samples with the same group number and the same team number have the same label; Samples with the same group number and different team numbers have different labels; There is no constraint relationship between samples with different group numbers.

10. The incremental clustering device according to claim 9, characterized in that, the human-computer interaction discrimination module is further configured to perform the following operations: Update the first clustering data according to the third-class clustering data; Resume the incremental clustering process and start a new incremental clustering process.

11. The incremental clustering device according to claim 10, characterized in that, the human-computer interaction discrimination module is further configured to perform the following operations: Obtain the temporarily stored category discrimination result; When the category discrimination result is an impure category, perform group attribute splitting on the samples of the suspected merged data pair to obtain the third type of clustering data, and go to "check whether all the temporarily stored category discrimination results have been processed"; When the category discrimination result is a non-merging category, perform first group attribute assignment on the group attributes of the samples of the suspected merged data pair to obtain the third type of clustering data, and go to "check whether all the temporarily stored category discrimination results have been processed"; When the discrimination result is a merging category, follow the group attribute constraints and execute the process for handling the merging category to obtain the third type of clustering data; Check whether all the temporarily stored category discrimination results have been processed; If not, obtain and process the unprocessed temporarily stored category discrimination results; If so, go to "update the first clustering data according to the third type of clustering data".

12. The incremental clustering device according to claim 11, wherein, the human-computer interaction discrimination module specifically performs the following operations: Divide the samples in the suspected merged data pairs with the same group number and the same team number into the first split category; Divide the samples in the remaining suspected merged data pairs into the second split category; Set new labels for the first split category and the second split category respectively.

13. The incremental clustering device according to claim 11, wherein, the human-computer interaction discrimination module specifically performs the following operations: Assign a brand-new and identical group number to the samples in the suspected merged data pair; Assign different team numbers to the samples corresponding to the two types of labels in the suspected merged data pair respectively.

14. The incremental clustering device according to claim 11, wherein, the human-computer interaction discrimination module specifically performs the following operations: Perform second group attribute assignment on the group attributes of the samples in the suspected merged data pair; Obtain the latest label of the sample and the sample corresponding to the latest label according to the samples in the suspected merged data pair; Check whether there is a conflict in the group attribute constraints of the sample corresponding to the latest label; If there is a conflict, do not merge; If there is no conflict, merge; Among them, the step of "performing second group attribute assignment on the group attributes of the samples in the suspected merged data pair" specifically includes: Assign a brand-new and identical group number to the samples of the suspected merged data pair; Assign the same team number to the samples of the suspected merged data pair.

15. The incremental clustering device according to claim 14, wherein, the human-computer interaction discrimination module specifically performs the following operations: Select all the samples with a group number not equal to 0; When the group numbers are the same and the team numbers are the same, the group attributes do not conflict; When the group numbers are the same but the team numbers are different, the group attributes conflict.

16. The incremental clustering device according to claim 9, wherein, the suspected merged data management module specifically performs the following operations: When the number of suspected merge data pairs obtained during the incremental clustering process reaches a preset quantity threshold, or when the cumulative statistical incremental clustering completion time during the incremental clustering process reaches a preset time threshold, the process proceeds to the category discrimination result processing flow.

17. A computer device, comprising a processor and a storage device, the storage device being adapted to store multiple program codes, characterized in that the program codes are adapted to be loaded and run by the processor to execute the incremental clustering method according to any one of claims 1 to 8.

18. A storage medium, the storage medium being adapted to store multiple program codes, characterized in that the program codes are adapted to be loaded and run by a processor to execute the incremental clustering method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Failure data management method based on cluster estimation

    KR102141391B1