Vector clustering training method and device

Through a vector clustering training method, the cluster queue and process allocation mechanism are used to solve the problem of long training time and lock conflict in asynchronous Kmeans clustering training, and efficient clustering training and timely update of clustering results are achieved.

CN113298103BActive Publication Date: 2025-06-20ALIBABA CLOUD COMPUTING CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202010461281.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-05-27
Publication Date
2025-06-20
Estimated Expiration
2040-05-27

AI Technical Summary

Technical Problem

The existing asynchronous Kmeans clustering training method is triggered after the training vector reaches a certain scale. The training time is long and a large amount of IO reading data is generated, which affects efficiency. When synchronous Kmeans clustering training, writing vectors requires updating the training state, resulting in serious lock conflicts between multiple training processes, affecting efficiency.

Method used

A vector clustering training method is provided. By continuously receiving the vector to be clustered and determining the target cluster queue based on the vector identification, the vector to be clustered is added to the tail of the queue, and the clustering process is used to read the vector in first-in-first order for clustering training until the preset threshold is reached.

Benefits of technology

By allocating a separate clustering process for each cluster queue, the lock overhead problem during multi-process writing is solved, the computational efficiency of training is improved, and the clustering results are updated in time, solving the problem of long training time and lagging clustering results of asynchronous clustering training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113298103B_ABST
    Figure CN113298103B_ABST
Patent Text Reader

Abstract

The embodiments of this specification provide a vector clustering training method and apparatus. The vector clustering training method includes: continuously receiving vectors to be clustered, where the vectors to be clustered include vector identifiers; determining the target clustering queue of the vectors to be clustered according to the vector identifiers, and adding the vectors to be clustered to the end of the target clustering queue; using the clustering process corresponding to the target clustering queue to sequentially read the vectors to be clustered in the target clustering queue in the first-in, first-out order for clustering training, obtaining a clustering result, and counting the number of vectors to be clustered participating in the clustering training; and in the case where the number of vectors to be clustered participating in the clustering training reaches a preset threshold, determining the obtained clustering result as the clustering result of the target clustering queue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of computer technology, and particularly to a vector clustering training method. One or more embodiments of this specification also relate to a vector clustering training device, a computing device, and a computer-readable storage medium. Background Art

[0002] With the development of computer technology, Approximate Nearest Neighbor Search (ANN search) has also been widely applied to analytical databases with a multi-process architecture. ANN search is to quickly retrieve the nearest N adjacent vectors in a high-dimensional space through a pre-constructed index. This method can return approximately accurate results but cannot guarantee accuracy.

[0003] Kmeans clustering is a vector quantization method originating from signal processing and is also a classic clustering analysis method in the field of data mining. Kmeans clustering is a very important auxiliary tool in ANN search. Existing asynchronous Kmeans clustering training methods usually trigger after the training vectors reach a certain scale, and the training process lasts for a very long time. A large amount of IO data is read during training, affecting the usage efficiency. When directly obtaining the written vectors for synchronous Kmeans clustering training, each written vector needs to update the training status, resulting in very serious lock conflicts between multiple training processes and affecting the training efficiency.

[0004] Therefore, how to solve the above problems, improve the efficiency of Kmeans clustering training, and quickly obtain available cluster center vectors has become an urgent problem for current technicians to solve. Summary of the Invention

[0005] In view of this, the embodiments of this specification provide a vector clustering training method. One or more embodiments of this specification also relate to a vector clustering training device, a computing device, and a computer-readable storage medium to solve the technical defects existing in the prior art.

[0006] According to the first aspect of the embodiments of this specification, a vector clustering training method is provided, including:

[0007] Continuously receive vectors to be clustered, where the vectors to be clustered include vector identifiers;

[0008] Determine the target clustering queue of the vectors to be clustered according to the vector identifiers, and add the vectors to be clustered to the end of the target clustering queue;

[0009] Using the clustering process corresponding to the target clustering queue, read the vectors to be clustered in the target clustering queue in the first-in-first-out order one by one for clustering training, obtain the clustering result, and count the number of vectors to be clustered participating in the clustering training;

[0010] When the number of vectors to be clustered participating in the clustering training reaches a preset threshold, determine the obtained clustering result as the clustering result of the target clustering queue.

[0011] Optionally, determining the target clustering queue of the vector to be clustered according to the vector identifier includes:

[0012] Compare the vector identifier with the queue identifiers of each clustering queue in the preset clustering queue set, allocate corresponding numbers of clustering queues in the shared memory according to the preset queue configuration parameters, configure a corresponding clustering process for each clustering queue, and allocate shared caches for each clustering process in the shared memory;

[0013] When there is a queue identifier identical to the vector identifier, determine the clustering queue corresponding to the queue identifier as the target clustering queue.

[0014] Optionally, the method further includes:

[0015] When there is no queue identifier identical to the vector identifier, search for a clustering queue with an empty queue identifier in the clustering queue set;

[0016] When at least one clustering queue with an empty queue identifier is found, determine any one of the clustering queues with an empty queue identifier as the target clustering queue, and set the queue identifier of the target clustering queue to be the same as the vector identifier.

[0017] Optionally, the method further includes:

[0018] When no clustering queue with an empty queue identifier is found, return a prompt message to expand the preset queue configuration parameters.

[0019] Optionally, using the clustering process corresponding to the target clustering queue to read the vectors to be clustered in the target clustering queue in the first-in-first-out order one by one for clustering training includes:

[0020] Using the clustering process corresponding to the target clustering queue to read the vectors to be clustered in the target clustering queue in the first-in-first-out order;

[0021] Obtain the vector distances between the vector to be clustered and the centroid vectors of each cluster in the clustering result corresponding to the target clustering queue, and determine the centroid vector with the smallest vector distance as the target centroid vector, and the cluster corresponding to the target centroid vector is the target cluster;

[0022] Judge whether the vector distance between the vector to be clustered and the target centroid vector is greater than a preset threshold:

[0023] If so, construct a cluster of the clustering result corresponding to the target clustering queue with the vector to be clustered as the centroid vector;

[0024] If not, add the vector to be clustered to the target cluster and update the centroid vector of the target cluster;

[0025] Continue to execute the step of reading the vector to be clustered in the target clustering queue.

[0026] According to the second aspect of the embodiments of the present specification, there is provided a vector clustering training device, including:

[0027] A receiving module, configured to continuously receive vectors to be clustered, where the vectors to be clustered include vector identifiers;

[0028] An adding module, configured to determine the target clustering queue of the vector to be clustered according to the vector identifier, and add the vector to be clustered to the end of the target clustering queue;

[0029] A training module, configured to perform clustering training on the vectors to be clustered in the target clustering queue in the first-in, first-out order by using the clustering process corresponding to the target clustering queue, obtain a clustering result, and count the number of vectors to be clustered participating in the clustering training;

[0030] A determining module, configured to determine the obtained clustering result as the clustering result of the target clustering queue when the number of vectors to be clustered participating in the clustering training reaches a preset threshold.

[0031] Optionally, the adding module is further configured to compare the vector identifier with the queue identifiers of each clustering queue in a preset clustering queue set, allocate a corresponding number of clustering queues in the shared memory according to preset queue configuration parameters, configure a corresponding clustering process for each clustering queue, and allocate shared caches for each clustering process in the shared memory; when there is a queue identifier identical to the vector identifier, determine the clustering queue corresponding to the queue identifier as the target clustering queue.

[0032] Optionally, the device further includes:

[0033] A search module, configured to search for a clustering queue with an empty queue identifier in the set of clustering queues when there is no queue identifier identical to the vector identifier; when at least one clustering queue with an empty queue identifier is found, determine an arbitrary clustering queue with an empty queue identifier as the target clustering queue, and set the queue identifier of the target clustering queue to be the same as the vector identifier.

[0034] Optionally, the apparatus further includes:

[0035] A prompt module, configured to return a prompt message for expanding preset queue configuration parameters when no clustering queue with an empty queue identifier is found.

[0036] Optionally, the training module is further configured to use the clustering process corresponding to the target clustering queue to read the vectors to be clustered in the target clustering queue in a first-in-first-out order; obtain the vector distances between the vectors to be clustered and the centroid vectors of each clustering cluster in the clustering result corresponding to the target clustering queue, and determine the centroid vector with the smallest vector distance as the target centroid vector, and the clustering cluster corresponding to the target centroid vector is the target clustering cluster; determine whether the vector distance between the vector to be clustered and the target centroid vector is greater than a preset threshold: if so, construct a clustering cluster of the clustering result corresponding to the target clustering queue with the vector to be clustered as the centroid vector; if not, add the vector to be clustered to the target clustering cluster and update the centroid vector of the target clustering cluster; continue to execute the step of reading the vectors to be clustered in the target clustering queue.

[0037] According to a third aspect of the embodiments of the present specification, a computing device is provided, including:

[0038] A memory and a processor;

[0039] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions:

[0040] Continuously receive vectors to be clustered, where the vectors to be clustered include vector identifiers;

[0041] Determine the target clustering queue of the vector to be clustered according to the vector identifier, and add the vector to be clustered to the end of the target clustering queue;

[0042] Use the clustering process corresponding to the target clustering queue to sequentially read the vectors to be clustered in the target clustering queue in a first-in-first-out order for clustering training, obtain a clustering result, and count the number of vectors to be clustered participating in the clustering training;

[0043] When the number of vectors to be clustered participating in the clustering training reaches a preset threshold, determine the obtained clustering result as the clustering result of the target clustering queue.

[0044] According to the fourth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer-executable instructions that, when executed by a processor, implement the steps of the vector clustering training method.

[0045] In one embodiment of the present specification, by continuously receiving vectors to be clustered, where the vectors to be clustered include vector identifiers; determining the target clustering queue of the vectors to be clustered according to the vector identifiers, and adding the vectors to be clustered to the end of the target clustering queue; using the clustering process corresponding to the target clustering queue to sequentially read the vectors to be clustered in the target clustering queue in the first-in-first-out order for clustering training, obtaining a clustering result, and counting the number of vectors to be clustered participating in the clustering training; when the number of vectors to be clustered participating in the clustering training reaches a preset threshold, determine the obtained clustering result as the clustering result of the target clustering queue. Through this method, the write vector only needs to be added to the corresponding clustering queue, and each clustering queue is configured with a separate clustering process, so that each training process can be executed independently and efficiently in batches, solving the problem of lock overhead caused by multiple training processes accessing the write vector simultaneously in the case of multi-process writing. The computing efficiency of the training is improved. Description of the Drawings

[0046] Figure 1 is a processing flowchart of a vector clustering training method provided by an embodiment of the present specification;

[0047] Figure 2 is a flowchart of a clustering training method provided by an embodiment of the present specification;

[0048] Figure 3 is a processing process flowchart of a vector clustering training method provided by an embodiment of the present specification;

[0049] Figure 4 is a schematic structural diagram of a shared memory provided by an embodiment of the present specification;

[0050] Figure 5 is a schematic diagram of a vector clustering training device provided by an embodiment of the present specification;

[0051] Figure 6 is a structural block diagram of a computing device provided by an embodiment of the present specification. Detailed Embodiments

[0052] In the following description, numerous specific details are set forth to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.

[0053] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0054] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0055] First, the noun terms related to one or more embodiments of this specification are explained.

[0056] ANN retrieval: Approximate Nearest Neighbor Search is to quickly retrieve the nearest N adjacent vectors in a high-dimensional space through a pre-constructed index, but only approximate accurate results can be returned, and accuracy cannot be guaranteed.

[0057] Kmeans clustering: Kmeans clustering originated from a vector quantization method in signal processing and is also a classic clustering analysis method in the field of data mining.

[0058] In this specification, a vector clustering training method is provided. This specification also relates to a vector clustering training device, a computing device, and a computer-readable storage medium, which will be described in detail one by one in the following embodiments.

[0059] Figure 1 The processing flow chart of a vector clustering training method provided according to an embodiment of this specification is shown, including steps 102 to 108.

[0060] Step 102: Continuously receive the vectors to be clustered, where the vectors to be clustered include vector identifiers.

[0061] The vectors to be clustered are the vectors that need to be subjected to clustering training. The user can write the vectors to be clustered into the computing nodes or servers of the distributed system through an insertion statement. The method of writing the vectors to be clustered shall be subject to the actual situation. In this specification, the method for the user to write the vectors to be clustered is not limited.

[0062] The vector identifier of the vector to be clustered is used to identify the category of the vector to be clustered. The vector identifier of the vector to be clustered can be the algorithm identifier for generating the vector to be clustered, or the source identifier of the vector to be clustered, or the usage identifier of the vector to be clustered. In this application, the specific form of the vector identifier is not limited.

[0063] In a specific embodiment provided in this specification, taking the face recognition task as an example, the vector to be clustered is the feature vector obtained after the face image is subjected to feature extraction, and the vector identifier is the algorithm "A" for obtaining the vector to be clustered.

[0064] In another specific embodiment provided in this specification, taking the speech recognition task as an example, the vector to be clustered is the feature vector obtained after the speech information is subjected to feature extraction, and the vector identifier is the task type "speech recognition" of the vector to be clustered.

[0065] Step 104: Determine the target clustering queue of the vector to be clustered according to the vector identifier, and add the vector to be clustered to the end of the target clustering queue.

[0066] After receiving the vector to be clustered, determine the target clustering queue for storing the vector to be clustered according to the vector identifier of the vector to be clustered, and add the vector to be clustered to the end of the target clustering queue. At this time, the grouping of the vectors to be clustered is completed, and the next vector to be clustered is continuously received.

[0067] Optionally, determining the target clustering queue of the vector to be clustered according to the vector identifier includes: comparing the vector identifier with the queue identifiers of each clustering queue in the preset clustering queue set, allocating corresponding numbers of clustering queues in the shared memory according to the preset queue configuration parameters, configuring a corresponding clustering process for each clustering queue, and allocating shared caches for each clustering process in the shared memory; in the case where there is a queue identifier identical to the vector identifier, determining the clustering queue corresponding to the queue identifier as the target clustering queue.

[0068] In practical applications, users will preset queue configuration parameters in advance. The queue configuration parameters are used to determine the number of clustering queues to be trained simultaneously. According to the queue configuration parameters, the same number of clustering queues are allocated in the multi-process shared memory, and respective corresponding clustering processes are started for each clustering queue. Each clustering process corresponds to a clustering queue, and the clustering process is used to train the data in the corresponding clustering queue. A shared cache is allocated for each clustering process in the multi-process shared memory, and the shared cache can minimize the impact of the synchronous training of each clustering process on the writing process.

[0069] When allocating clustering queues, a corresponding queue identifier is set for each clustering queue. The initial value of the queue identifier of each clustering queue is empty. After the queue identifier of the clustering queue is bound to the vector identifier of the vector to be clustered, the clustering queue can only add vectors to be clustered with the same vector identifier as the queue identifier.

[0070] In a specific embodiment provided in this specification, taking three clustering queues in the clustering queue set as an example, namely: clustering queue 1 - queue identifier "A", clustering queue 2 - queue identifier "B", clustering queue 3 - queue identifier "C". According to the vector identifier "A" of the vector to be clustered 1, search and determine in the clustering queue set that the queue identifier of clustering queue 1 is the same as the vector identifier of the vector to be clustered 1, and then determine that clustering queue 1 is the target clustering queue of the vector to be clustered 1.

[0071] A clustering queue is a special linear list with restricted operations. The special feature is that the clustering queue only allows deletion operations at the front end of the list and insertion operations at the back end of the list. For a certain clustering queue, the vectors to be clustered belonging to the clustering queue are added to the tail of the clustering queue, and the corresponding clustering process of the clustering queue continuously obtains the vectors to be clustered from the head of the clustering queue using the data stream processing method.

[0072] Optionally, in the case where there is no queue identifier the same as the vector identifier, search for a clustering queue with an empty queue identifier in the clustering queue set; in the case where at least one clustering queue with an empty queue identifier is found, determine any one of the clustering queues with an empty queue identifier as the target clustering queue, and set the queue identifier of the target clustering queue to be the same as the vector identifier.

[0073] Specifically, if there is no clustering queue in the preset clustering queue set that has the same queue identifier as the vector identifier of the vector to be clustered, it indicates that there is no clustering queue bound to the vector identifier of the vector to be clustered in the clustering queue set. At this time, it is necessary to search for a clustering queue with an empty queue identifier in the clustering queue set. If there are at least two clustering queues with empty queue identifiers in the clustering queue set, any one of the clustering queues with an empty queue identifier is taken as the target clustering queue, and at the same time, the queue identifier of the target clustering queue is set to be the same as the vector identifier, binding the queue identifier and the vector identifier.

[0074] In a specific embodiment provided in this specification, taking the example that there are five clustering queues in the clustering queue set, namely: clustering queue 1 - queue identifier "A", clustering queue 2 - queue identifier "B", clustering queue 3 - queue identifier "C", clustering queue 4 - queue identifier "empty", clustering queue 5 - queue identifier "empty", and the vector identifier of the vector to be clustered 2 is "D". Search according to the vector identifier "D" in the clustering queue set. Since there is no clustering queue with the queue identifier "D" in the clustering queue set, continue to search for a clustering queue with the queue identifier "empty" in the clustering queue set. At this time, two clustering queues with the queue identifier "empty" are found, namely clustering queue 4 and clustering queue 5. Arbitrarily select clustering queue 4 as the target clustering queue from the two clustering queues, and set the queue identifier of clustering queue 4 to be the same as the vector identifier of the vector to be clustered 2, that is, set the queue identifier of clustering queue 4 to "D".

[0075] Optionally, in the case where no clustering queue with an empty queue identifier is found, a prompt message for expanding the preset queue configuration parameters is returned.

[0076] Specifically, in the case where no clustering queue with an empty queue identifier is found, it indicates that the number of currently allocated clustering queues is less than the number of types of vectors to be clustered, and the number of clustering queues needs to be increased. Then, a prompt message for expanding the preset queue configuration parameters is returned to prompt the user to expand the queue configuration parameters and increase the number of clustering queues to bind new vector identifiers.

[0077] In a specific embodiment provided in this specification, taking the case where there are five clustering queues in the clustering queue set as an example, they are respectively: clustering queue 1 - queue identifier "A", clustering queue 2 - queue identifier "B", clustering queue 3 - queue identifier "C", clustering queue 4 - queue identifier "D", clustering queue 5 - queue identifier "E", and the vector identifier of the vector to be clustered 3 is "F". After searching, there is no clustering queue with the queue identifier "F" in the clustering queue set, nor is there a clustering queue with the queue identifier "empty". Therefore, the vector to be clustered 3 cannot be added to the existing clustering queues, and a prompt of "There is no clustering queue that can be bound. Please add a new clustering queue" is returned to prompt the user to expand the preset queue configuration parameters and add a new clustering queue with an empty queue identifier to bind the vector identifier "F" of the vector to be clustered 3.

[0078] Step 106: Use the clustering process corresponding to the target clustering queue to sequentially read the vectors to be clustered in the target clustering queue in the first-in, first-out order for clustering training, obtain the clustering result, and count the number of vectors to be clustered participating in the clustering training.

[0079] Each clustering queue has a corresponding clustering process. Through the clustering process corresponding to the target clustering queue, according to the characteristics of the queue, the vectors to be clustered saved in the target clustering queue are sequentially read in the first-in, first-out order for clustering training. The vectors to be clustered in the clustering queue are clustered through a one-pass clustering algorithm, and the number of vectors to be clustered participating in the clustering training is counted. Each vector to be clustered only participates in one training calculation, and the number of participants in each round of iteration increases geometrically, that is, the i-th iteration training requires 2 i *k vectors to be clustered, where k is the number of clusters in the clustering result.

[0080] Optionally, refer to Figure 2 , Figure 2 shows the flowchart of the clustering training method provided in an embodiment of this specification, including steps 202 to 210.

[0081] Step 202: Use the clustering process corresponding to the target clustering queue to read the vectors to be clustered in the target clustering queue in the first-in, first-out order.

[0082] Use the clustering process corresponding to the target clustering queue to read the vectors to be clustered in the target clustering queue in the first-in, first-out order, and only obtain one vector to be clustered each time.

[0083] In an embodiment provided in this specification, taking the case where there are 10 vectors to be clustered in the target clustering queue as an example for explanation, the vectors to be clustered in the target clustering queue L1 are W1, W2,... W10 The vectors to be clustered are arranged in the order of the subscripts, and the vector to be clustered W1 at the head of the target clustering queue L1 is obtained.

[0084] Step 204: Obtain the vector distances between the vector to be clustered and the centroid vectors of each cluster in the clustering result corresponding to the target clustering queue, and determine the centroid vector with the smallest vector distance as the target centroid vector. The cluster corresponding to the target centroid vector is the target cluster.

[0085] Obtain the clustering result corresponding to the target clustering queue. The clustering result includes at least one cluster, and there is a center point, i.e., a centroid vector, in each cluster. Calculate the vector distances between the vector to be clustered and each centroid vector, and determine the centroid vector with the smallest vector distance as the target centroid vector. The cluster corresponding to the target centroid vector is the target cluster.

[0086] There are many methods for calculating the vector distance between the vector to be clustered and the centroid vector, such as Euclidean distance, Manhattan distance, Chebyshev distance, etc. In this specification, the method for calculating the vector distance is not limited.

[0087] In an embodiment provided in this specification, following the above example, obtain the clustering result corresponding to the target clustering queue L1. The clustering result includes three clusters C1, C2, and C3, and the centroid vectors of each cluster are K1, K2, and K3 respectively. Calculate the Euclidean distances between the vector to be clustered W1 and each centroid vector as r1, r2, and r3 respectively, where r2 > r3 > r1. Then K1 is the target centroid vector, and C1 is the target cluster.

[0088] It should be noted that when the vector to be clustered is the first vector to be clustered, there is no corresponding clustering result for the target clustering queue yet. Then, a cluster of the target clustering queue is constructed with the vector to be clustered as the centroid vector.

[0089] Step 206: Determine whether the vector distance between the vector to be clustered and the target centroid vector is greater than a preset threshold. If so, execute 208; if not, execute step 210.

[0090] The preset threshold is used to determine whether the vector to be clustered belongs to the target cluster. If the vector distance between the vector to be clustered and the target centroid vector is greater than the preset threshold, it means that the vector to be clustered does not belong to the target cluster, and execute step 208. If the vector distance between the vector to be clustered and the target centroid vector is less than or equal to the preset threshold, it means that the vector to be clustered belongs to the target cluster, and execute step 210.

[0091] In an embodiment provided in this specification, following the previous example, the Euclidean distance r1 between the vector W1 to be clustered and the centroid vector K1 of the target cluster C1 is less than the preset threshold, indicating that the vector W1 to be clustered belongs to the target cluster C1, and step 210 needs to be executed.

[0092] Taking the vector W5 to be clustered as an example for further explanation, the Euclidean distance between the vector W5 to be clustered and the centroid vector K3 of the target cluster C3 is greater than the preset threshold, indicating that the vector W5 to be clustered does not belong to the target cluster C3, and step 208 needs to be executed.

[0093] Step 208: Construct a cluster of the clustering result corresponding to the target clustering queue with the vector to be clustered as the centroid vector, and execute step 202.

[0094] Use the vector to be clustered as the centroid vector to create a new cluster for the target clustering queue, and continue to execute the step of taking the vector to be clustered in the target clustering queue.

[0095] In an embodiment provided in this specification, taking the vector W5 to be clustered as an example, use the vector W5 to be clustered as the centroid vector to create a new cluster C4 for the target clustering queue L1. At this time, the target clustering queue L1 includes 4 clusters C1, C2, C3, and C4, and continue to execute the step of taking the vector W6 to be clustered in the target clustering queue L1.

[0096] Step 210: Add the vector to be clustered to the target cluster, update the centroid vector of the target cluster, and execute step 202.

[0097] Add the vector to be clustered to the target cluster, update the centroid vector of the target cluster according to the vectors in the target cluster, and continue to execute the step of taking the vector to be clustered in the target clustering queue.

[0098] In an embodiment provided in this specification, taking the vector W1 to be clustered as an example, add the vector W1 to be clustered to the target cluster C1, and update the centroid vector K1 of the target cluster C1 according to the vectors to be clustered that already exist in the target cluster C1, and continue to execute the step of taking the vector W2 to be clustered in the target clustering queue L1.

[0099] After reading a vector to be clustered and performing clustering training each time, count the number of vectors to be clustered participating in the clustering training. In the clustering training, the number of vectors to be clustered participating in each round of iteration grows geometrically. The i-th iteration training requires 2 i *k vectors to be clustered, where k is the number of clusters in the clustering result.

[0100] Step 108: When the number of clustering vectors to be clustered participating in the clustering training reaches a preset threshold, determine the obtained clustering result as the clustering result of the target clustering queue.

[0101] When the number of clustering vectors to be clustered participating in the clustering training reaches a preset threshold, the obtained clustering result at this time is the clustering result of the target clustering queue. Through the clustering training method provided in this specification, after each data to be clustered participates in the clustering training, the clustering result will be updated in a timely manner, solving the problems of long training time and lagging clustering result in asynchronous clustering training.

[0102] In an embodiment provided in this specification, the preset threshold is 10,000 clustering vectors to be clustered. When the number of clustering vectors to be clustered participating in the clustering training reaches 10,000, the obtained clustering result includes 10 clustering clusters, and this clustering result is used as the clustering result of the target clustering queue.

[0103] An embodiment of this specification continuously receives clustering vectors to be clustered, where the clustering vectors to be clustered include vector identifiers; determines the target clustering queue of the clustering vectors to be clustered according to the vector identifiers, and adds the clustering vectors to be clustered to the end of the target clustering queue, so that the clustering vectors to be clustered can be assigned to their corresponding clustering queues. Each clustering queue is configured with a separate clustering process, enabling each training process to be executed batchwise, independently, and efficiently, solving the problem of lock overhead caused by multiple training processes accessing and writing vectors simultaneously in the case of multi-process writing. The clustering process corresponding to the target clustering queue reads the clustering vectors to be clustered in the target clustering queue in the first-in, first-out order to perform clustering training, obtains the clustering result, and counts the number of clustering vectors to be clustered participating in the clustering training; when the number of clustering vectors to be clustered participating in the clustering training reaches a preset threshold, determines the obtained clustering result as the clustering result of the target clustering queue, so that each clustering vector to be clustered only needs to participate in clustering training once and does not need to read historical data, expanding the data volume of clustering training and improving the computational efficiency of training.

[0104] Secondly, by updating the clustering result in a timely manner after each data to be clustered participates in the clustering training and updating the cluster center vector of the clustering result in real time, a more accurate cluster center vector can be obtained, solving the problems of long training time and lagging clustering result in asynchronous clustering training.

[0105] The following combines Figure 3 and Figure 4 , taking the application of the vector clustering training method provided in this specification in face recognition as an example, to further illustrate the vector clustering training method. Among them, Figure 3 shows the processing process flowchart of a vector clustering training method provided by an embodiment of this specification.Figure 4 It is a schematic diagram of the structure of the shared memory in a certain training node of a distributed system to which this method is applied. The specific steps include Step 302 to Step 322.

[0106] Step 302: Receive preset queue configuration parameters, and allocate corresponding numbers of clustering queues in the multi-process shared memory according to the queue configuration parameters, and start corresponding clustering processes for each clustering queue.

[0107] In the embodiment provided in this specification, the preset queue configuration parameter is 5. As Figure 4 shown, 5 clustering queues are allocated in the multi-process shared memory, and the size of each clustering queue is set to 10MB. At the same time, 5 clustering processes are started, and each clustering process corresponds to a clustering queue. Among them, the Wal cache is a cache area for temporarily storing database changes, and the content stored in the Wal cache will be written into the disk WAl file according to the requirements of the predefined time point parameters.

[0108] Step 304: Continuously receive the vectors to be clustered, where the vectors to be clustered include vector identifiers.

[0109] In the embodiment provided in this specification, through the vectorization process of the face image M, the corresponding vector to be clustered m is obtained, and the vector to be clustered m is written into the shared cache of the multi-process shared memory through the vector writing process. The shared cache is used to reduce disk I / O. The vector identifiers of the vectors to be clustered obtained through the vectorization process of the face image are all set to T.

[0110] Step 306: Determine the target clustering queue of the vector to be clustered according to the vector identifier, and add the vector to be clustered to the end of the target clustering queue.

[0111] In the embodiment provided in this specification, by traversing the clustering queue set through the vector identifier T of the vector to be clustered m, it is found that the queue identifier of clustering queue 3 is T, and clustering queue 3 is the target clustering queue of the vector to be clustered with the vector identifier T, and the vector to be clustered m is added to the end of clustering queue 3.

[0112] Step 308: Use the clustering process corresponding to the target clustering queue to sequentially read the vectors to be clustered in the target clustering queue in the first-in-first-out order.

[0113] In the embodiment provided in this specification, the vectors to be clustered in clustering queue 3 are a, b, c... m. The clustering process 3 corresponding to clustering queue 3 sequentially reads the vector to be clustered a, the vector to be clustered b,... the vector to be clustered m from clustering queue 3 in the order in the clustering queue.

[0114] Step 310: Obtain the vector distances between the vector to be clustered and the centroid vectors of each cluster in the clustering result corresponding to the target clustering queue, and determine the centroid vector with the minimum vector distance as the target centroid vector. The cluster corresponding to the target centroid vector is the target cluster.

[0115] In the embodiment provided in this specification, obtain the clustering result corresponding to the clustering queue 3. The clustering result includes 5 clusters, namely C1, C2, C3, C4, and C5. The centroid vectors of each cluster are K1, K2, K3, K4, and K5 respectively. Taking the vector a to be clustered as an example, calculate the Euclidean distances between the vector a to be clustered and each centroid vector respectively, and determine that the Euclidean distance from the centroid vector K5 is the smallest. That is, the target centroid vector of the vector a to be clustered is K5, and the target cluster is C5.

[0116] Step 312: Determine whether the vector distance between the vector to be clustered and the target centroid vector is greater than a preset threshold. If so, execute step 314; if not, execute step 316.

[0117] In the embodiment provided in this specification, taking the vector a to be clustered as an example, the Euclidean distance between the vector a to be clustered and the target centroid vector K5 is R a , which is greater than the preset threshold, so execute step 314.

[0118] In the embodiment provided in this specification, taking the vector d to be clustered as an example, the Euclidean distance between the vector d to be clustered and the target centroid vector K2 is R d , which is less than the preset threshold, so execute step 316.

[0119] Step 314: Use the vector to be clustered as the centroid vector to construct the cluster in the clustering result corresponding to the target clustering queue, and execute step 318.

[0120] In the embodiment provided in this specification, taking the vector a to be clustered as an example, since the Euclidean distance between the vector a to be clustered and the target centroid vector K5 is R a which is greater than the preset threshold, use the vector a to be clustered as the centroid vector of the new cluster, and construct the cluster K6 in the center of the clustering result corresponding to the target clustering queue. Then execute step 318.

[0121] Step 316: Add the vector to be clustered to the target cluster, update the centroid vector of the target cluster, and execute step 318.

[0122] In the embodiment provided in this specification, taking the vector d to be clustered as an example, since the Euclidean distance between the vector d to be clustered and the target centroid vector K2 is R dIf it is greater than the preset threshold, the vector d to be clustered is added to the target cluster C2, and the cluster center vector K2 of the target cluster C2 is updated according to the vectors to be clustered in the target cluster C2, and then step 318 is executed.

[0123] Step 318: Count the number of vectors to be clustered participating in the clustering training in the target clustering queue.

[0124] Count the number of vectors to be clustered participating in the clustering training. The number of vectors to be clustered participating in the training grows geometrically each time. The i-th iteration of training requires 2 i *k vectors to be clustered, where k is the number of clusters.

[0125] Step 320: Determine whether the number of vectors to be clustered participating in the clustering training is equal to the preset threshold. If so, execute step 322; if not, execute step 308.

[0126] Determine whether the number of vectors to be clustered participating in the clustering training is equal to the preset threshold. If so, the training ends and step 322 is executed; if not, continue to execute step 308.

[0127] Step 322: The obtained clustering result is determined as the clustering result of the target clustering queue.

[0128] Take the clustering result corresponding to the target clustering queue at this time as the clustering result of the target clustering queue, and complete the clustering training.

[0129] In one embodiment of the present specification, by continuously receiving vectors to be clustered, where the vectors to be clustered include vector identifiers; determining the target clustering queue of the vectors to be clustered according to the vector identifiers, and adding the vectors to be clustered to the end of the target clustering queue, the vectors to be clustered can be assigned to their corresponding clustering queues. Each clustering queue is configured with a separate clustering process, so that each training process can be executed batchwise and independently and efficiently, solving the problem of lock overhead caused by multiple training processes accessing and writing vectors simultaneously in the case of multi-process writing. Using the clustering process corresponding to the target clustering queue to sequentially read the vectors to be clustered in the target clustering queue from the target clustering queue in the first-in, first-out order for clustering training, obtaining the clustering result, and counting the number of vectors to be clustered participating in the clustering training; when the number of vectors to be clustered participating in the clustering training reaches the preset threshold, the obtained clustering result is determined as the clustering result of the target clustering queue, so that each vector to be clustered only needs to participate in clustering training once and does not need to read historical data, expanding the data volume of clustering training and improving the computational efficiency of training.

[0130] Secondly, by updating the clustering results in a timely manner after each piece of data to be clustered participates in the clustering training and updating the cluster center vectors of the clustering results in real time, more accurate cluster center vectors can be obtained, solving the problems of long training time and lagging clustering results in asynchronous clustering training.

[0131] Corresponding to the above method embodiments, this specification also provides embodiments of a vector clustering training device. Figure 5 The structural schematic diagram of a vector clustering training device provided by an embodiment of this specification is shown. As Figure 5 shown, the device includes:

[0132] A receiving module 502, configured to continuously receive vectors to be clustered, where the vectors to be clustered include vector identifiers.

[0133] An adding module 504, configured to determine the target clustering queue of the vectors to be clustered according to the vector identifiers, and add the vectors to be clustered to the end of the target clustering queue.

[0134] A training module 506, configured to use the clustering process corresponding to the target clustering queue to sequentially read the vectors to be clustered in the target clustering queue in the first-in-first-out order for clustering training, obtain clustering results, and count the number of vectors to be clustered participating in the clustering training.

[0135] A determining module 508, configured to, when the number of vectors to be clustered participating in the clustering training reaches a preset threshold, determine the obtained clustering results as the clustering results of the target clustering queue.

[0136] Optionally, the adding module 504 is further configured to compare the vector identifiers with the queue identifiers of each clustering queue in a preset clustering queue set, where a corresponding number of clustering queues are allocated in the shared memory according to preset queue configuration parameters, each clustering queue is configured with a corresponding clustering process, and a shared cache is allocated for each clustering process in the shared memory; in the case where there is a queue identifier identical to the vector identifier, the clustering queue corresponding to the queue identifier is determined as the target clustering queue.

[0137] Optionally, the device further includes:

[0138] A searching module, configured to, in the case where there is no queue identifier identical to the vector identifier, search for a clustering queue with an empty queue identifier in the clustering queue set; in the case where at least one clustering queue with an empty queue identifier is found, determine any one of the clustering queues with an empty queue identifier as the target clustering queue, and set the queue identifier of the target clustering queue to be the same as the vector identifier.

[0139] Optionally, the device further includes:

[0140] A prompt module, configured to return a prompt message for expanding preset queue configuration parameters when a clustering queue with an empty queue identifier is not found.

[0141] Optionally, the training module 506 is further configured to use the clustering process corresponding to the target clustering queue to read the vectors to be clustered in the target clustering queue in a first-in, first-out order; obtain the vector distances between the vectors to be clustered and the center vectors of each clustering cluster in the clustering result corresponding to the target clustering queue, and determine the center vector with the smallest vector distance as the target center vector, and the clustering cluster corresponding to the target center vector is the target clustering cluster; determine whether the vector distance between the vector to be clustered and the target center vector is greater than a preset threshold: if so, use the vector to be clustered as the center vector to construct a clustering cluster of the clustering result corresponding to the target clustering queue; if not, add the vector to be clustered to the target clustering cluster and update the center vector of the target clustering cluster; continue to execute the step of reading the vectors to be clustered in the target clustering queue.

[0142] The vector clustering training device provided by an embodiment of this specification continuously receives vectors to be clustered, where the vectors to be clustered include vector identifiers; determines the target clustering queue of the vectors to be clustered according to the vector identifiers, and adds the vectors to be clustered to the end of the target clustering queue, so that the vectors to be clustered can be assigned to their corresponding clustering queues. Each clustering queue is configured with a separate clustering process, enabling each training process to be executed independently and efficiently in batches, solving the problem of lock overhead caused by multiple training processes simultaneously accessing and writing vectors in the case of multi-process writing. Use the clustering process corresponding to the target clustering queue to sequentially read the vectors to be clustered in the target clustering queue in a first-in, first-out order for clustering training, obtain the clustering result, and count the number of vectors to be clustered participating in the clustering training; when the number of vectors to be clustered participating in the clustering training reaches a preset threshold, determine the obtained clustering result as the clustering result of the target clustering queue, so that each vector to be clustered only needs to participate in clustering training once and does not need to read historical data, expanding the data volume of clustering training and improving the calculation efficiency of training.

[0143] Secondly, by updating the clustering result in a timely manner after each vector to be clustered participates in the clustering training and updating the center vector of the clustering result in real time, a more accurate center vector can be obtained, solving the problems of long training time and lagging clustering result in asynchronous clustering training.

[0144] The above is a schematic solution of a vector clustering training device according to this embodiment. It should be noted that the technical solution of this vector clustering training device and the technical solution of the above vector clustering training method belong to the same concept. For the details not described in the technical solution of the vector clustering training device, reference can be made to the description of the technical solution of the above vector clustering training method.

[0145] Figure 6 FIG. 4 shows a block diagram of a computing device 600 according to an embodiment of the present specification. The components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 through a bus 630, and a database 650 is used to store data.

[0146] The computing device 600 further includes an access device 640, and the access device 640 enables the computing device 600 to communicate via one or more networks 660. Examples of these networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 640 may include one or more of any type of wired or wireless network interfaces (e.g., Network Interface Card (NIC)), such as IEEE802.11 Wireless Local Area Network (WLAN) wireless interface, Worldwide Interoperability for Microwave Access (Wi-MAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, Near Field Communication (NFC) interface, and so on.

[0147] In an embodiment of the present specification, the above components of the computing device 600 and Figure 6 other components not shown in FIG. 4 may also be connected to each other, for example, through a bus. It should be understood that Figure 6 the block diagram of the computing device shown in FIG. 4 is only for illustrative purposes and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed.

[0148] The computing device 600 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smart phones), wearable computing devices (e.g., smart watches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or PCs. The computing device 600 can also be a mobile or stationary server.

[0149] Wherein, the memory 610 is used to store computer-executable instructions, and the processor 620 is used to execute the following computer-executable instructions:

[0150] Continuously receive vectors to be clustered, where the vectors to be clustered include vector identifiers;

[0151] Determine the target clustering queue of the vectors to be clustered according to the vector identifiers, and add the vectors to be clustered to the end of the target clustering queue;

[0152] Use the clustering process corresponding to the target clustering queue to sequentially read the vectors to be clustered in the target clustering queue in the first-in, first-out order for clustering training, obtain the clustering result, and count the number of vectors to be clustered participating in the clustering training;

[0153] When the number of vectors to be clustered participating in the clustering training reaches a preset threshold, determine the obtained clustering result as the clustering result of the target clustering queue.

[0154] The above is a schematic solution of a computing device in this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above vector clustering training method belong to the same concept. For the details not described in the technical solution of the computing device, reference can be made to the description of the technical solution of the above vector clustering training method.

[0155] An embodiment of this specification also provides a computer-readable storage medium, which stores computer instructions that are executed by a processor to implement the steps of the vector clustering training method.

[0156] The above is a schematic solution of a computer-readable storage medium in this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above vector clustering training method belong to the same concept. For the details not described in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above vector clustering training method.

[0157] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.

[0158] The computer instructions include computer program code, which may be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, mobile hard disks, magnetic disks, optical disks, computer memories, read-only memories (ROMs), random access memories (RAMs), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0159] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the embodiments of this specification are not limited by the described action sequence, because according to the embodiments of this specification, certain steps may be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.

[0160] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0161] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. A vector clustering training method, comprising: Continuously receive vectors to be clustered, where the vectors to be clustered include vector identifiers, and the vector identifiers are used to identify the categories of the vectors to be clustered; Determine the target clustering queue of the vectors to be clustered according to the vector identifiers, and add the vectors to be clustered to the end of the target clustering queue. The target clustering queue is determined by comparing the vector identifiers with the queue identifiers of each of the multiple clustering queues in the clustering queue set. The vector identifiers of the vectors to be clustered added to the target clustering queue are the same as the queue identifier of the target clustering queue, and each clustering queue is configured with a separate clustering process; Use the clustering process corresponding to the target clustering queue to sequentially read the vectors to be clustered in the target clustering queue in the first-in-first-out order for clustering training, obtain the clustering result, and count the number of vectors to be clustered participating in the clustering training; When the number of vectors to be clustered participating in the clustering training reaches a preset threshold, determine the obtained clustering result as the clustering result of the target clustering queue.

2. The vector clustering training method according to claim 1, determining a target clustering queue of the vector to be clustered according to the vector identifier, comprising: Compare the vector identifiers with the queue identifiers of each clustering queue in the preset clustering queue set. According to the preset queue configuration parameters, allocate a corresponding number of clustering queues in the shared memory. Each clustering queue is configured with a corresponding clustering process, and a shared cache is allocated for each clustering process in the shared memory; When there is a queue identifier identical to the vector identifier, determine the clustering queue corresponding to the queue identifier as the target clustering queue.

3. The vector clustering training method according to claim 2, further comprising: When there is no queue identifier identical to the vector identifier, search for a clustering queue with an empty queue identifier in the clustering queue set; When at least one clustering queue with an empty queue identifier is found, determine any one of the clustering queues with an empty queue identifier as the target clustering queue, and set the queue identifier of the target clustering queue to be the same as the vector identifier.

4. The vector clustering training method according to claim 3, further comprising: When no clustering queue with an empty queue identifier is found, return a prompt message to expand the preset queue configuration parameters.

5. The vector clustering training method according to claim 1, using a clustering process corresponding to the target clustering queue to sequentially read vectors to be clustered in the target clustering queue in a first-in-first-out order for clustering training, comprising: Use the clustering process corresponding to the target clustering queue to read the vectors to be clustered in the target clustering queue in the first-in-first-out order; Obtain the vector distances between the vector to be clustered and the centroid vectors of each clustering cluster in the clustering result corresponding to the target clustering queue, and determine the centroid vector with the smallest vector distance as the target centroid vector. The clustering cluster corresponding to the target centroid vector is the target clustering cluster; Judge whether the vector distance between the vector to be clustered and the target centroid vector is greater than a preset threshold: If so, use the vector to be clustered as the centroid vector to construct the clustering cluster of the clustering result corresponding to the target clustering queue; If not, add the vector to be clustered to the target clustering cluster and update the centroid vector of the target clustering cluster; Continue to execute the step of reading the vectors to be clustered in the target clustering queue.

6. A vector clustering training device, comprising: A receiving module, configured to continuously receive vectors to be clustered, where the vectors to be clustered include vector identifiers, and the vector identifiers are used to identify the categories of the vectors to be clustered; An adding module, configured to determine a target clustering queue of the vector to be clustered according to the vector identifier, and add the vector to be clustered to the end of the target clustering queue, where the target clustering queue is determined by comparing the vector identifier with the queue identifier of each clustering queue in a set of clustering queues, the vector identifier of the vector to be clustered added to the target clustering queue is the same as the queue identifier of the target clustering queue, and each clustering queue is configured with a separate clustering process; A training module, configured to use the clustering process corresponding to the target clustering queue to sequentially read the vectors to be clustered in the target clustering queue in a first-in, first-out order for clustering training, obtain a clustering result, and count the number of vectors to be clustered participating in the clustering training; A determining module, configured to, when the number of vectors to be clustered participating in the clustering training reaches a preset threshold, determine the obtained clustering result as the clustering result of the target clustering queue.

7. The vector clustering training device according to claim 6, The adding module is further configured to compare the vector identifier with the queue identifier of each clustering queue in a preset clustering queue set, wherein, Allocate a corresponding number of clustering queues in the shared memory according to preset queue configuration parameters, configure a corresponding clustering process for each clustering queue, and allocate shared caches for each clustering process in the shared memory; When there is a queue identifier identical to the vector identifier, determine the clustering queue corresponding to the queue identifier as the target clustering queue.

8. The vector clustering training device according to claim 7, further comprising: A searching module, configured to, when there is no queue identifier identical to the vector identifier, search for a clustering queue with an empty queue identifier in the set of clustering queues; when at least one clustering queue with an empty queue identifier is found, determine any one of the clustering queues with an empty queue identifier as the target clustering queue, and set the queue identifier of the target clustering queue to be the same as the vector identifier.

9. The vector clustering training device according to claim 8, further comprising: A prompting module, configured to, when no clustering queue with an empty queue identifier is found, return a prompt message for expanding the preset queue configuration parameters.

10. The vector clustering training device according to claim 6, The training module is further configured to read the vectors to be clustered in the target clustering queue from the target clustering queue in the first-in-first-out order by using the clustering process corresponding to the target clustering queue; obtain the vector distances between the vectors to be clustered and the centroid vectors of each clustering cluster in the clustering result corresponding to the target clustering queue, and determine the centroid vector with the smallest vector distance as the target centroid vector, and the clustering cluster corresponding to the target centroid vector is the target clustering cluster; determine whether the vector distance between the vector to be clustered and the target centroid vector is greater than a preset threshold: if so, construct a clustering cluster of the clustering result corresponding to the target clustering queue with the vector to be clustered as the centroid vector; if not, add the vector to be clustered to the target clustering cluster and update the centroid vector of the target clustering cluster; continue to execute the step of reading the vectors to be clustered in the target clustering queue.

11. A computing device, comprising: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions: Continuously receive vectors to be clustered, where the vector to be clustered includes a vector identifier, and the vector identifier is used to identify the category of the vector to be clustered; Determine a target clustering queue of the vector to be clustered according to the vector identifier, and add the vector to be clustered to the end of the target clustering queue, where the target clustering queue is determined by comparing the vector identifier with the queue identifier of each clustering queue in a set of clustering queues, the vector identifier of the vector to be clustered added to the target clustering queue is the same as the queue identifier of the target clustering queue, and each clustering queue is configured with a separate clustering process; Use the clustering process corresponding to the target clustering queue to sequentially read the vectors to be clustered in the target clustering queue in a first-in, first-out order for clustering training, obtain a clustering result, and count the number of vectors to be clustered participating in the clustering training; When the number of vectors to be clustered participating in the clustering training reaches a preset threshold, the obtained clustering result is determined as the clustering result of the target clustering queue.

12. A computer-readable storage medium storing computer instructions, which when executed by a processor, implement the steps of the vector clustering training method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Processing system and processing method of large-scale concurrent data stream

    CN102200906A

  • Clustering implementation method and apparatus

    CN106778812A

  • Word multi-prototype vector representation and word sense disambiguation method based on CRP clustering

    CN109033307A

  • A method and apparatus for clustering log streams

    CN109388711A