Clustering Method, Device, Electronic Device and Computer Readable Medium for Target Image

Through feature extraction and clustering schemes, the transfer probability value is adjusted, combined with deep learning and data augmentation technology, the clustering of face images is optimized, and the problem of low clustering accuracy is solved, and more efficient and accurate clustering results are achieved.

CN112749668BActive Publication Date: 2025-07-08SHANGHAI MININGLAMP ARTIFICIAL INTELLIGENCE GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110064499.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-18
Publication Date
2025-07-08
Estimated Expiration
2041-01-18

AI Technical Summary

Technical Problem

The existing clustering algorithm has low accuracy in face classification, and it is impossible to effectively judge the clustering effect, resulting in the inability to improve clustering accuracy.

Method used

The feature matrix of the target image is obtained through the feature extraction model, the feature matrix is processed using the first clustering scheme, and the transfer probability value is adjusted to control the ratio of the isolated vector to the block within the preset numerical interval, combining deep learning and data enhancement technology to optimize the clustering results.

Benefits of technology

The quantification and accuracy of clustering results are achieved, the accuracy and efficiency of clustering are improved, the equipment occupation time is reduced, and the running time of deep learning algorithms is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112749668B_ABST
    Figure CN112749668B_ABST
Patent Text Reader

Abstract

The present application provides a clustering method, apparatus, electronic device, and computer-readable medium for a target image, belonging to the field of image technology. The method includes: inputting the acquired target image into a feature extraction model to obtain a feature matrix of the target image output by the feature extraction model; processing the feature matrix through a first clustering scheme to obtain isolated vectors in the feature matrix; in the case where the ratio of the number of isolated vectors to the number of blocks does not fall within a preset numerical range, adjusting the transition probability value in the first clustering scheme until the ratio of the number of isolated vectors to the number of blocks falls within the preset numerical range, where the transition probability value is applied in the process of processing the feature matrix. The present application improves the accuracy of clustering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image technology, and in particular, to a clustering method, device, electronic device, and computer-readable medium for target images. Background Art

[0002] Face clustering is a technical means for assigning IDs to each individual for unknown (unregistered) batch face data. In places where users visit frequently, such as public places like car 4S stores, restaurants, and hotels, surveillance videos can be used to collect the facial information of the visiting users, which is processed into feature vectors and saved as a feature matrix. Then, the features of the frequently visiting users are found from the feature matrix and assigned independent IDs. When the same ID user visits again, the system can quickly perceive and prepare in advance the services required by the user.

[0003] Currently, clustering algorithms are generally used to classify faces. Clustering analysis, also known as cluster analysis, is a statistical analysis method for studying (sample or index) classification problems and is also an important algorithm in data mining. Clustering algorithms classify faces based on the similarity of facial pictures of the same user. The facial features within the same cluster have more similarities than those not in the same cluster. However, since clustering algorithms belong to unsupervised learning, their clustering results are unknown, and it is impossible to judge the final clustering effect, so it is also impossible to improve the clustering accuracy. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a clustering method, device, electronic device, and computer-readable medium for target images to solve the problem of low clustering accuracy. The specific technical solutions are as follows:

[0005] In a first aspect, a clustering method for target images is provided. The method includes:

[0006] Input the obtained target images into a feature extraction model to obtain the feature matrix of the target images output by the feature extraction model, where the feature matrix contains the feature vectors of each target image;

[0007] Process the feature matrix through a first clustering scheme to obtain the isolated vectors in the feature matrix, where the processed feature matrix contains multiple blocks, each block includes at least one feature vector, each block represents a type of target image, each feature vector represents a target image, and the isolated vector means that there is exactly one vector in the block;

[0008] In the case where the ratio of the number of the isolated vectors to the number of the blocks does not fall within a preset numerical range, adjust the transition probability value in the first clustering scheme until the ratio of the number of the isolated vectors to the number of the blocks falls within the preset numerical range, where the transition probability value is applied in the process of processing the feature matrix.

[0009] Optionally, the process of obtaining the isolated vectors in the feature matrix by processing the feature matrix through the first clustering scheme includes:

[0010] Determine multiple nearest neighboring points of each feature vector in the feature matrix through a second clustering scheme;

[0011] Divide the feature matrix into different blocks according to the multiple nearest neighboring points to obtain a nearest neighbor matrix;

[0012] Process the nearest neighbor matrix through the first clustering scheme to obtain the isolated vectors in the nearest neighbor matrix.

[0013] Optionally, after adjusting the transition probability value in the first clustering scheme until the ratio of the number of the isolated vectors to the number of the blocks falls within the preset numerical range, the method further includes:

[0014] Generate a target matrix according to the adjusted isolated vectors and blocks, where the target matrix includes multiple blocks, and each block corresponds to at least one feature vector;

[0015] Store the blocks and the feature vectors corresponding to the blocks into a feature index library.

[0016] Optionally, after storing the blocks and the feature vectors corresponding to the blocks into the feature index library, the method further includes:

[0017] Determine the isolated vectors in the target matrix and the first images corresponding to the isolated vectors;

[0018] Perform data augmentation on the first images to obtain second images, and extract image vectors in the second images through a deep learning scheme;

[0019] Obtain a confidence level according to the image vectors and the feature vectors in the feature index library;

[0020] In the case where the confidence level is greater than a preset threshold, incorporate the isolated vectors into the corresponding blocks in the index library.

[0021] Optionally, the obtaining the confidence level according to the image vectors and the feature vectors in the feature index library includes:

[0022] Determine the distance between the image vector and the feature vectors in the feature index library;

[0023] Use the distance as the confidence level.

[0024] Optionally, after obtaining the confidence level based on the image vector and the feature vectors in the feature index library, the method further includes:

[0025] Discard the isolated vector when the confidence level is greater than the preset threshold.

[0026] Optionally, when the ratio of the number of isolated vectors to the number of blocks does not fall within a preset numerical range, adjusting the transition probability value in the first clustering scheme includes:

[0027] Increase the transition probability value in the first clustering scheme when the ratio of the number of isolated vectors to the number of blocks is less than the preset numerical range;

[0028] Decrease the transition probability value in the first clustering scheme when the ratio of the number of isolated vectors to the number of blocks is greater than the preset numerical range.

[0029] In a second aspect, a clustering device for target images is provided. The device includes:

[0030] An input module, configured to input the acquired target image into a feature extraction model to obtain a feature matrix of the target image output by the feature extraction model, where the feature matrix contains feature vectors of each target image;

[0031] A processing module, configured to process the feature matrix through a first clustering scheme to obtain isolated vectors in the feature matrix, where the processed feature matrix contains multiple blocks, each block includes at least one feature vector, each block represents a type of target image, each feature vector represents a target image, and the isolated vector represents that there is exactly one vector in the block;

[0032] An adjustment module, configured to adjust the transition probability value in the first clustering scheme when the ratio of the number of isolated vectors to the number of blocks does not fall within a preset numerical range until the ratio of the number of isolated vectors to the number of blocks falls within the preset numerical range, where the transition probability value is applied in the process of processing the feature matrix.

[0033] In a third aspect, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus. The processor, the communication interface, and the memory communicate with each other through the communication bus;

[0034] A memory for storing a computer program;

[0035] A processor for implementing any of the method steps when executing the program stored in the memory.

[0036] In a fourth aspect, a computer-readable storage medium is provided, in which a computer program is stored, and when the computer program is executed by a processor, any of the method steps is implemented.

[0037] Advantageous effects of the embodiments of the present application:

[0038] Through the relationship between the ratio and the preset numerical range, the present application performs continuous iterative clustering to quantify the clustering result, so that the clustering result meets the requirements and improves the accuracy of clustering.

[0039] Of course, when implementing any product or method of the present application, it is not necessarily required to achieve all the above-mentioned advantages at the same time. Description of the Drawings

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0041] Figure 1 Schematic diagram of the hardware environment of a clustering method for a target image provided by an embodiment of the present application;

[0042] Figure 2 Flowchart of a method for clustering a target image provided by an embodiment of the present application;

[0043] Figure 3 Schematic diagram of the clustering process of a target image provided by an embodiment of the present application;

[0044] Figure 4 Schematic diagram of the structure of a clustering device for a target image provided by an embodiment of the present application;

[0045] Figure 5 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed Embodiments

[0046] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Apparently, the described embodiments are some but not all of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.

[0047] In subsequent descriptions, the suffixes such as "module", "component", or "unit" used to represent elements are only for the convenience of description of this application, and they have no specific meaning in themselves. Therefore, "module" and "component" can be used interchangeably.

[0048] To solve the problems mentioned in the background art, according to one aspect of the embodiments of this application, an embodiment of a clustering method for target images is provided.

[0049] Optionally, in the embodiments of this application, the above-mentioned clustering method for target images can be applied to a Figure 1 hardware environment composed of a terminal 101 and a server 103 as shown. As Figure 1 shown, the server 103 is connected to the terminal 101 through a network and can be used to provide services for the terminal or the client installed on the terminal. A database 105 can be set up on the server or independently of the server to provide data storage services for the server 103. The above-mentioned network includes but is not limited to: wide area network, metropolitan area network, or local area network. The terminal 101 includes but is not limited to a PC, mobile phone, tablet computer, etc.

[0050] The clustering method for target images in the embodiments of this application can be executed by the server 103, or can be jointly executed by the server 103 and the terminal 101.

[0051] The embodiments of this application provide a clustering method for target images, which can be applied to a server and is used to accurately classify human images.

[0052] The following will combine specific implementation manners to detail the clustering method for target images provided by the embodiments of this application. As Figure 2 shown, the specific steps are as follows:

[0053] Step 201: Input the obtained target image into a feature extraction model to obtain the feature matrix of the target image output by the feature extraction model.

[0054] Among them, the feature matrix contains the feature vectors of each target image.

[0055] In an embodiment of the present application, the server continuously captures through a capturing device to obtain multiple target images within a period of time, and then inputs the target images into a feature extraction model to obtain a feature matrix of the target images output by the feature extraction model, where the feature matrix includes multiple feature vectors, and each feature vector represents a target image.

[0056] Exemplarily, the target image can be a face image. The server obtains multiple face images within a period of time through a capturing device, and then inputs the face images into a face feature extraction model to obtain a feature matrix of the face images output by the face feature extraction model. Each feature vector in the feature matrix represents each face image.

[0057] Step 202: Process the feature matrix through a first clustering scheme to obtain the isolated vectors in the feature matrix.

[0058] Wherein, the processed feature matrix contains multiple blocks, each block includes at least one feature vector, each block represents a type of target image, each feature vector represents a target image, and the isolated vector means that there is exactly one vector in the block.

[0059] The server processes the feature matrix through the first clustering scheme, specifically by clustering the feature vectors in the feature matrix to obtain multiple blocks, each block includes at least one feature vector, each block represents a type of target image, and each feature vector represents a target image.

[0060] If the target image is a face image, then each block represents a user ID, and one user ID can only represent one user. At least one feature vector in each block represents at least one face image of the same user ID. If there is exactly one vector in the block, then the feature vector is an isolated vector, that is, one user ID has only one face image.

[0061] Step 203: In the case where the ratio of the number of isolated vectors to the number of blocks does not fall within a preset numerical range, adjust the transition probability value in the first clustering scheme until the ratio of the number of isolated vectors to the number of blocks falls within the preset numerical range.

[0062] Wherein, the transition probability value is applied in the process of processing the feature matrix.

[0063] The number of isolated vectors is the number of target images where there is only one target image for one image type, and the number of blocks is the number of all types of target images. The server determines the ratio of the number of isolated vectors to the number of blocks, and then determines whether this ratio is within a preset numerical range. If the server determines that this ratio is within the preset numerical range, indicating that the number of isolated vectors is appropriate, a feature index library is established based on this feature matrix. Exemplarily, the preset numerical range is 0.4 - 0.5.

[0064] If the server determines that this ratio is not within the preset numerical range, indicating that the number of isolated vectors is too large or too small. If the ratio of the number of isolated vectors to the number of blocks is greater than the preset numerical range, it means that the number of target images occupied by one image type is too large, indicating that the image appears few times. If the ratio of the number of isolated vectors to the number of blocks is less than the preset numerical range, it means that the number of target images occupied by one image type is too small, indicating that there may be a clustering error for this image. The transfer probability value is used in the process of processing the feature matrix in the first clustering scheme, and this transfer probability value can be used to adjust the number of isolated vectors and blocks. Therefore, the server adjusts the transfer probability value in the first clustering scheme until the ratio of the number of isolated vectors to the number of blocks is within the preset numerical range.

[0065] Optionally, when the server determines that the ratio of the number of isolated vectors to the number of blocks is less than the preset numerical range, it increases the transfer probability value in the first clustering scheme; when the server determines that the ratio of the number of isolated vectors to the number of blocks is greater than the preset numerical range, it decreases the transfer probability value in the first clustering scheme. Exemplarily, the first clustering scheme can be the informap algorithm.

[0066] Exemplarily, an isolated vector means that one user ID contains only one target image, and a block represents the number of different user IDs. If the ratio of the number of isolated vectors to the number of blocks is greater than the preset numerical range, it indicates that the number of target images contained in one user ID is too large, and this user is an infrequent visitor, such as a passerby or a courier, etc., and there is no need to store and register such users. If the ratio of the number of isolated vectors to the number of blocks is less than the preset numerical range, it indicates that the number of target images contained in one user ID is too small. Generally speaking, within a certain period of time, the imaging device will obtain multiple images of most users. If the situation where a user ID has only one image is too few, there may be an image clustering error. For example, the only image of user A becomes one of the multiple images of user B, user A has no ID, and one of the images of user B contains an image that does not belong to user B. If the ratio of the number of isolated vectors to the number of blocks is within the preset numerical range, then the users represented by the blocks containing at least one feature vector are registered. These users are frequent visitors. When this user is detected again, it can be quickly sensed and a reminder can be sent to prepare in advance to provide the services required by this user.

[0067] In the present application, when it is detected that the ratio of the number of isolated vectors to the number of blocks is not within a preset numerical interval, the allocation of isolated vectors and blocks can be adjusted by adjusting the transition probability value to achieve clustering of feature vectors relative to blocks. The present application performs continuous iterative clustering based on the relationship between the ratio and the preset numerical interval to achieve quantification of clustering results, so that the clustering results meet the requirements and improve the accuracy of clustering.

[0068] As an optional implementation, the feature matrix is ​​processed by the first clustering scheme to obtain isolated vectors in the feature matrix, including: determining multiple nearest neighboring points of each feature vector in the feature matrix by the second clustering scheme; dividing the feature matrix into different blocks according to the multiple nearest neighboring points to obtain a nearest neighbor matrix; and processing the nearest neighbor matrix by the first clustering scheme to obtain isolated vectors in the nearest neighbor matrix.

[0069] In an embodiment of the present application, after the server obtains the feature matrix, it first determines each eigenvector in the feature matrix, and then determines the multiple nearest neighboring points of each eigenvector through a second clustering scheme. The types of the multiple neighboring points are consistent with the types of the feature vectors closest to the multiple neighboring points. Therefore, the server divides the feature matrix into different blocks according to the multiple nearest neighboring points. Each block includes the feature vectors of multiple nearest neighboring points. Each block represents the types of multiple nearest neighboring points. Multiple blocks constitute a nearest neighbor matrix. The accuracy of the feature vectors in the blocks obtained by the second clustering scheme is low. Therefore, the present application adopts the first clustering scheme to continue processing the nearest neighbor matrix, clusters and upgrades the feature vectors in the blocks, and organizes the division of the blocks to obtain isolated vectors in the nearest neighbor matrix.

[0070] Exemplarily, the second clustering scheme is the KNN algorithm. The server determines the K nearest neighbor points of each feature vector through the KNN algorithm, generates a nearest neighbor matrix through the K nearest neighbor points, and then processes the nearest neighbor matrix through the informap algorithm to obtain isolated vectors in the nearest neighbor matrix. The K value can be determined by the user's visit duration and the shooting interval, and this parameter has little effect on the clustering result.

[0071] Optionally, the server may also introduce a GPU (Graphics Processing Unit) for acceleration during the process of generating the nearest neighbor matrix, thereby processing larger amounts of data in a shorter period of time and improving the data processing rate.

[0072] In this application, the server first generates a nearest neighbor matrix through the second clustering scheme to obtain an initial clustering result, and then decomposes and reconstructs the nearest neighbor matrix through the first clustering scheme, optimizing the clustering result through a quantifiable clustering scheme, which improves the clustering efficiency and accuracy. This application can adjust the transfer probability value through an iterative clustering method to screen out isolated vectors, thereby enhancing the clustering ability and effect.

[0073] Compared with a single clustering process, this application obtains the clustered face categories through two clustering schemes, reduces the number of face categories, and improves the recall rate of the clustering result; compared with the clustering method of brute-force search, this application introduces the number of isolated vectors to provide a reference for the clustering iteration direction; compared with the method of directly comparing with a face recognition library, this application can perform deep learning feature extraction after establishing a feature index library, greatly reducing the running time of the deep learning algorithm and the device occupancy time.

[0074] As an alternative implementation, after adjusting the transfer probability value in the first clustering scheme until the ratio of the number of isolated vectors to the number of blocks is within a preset value range, the method further includes: generating a target matrix based on the adjusted isolated vectors and blocks, where the target matrix includes multiple blocks, and each block corresponds to at least one feature vector; storing the blocks and the feature vectors corresponding to the blocks in a feature index library.

[0075] After the server determines that the ratio of the number of isolated vectors to the number of blocks is within the preset value range, it indicates that the matrix clustering result is accurate. Then, a target matrix is generated through the blocks and the feature vectors included in each block. The clustering result of the target matrix is more accurate than that of the nearest neighbor matrix. The target matrix includes multiple blocks, and each block corresponds to at least one feature vector. The server establishes a correspondence between the blocks and the feature vectors included in the blocks, and then stores this correspondence in the feature index library.

[0076] Exemplarily, the server establishes a correspondence between the user category and the feature vectors of multiple images under this user category, and then stores this correspondence in the feature index library.

[0077] As an alternative implementation, after storing the blocks and the feature vectors corresponding to the blocks in the feature index library, the method further includes: determining the isolated vectors in the target matrix and the first images corresponding to the isolated vectors; performing data augmentation on the first images to obtain second images, and extracting image vectors from the second images through a deep learning scheme; obtaining a confidence level based on the image vectors and the feature vectors in the feature index library; and incorporating the isolated vectors into the corresponding blocks in the index library when the confidence level is greater than a preset threshold.

[0078] In an embodiment of the present application, for isolated vectors that cannot be clustered, the server determines the first image corresponding to the isolated vector, and then performs data augmentation on the first image to obtain a second image. Data augmentation includes means such as enhancing lighting, scaling, adding noise, and translation. Deep learning algorithms are used to extract the image features in the second image. Compared with the first image, the image features in the second image after data augmentation are more obvious.

[0079] The server determines the distance between the image vector in the second image and the feature vectors in the feature index library, and uses this distance as the confidence level. The distance is inversely proportional to the confidence level. The smaller the distance, the closer the image vector and the feature vectors in the feature index library are, and the higher the confidence level. Among them, the method for calculating the distance between vectors can adopt the Euclidean distance calculation method. The present application does not limit the specific distance calculation method.

[0080] The server determines whether the confidence level is greater than a preset threshold. If the server determines that the confidence level is greater than the preset threshold, it indicates that the confidence level is relatively high, and the image vector and the feature vectors in the feature index library are very similar. Then, the image vector can be incorporated into the block where the feature vector is located, that is, the isolated face corresponding to the image vector is incorporated into the corresponding user ID. If the server determines that the confidence level is not greater than the preset threshold, it indicates that the confidence level is relatively low, and the image vector and the feature vectors in the feature index library differ significantly. Then, the user corresponding to the image vector is a non-regular visitor, and the image vector can be discarded, that is, the non-regular visitor is not registered, avoiding excessive data volume in the feature index library and reducing data storage consumption, which is beneficial to improving the retrieval efficiency of the feature index library.

[0081] Optionally, the embodiment of the present application also provides a processing flow for clustering target images, as Figure 3 shown, and the specific steps are as follows.

[0082] 1. Input the obtained target image into the feature extraction model to obtain the feature matrix of the target image output by the feature extraction model.

[0083] 2. Use the KNN algorithm to determine the K nearest neighbor points of each feature vector to generate a nearest neighbor matrix.

[0084] 3. Process the nearest neighbor matrix through the infomap algorithm to obtain the isolated vectors in the nearest neighbor matrix.

[0085] 4. Determine whether the ratio of the number of isolated vectors to the number of blocks is within a preset numerical range. If so, execute step 5; if not, execute step 6.

[0086] 5. Establish a feature index library based on the feature matrix.

[0087] 6. Adjust the transition probability value in the first clustering scheme until the ratio of the number of isolated vectors to the number of blocks is within a preset numerical range.

[0088] 7. Perform data augmentation on the first image corresponding to the isolated vector to obtain a second image.

[0089] 8. Extract the image vector in the second image through a deep learning scheme, and obtain the confidence level based on the image vector and the feature vectors in the feature index library.

[0090] 9. Determine whether the confidence level is greater than a preset threshold. If so, execute step 10; if not, execute step 11.

[0091] 10. Incorporate the isolated vector into the corresponding block in the index library.

[0092] 11. Discard the image vector (isolated vector).

[0093] Among them, steps 4 and 7 are parallel steps.

[0094] Based on the same technical concept, an embodiment of the present application further provides a clustering device for a target image, as Figure 4 shown. The device includes:

[0095] An input module 401, configured to input the acquired target image into a feature extraction model to obtain a feature matrix of the target image output by the feature extraction model, where the feature matrix contains feature vectors of each target image;

[0096] A processing module 402, configured to process the feature matrix through a first clustering scheme to obtain isolated vectors in the feature matrix, where the processed feature matrix includes multiple blocks, each block includes at least one feature vector, each block represents a type of target image, each feature vector represents a target image, and the isolated vector means that there is exactly one vector in the block;

[0097] An adjustment module 403, configured to adjust the transition probability value in the first clustering scheme until the ratio of the number of isolated vectors to the number of blocks is within a preset numerical range when the ratio of the number of isolated vectors to the number of blocks is not within the preset numerical range, where the transition probability value is applied in the process of processing the feature matrix.

[0098] Optionally, the processing module 402 includes:

[0099] A first determination unit, configured to determine multiple nearest neighbor points of each feature vector in the feature matrix through a second clustering scheme;

[0100] A partitioning unit, configured to partition the feature matrix into different blocks according to the multiple nearest neighbor points to obtain a nearest neighbor matrix;

[0101] A processing unit for processing the nearest neighbor matrix through a first clustering scheme to obtain isolated vectors in the nearest neighbor matrix.

[0102] Optionally, the device further includes:

[0103] A generation module for generating a target matrix according to the adjusted isolated vectors and blocks, where the target matrix includes multiple blocks, and each block corresponds to at least one feature vector;

[0104] A storage module for storing the blocks and the feature vectors corresponding to the blocks into the feature index library.

[0105] Optionally, the device further includes:

[0106] A determination module for determining the isolated vectors in the target matrix and the first images corresponding to the isolated vectors;

[0107] An enhancement module for performing data enhancement on the first images to obtain second images, and extracting image vectors from the second images through a deep learning scheme;

[0108] An obtaining module for obtaining a confidence level according to the image vectors and the feature vectors in the feature index library;

[0109] An inclusion module for, when the confidence level is greater than a preset threshold, including the isolated vectors into the corresponding blocks in the index library.

[0110] Optionally, the obtaining module includes:

[0111] A second determination unit for determining the distance between the image vectors and the feature vectors in the feature index library;

[0112] A serving as unit for using the distance as the confidence level.

[0113] Optionally, the device further includes:

[0114] A discarding module for, when the confidence level is greater than a preset threshold, discarding the isolated vectors.

[0115] Optionally, the adjustment module 403 includes:

[0116] An increasing unit for, when the ratio of the number of isolated vectors to the number of blocks is less than a preset numerical range, increasing the transition probability value in the first clustering scheme;

[0117] A decreasing unit for, when the ratio of the number of isolated vectors to the number of blocks is greater than a preset numerical range, decreasing the transition probability value in the first clustering scheme.

[0118] According to another aspect of the embodiments of the present application, the present application provides an electronic device, such as Figure 5 as shown, which includes a memory 503, a processor 501, a communication interface 502, and a communication bus 504. A computer program that can run on the processor 501 is stored in the memory 503. The memory 503 and the processor 501 communicate through the communication interface 502 and the communication bus 504. When the processor 501 executes the computer program, the steps of the above method are implemented.

[0119] In the above electronic device, the memory and the processor communicate through a communication bus and a communication interface. The communication bus may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc.

[0120] The memory may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0121] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0122] According to still another aspect of the embodiments of the present application, there is also provided a computer-readable medium having non-volatile program code executable by a processor.

[0123] Optionally, in the embodiments of the present application, the computer-readable medium is configured to store program code for the processor to execute the above method:

[0124] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments, and will not be elaborated herein.

[0125] In the specific implementation of the embodiments of the present application, reference may be made to the above respective embodiments, and corresponding technical effects can be achieved.

[0126] It can be understood that the embodiments described herein can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in the present application, or a combination thereof.

[0127] For software implementation, the technologies described herein can be implemented by units that execute the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented within the processor or external to the processor.

[0128] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0129] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0130] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms.

[0131] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0132] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0133] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes. It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0134] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A clustering method for target images, characterized in that, The method includes: Input the obtained target image into a feature extraction model to obtain the feature matrix of the target image output by the feature extraction model, where the feature matrix contains the feature vectors of each target image; Process the feature matrix through a first clustering scheme to obtain the isolated vectors in the feature matrix, where the processed feature matrix contains multiple blocks, each block includes at least one feature vector, each block represents a type of target image, each feature vector represents a target image, and the isolated vector means that there is exactly one vector in the block; In the case where the ratio of the number of isolated vectors to the number of blocks does not fall within a preset numerical range, adjust the transition probability value in the first clustering scheme until the ratio of the number of isolated vectors to the number of blocks falls within the preset numerical range, where the transition probability value is applied in the process of processing the feature matrix; Among them, the process of obtaining the isolated vectors in the feature matrix by processing the feature matrix through the first clustering scheme includes: Determine multiple nearest neighbor points of each feature vector in the feature matrix through a second clustering scheme; Divide the feature matrix into different blocks according to the multiple nearest neighbor points to obtain a nearest neighbor matrix; Process the nearest neighbor matrix through the first clustering scheme to obtain the isolated vectors in the nearest neighbor matrix.

2. The method according to claim 1, characterized in that After adjusting the transition probability value in the first clustering scheme until the ratio of the number of isolated vectors to the number of blocks falls within the preset numerical range, the method further includes: Generate a target matrix according to the adjusted isolated vectors and blocks, where the target matrix includes multiple blocks, and each block corresponds to at least one feature vector; Store the blocks and the feature vectors corresponding to the blocks in a feature index library.

3. The method according to claim 2, wherein After obtaining the isolated vectors in the feature matrix, the method further includes: Determine the first image corresponding to the isolated vector; Perform data augmentation on the first image to obtain a second image, and extract the image vector in the second image through a deep learning scheme; Obtain a confidence level according to the image vector and the feature vectors in the feature index library; In the case where the confidence level is greater than a preset threshold, incorporate the isolated vector into the corresponding block in the feature index library.

4. The method according to claim 3, characterized in that, The obtaining the confidence level according to the image vector and the feature vectors in the feature index library includes: Determine the distance between the image vector and the feature vectors in the feature index library; Use the distance as the confidence level.

5. The method according to claim 3, characterized in that, After obtaining the confidence level according to the image vector and the feature vectors in the feature index library, the method further includes: In the case where the confidence level is not greater than the preset threshold, discard the isolated vector.

6. The method according to claim 1, characterized in that, The adjusting the transition probability value in the first clustering scheme in the case where the ratio of the number of isolated vectors to the number of blocks does not fall within a preset numerical range includes: When the ratio of the number of the isolated vectors to the number of the blocks is less than the preset numerical range, increase the transition probability value in the first clustering scheme; When the ratio of the number of the isolated vectors to the number of the blocks is greater than the preset numerical range, decrease the transition probability value in the first clustering scheme.

7. A clustering device for a target image, characterized in that, The device includes: An input module, configured to input the acquired target image into a feature extraction model to obtain a feature matrix of the target image output by the feature extraction model, where the feature matrix includes feature vectors of each of the target images; A processing module, configured to process the feature matrix through a first clustering scheme to obtain isolated vectors in the feature matrix, where the processed feature matrix includes a plurality of blocks, each block includes at least one feature vector, each block represents a type of the target image, each feature vector represents a target image, and the isolated vector means that there is exactly one vector in the block; An adjustment module, configured to adjust the transition probability value in the first clustering scheme until the ratio of the number of the isolated vectors to the number of the blocks is within the preset numerical range when the ratio of the number of the isolated vectors to the number of the blocks is not within the preset numerical range, where the transition probability value is applied in the process of processing the feature matrix; Wherein, the processing module is configured to: Determine a plurality of nearest neighbor points of each feature vector in the feature matrix through a second clustering scheme; Divide the feature matrix into different blocks according to the plurality of nearest neighbor points to obtain a nearest neighbor matrix; Process the nearest neighbor matrix through the first clustering scheme to obtain isolated vectors in the nearest neighbor matrix.

8. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used for storing a computer program; The processor is configured to implement the method steps according to any one of claims 1-6 when executing the program stored on the memory.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the method steps according to any one of claims 1-6 are implemented.

Citation Information

Patent Citations

  • Electromagnetic image extraction method and device, computer equipment and storage medium

    CN109727295A

  • Brain magnetic resonance image segmentation method and device based on unsupervised learning

    CN112233132A