Image recognition method and device, electronic equipment and storage medium

By performing image matting, feature extraction, clustering, and filtering on face authentication images, the system automatically identifies images with similar backgrounds, solving the problems of time-consuming, labor-intensive, and missed detections in manual spot checks in face authentication scenarios, and achieving efficient and intelligent violation identification.

CN116740392BActive Publication Date: 2026-05-19BEIJING 58 INFORMATION TTECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING 58 INFORMATION TTECH CO LTD
Filing Date
2023-06-29
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

In existing technologies, facial recognition systems suffer from issues such as multiple people authenticating under the same background and fraudulent check-ins. Furthermore, relying on manual sampling methods incurs significant labor costs and is prone to missed detections.

Method used

By performing background image extraction on face authentication images, extracting global and regional feature vectors, performing clustering and image filtering, and automatically identifying similar background images, automated recognition and push can be achieved.

Benefits of technology

It enables automated push notifications for background similarity recognition, saving labor costs, improving processing efficiency, avoiding missed detections of violations, and enhancing the accuracy and efficiency of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116740392B_ABST
    Figure CN116740392B_ABST
Patent Text Reader

Abstract

The application provides an image recognition method and device, electronic equipment and storage medium. The method comprises: performing cutout processing on a plurality of face authentication images in a first image set to obtain a second image set comprising a plurality of background images; performing feature extraction on the background images in the second image set to obtain a feature vector set comprising a global feature vector and at least one regional feature vector; clustering the first image set according to the regional feature vector to determine a plurality of image subsets; for each image subset, performing image filtering on the image subset based on the global feature vector to obtain a target image subset corresponding to each image subset, and the image background similarity of different images belonging to the same target image subset satisfies a first preset condition. The application can identify the background similarity in the overall data set, without a large amount of manual cost, and can avoid the problem of missing detection of illegal events caused by manual sampling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an image recognition method, apparatus, electronic device and storage medium. Background Technology

[0002] In recent years, deep learning algorithms in the field of image processing have developed rapidly, and their applications in various industries have become increasingly mature. Many real-life scenarios require authentication using selfie images, providing a prerequisite for the application of deep learning algorithms.

[0003] In facial recognition scenarios, the following violations may occur: multiple people authenticating under the same background may constitute fraudulent behavior; for attendance scenarios that require check-in at different locations (such as domestic service attendance), if one person checks in at the same location multiple times, it may constitute fraudulent check-in behavior.

[0004] Current methods for verifying facial recognition images with similar backgrounds primarily rely on manual sampling. This approach has the following drawbacks:

[0005] 1) The amount of data for face authentication is huge, and the image background recognition is relatively low. Using manual sampling to complete the image background similarity recognition will consume a lot of labor costs.

[0006] 2) The manual sampling inspection scheme requires checking not the entire set of images, but a subset of images that are sampled. Often, this sampled subset accounts for a small proportion of the entire set of images, which can lead to the omission of a large number of violations. Summary of the Invention

[0007] In view of the above problems, embodiments of this application provide an image recognition method, apparatus, electronic device, and storage medium that overcomes or at least partially solves the above problems.

[0008] In a first aspect, embodiments of this application provide an image recognition method, including:

[0009] Multiple face authentication images in the first image set are processed to remove background images, resulting in a second image set including multiple background images;

[0010] Feature extraction is performed on the background image in the second image set to obtain a feature vector set corresponding to the background image. The feature vector set includes a global feature vector and at least one region feature vector.

[0011] Based on the region feature vectors, the first image set is clustered to determine multiple image subsets;

[0012] For each image subset, image filtering is performed on the image subset based on the global feature vector corresponding to each image in the image subset to obtain the target image subsets corresponding to the multiple image subsets respectively. The background similarity of images corresponding to different images belonging to the same target image subset satisfies the first preset condition.

[0013] Secondly, embodiments of this application provide an image recognition device, comprising:

[0014] The processing and acquisition module is used to perform image cutout processing on multiple face authentication images in the first image set to obtain a second image set including multiple background images;

[0015] The acquisition module is used to extract features from the background image in the second image set and acquire a feature vector set corresponding to the background image. The feature vector set includes a global feature vector and at least one region feature vector.

[0016] The clustering determination module is used to cluster the first image set according to the region feature vector to determine multiple image subsets;

[0017] The filtering acquisition module is used to perform image filtering on each image subset based on the global feature vector corresponding to each image in the image subset, so as to obtain the target image subsets corresponding to the multiple image subsets respectively, and the background similarity of images corresponding to different images belonging to the same target image subset satisfies a first preset condition.

[0018] Thirdly, embodiments of this application provide an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the image recognition method as described in the first aspect above.

[0019] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the image recognition method described in the first aspect above.

[0020] The technical solution of this application embodiment obtains a second image set including multiple background images by performing background removal processing on the face authentication images in the first image set. Feature extraction is performed on the background images in the second image set to obtain the feature vector set corresponding to the background images. Based on the regional feature vectors in the feature vector set, the first image set is clustered to determine multiple image subsets. For each image subset, image filtering is performed based on the global feature vectors in the feature vector set to obtain the target image subset. This can identify background similarity situations in the overall data set without requiring a large amount of manual labor, and realizes automated push for background similarity recognition.

[0021] By identifying background similarities, violations related to similar backgrounds in different scenarios can be quickly resolved, saving manpower costs while improving processing efficiency. Furthermore, by recognizing the full amount of image data, the problem of missed violations caused by manual sampling can be avoided, enabling intelligent and automated image background similarity recognition and pushing of recognition results. Attached Figure Description

[0022] Figure 1 A schematic diagram illustrating the image recognition method provided in an embodiment of this application;

[0023] Figure 2a A schematic diagram illustrating the face authentication image provided in an embodiment of this application;

[0024] Figure 2b This is a schematic diagram illustrating the background image after image matting processing provided in the embodiments of this application;

[0025] Figure 3 This is a schematic diagram illustrating the region division of a background image provided in an embodiment of this application;

[0026] Figure 4 This diagram illustrates an overall implementation flowchart of the image recognition method provided in this application.

[0027] Figure 5 This is a schematic diagram illustrating the image recognition device provided in an embodiment of this application;

[0028] Figure 6 This is a schematic diagram of the electronic device structure provided in the embodiments of this application. Detailed Implementation

[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0030] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Multiple embodiments in this application may include two or more.

[0031] In the various embodiments of this application, it should be understood that the sequence number of each process described below does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0032] This application provides an image recognition method, see [link to relevant documentation]. Figure 1 As shown, it includes:

[0033] Step 101: Perform image cutout processing on multiple face authentication images in the first image set to obtain a second image set including multiple background images.

[0034] The image recognition method in this embodiment first obtains a first image set including multiple face authentication images, performs image matting processing on each face authentication image in the first image set to obtain the background image corresponding to the face authentication image, and determines a second image set corresponding to the first image set by aggregating multiple background images.

[0035] The first image set and the second image set correspond to the same number of images, and there is a one-to-one correspondence between the images in the first image set and the images in the second image set. For each face authentication image (which can be understood as the original image) in the first image set, there is a corresponding background image in the second image set. The background image can be regarded as the image obtained after removing the foreground object from the face authentication image.

[0036] Step 102: Extract features from the background image in the second image set to obtain a feature vector set corresponding to the background image. The feature vector set includes a global feature vector and at least one regional feature vector.

[0037] After performing image matting on the original images in the first image set and obtaining the second image set, feature extraction can be performed on each background image in the second image set to obtain the feature vector set corresponding to the background image.

[0038] When extracting features from a background image, the process includes global feature extraction and local feature extraction. Global feature vectors are obtained by global feature extraction, and regional feature vectors are obtained by local feature extraction. The set of feature vectors corresponding to the background image is determined based on the global feature vectors and regional feature vectors.

[0039] When performing local feature extraction on a background image, features can be extracted from one or more regions to obtain at least one region feature vector; and the region feature vector corresponding to the background image can be determined based on the richness of the image content corresponding to the background image, and the richness of the image content of the background image can be positively correlated with the number of region feature vectors corresponding to the background image.

[0040] Step 103: Based on the region feature vector, cluster the first image set to determine multiple image subsets.

[0041] After determining the feature vector set corresponding to each background image in the second image set, image clustering can be performed on the first image set based on the region feature vectors corresponding to the background images. Through image clustering, the first image set can be classified into categories, and multiple image subsets can be determined.

[0042] By performing image clustering on the first image set, K clusters can be obtained. Each cluster corresponds to a subset of images. Images in the same cluster have similar backgrounds. Through clustering, images with similar backgrounds in the first image set can be classified into one category.

[0043] Step 104: For each image subset, based on the global feature vector corresponding to each image in the image subset, perform image filtering on the image subset to obtain the target image subsets corresponding to the multiple image subsets respectively, and the background similarity of images corresponding to different images belonging to the same target image subset satisfies the first preset condition.

[0044] When multiple image subsets are determined through image clustering, image filtering can be performed on each image subset based on the global feature vectors corresponding to each image in the subset. This image filtering can obtain more accurate background similarity clusters and thus determine the target image subset corresponding to the image subset.

[0045] Among them, images belonging to the same image subset have similar backgrounds. By filtering the images in the image subset, images whose background similarity does not meet the condition can be removed. The background similarity of images corresponding to different images belonging to the same target image subset meets the first preset condition. Here, the first preset condition can be that the background similarity of images corresponding to different images is greater than a similarity threshold.

[0046] In this embodiment, each target image subset can be regarded as corresponding to an image category, and the backgrounds of images corresponding to the same image category are similar. By obtaining multiple target image subsets, the background similarity situation existing in the first image set can be determined, and the automatic push of background similarity recognition can be realized.

[0047] The above-described implementation scheme of this application obtains a second image set including multiple background images by performing background removal processing on the face authentication images in the first image set, extracting features from the background images in the second image set to obtain a set of feature vectors corresponding to the background images, clustering the first image set based on the regional feature vectors in the feature vector set to determine multiple image subsets, and filtering the images for each image subset based on the global feature vectors in the feature vector set to obtain the target image subset. This scheme can identify background similarity situations in the overall data set without requiring a large amount of manual labor, thus achieving automated push for background similarity recognition.

[0048] By identifying background similarities, violations related to similar backgrounds in different scenarios can be quickly resolved, saving manpower costs while improving processing efficiency. Furthermore, by recognizing the full amount of image data, the problem of missed violations caused by manual sampling can be avoided, enabling intelligent and automated image background similarity recognition and pushing of recognition results.

[0049] The process of obtaining a second image set based on a first image set is described below. This involves performing image cutout processing on multiple face authentication images in the first image set to obtain a second image set including multiple background images, including:

[0050] Obtain the first image set, which includes the multiple face authentication images to be reviewed in the target application scenario;

[0051] For each face authentication image, a human body matting algorithm is used to perform matting processing on the face authentication image to obtain the background image corresponding to the face authentication image.

[0052] In this embodiment, the first image set is the image set to be reviewed under the target application scenario. The first image set includes multiple face authentication images to be reviewed. The target application scenario can be a specific attendance scenario (an attendance scenario that requires clocking in at different locations), such as the attendance scenario of in-home housekeeping services. In this case, the first image set can be the newly added image data to be reviewed for a certain day. The target application scenario can also be a scenario where multiple people are authenticated under the same background, such as multiple people registering accounts under the same background.

[0053] Given a first image set, for each face authentication image in the first image set, a human body matting algorithm is used to perform matting processing on the face authentication image. This matting process is used to obtain the background image corresponding to the face authentication image. The purpose of removing the human image is to highlight key information in the background. See also... Figure 2a The image shown is the face authentication image before image cutout. (See attached image) Figure 2b The image shown is the background image after the image cutout process.

[0054] By processing the facial recognition image to be reviewed using a human body matting algorithm to remove the human figure, the background image can be obtained. This allows for the extraction of key information about the background, which in turn facilitates feature extraction and provides a basis for subsequent processing.

[0055] As an optional embodiment, when performing feature extraction on the background images in the second image set to obtain the feature vector set corresponding to the background images, the process includes:

[0056] The background image is divided into regions to obtain key regions corresponding to the background image, and the corresponding region feature vectors are determined based on the key regions, with each key region corresponding to a region feature vector.

[0057] Global feature extraction is performed on the background image to obtain the global feature vector corresponding to the background image.

[0058] After obtaining the second image set, feature extraction can be performed on each background image in the second image set to obtain a feature vector set corresponding to the background image, including regional feature vectors and global feature vectors.

[0059] When performing regional feature extraction and obtaining corresponding regional feature vectors for a background image, the background image can first be divided into regions according to a preset division rule. Key regions are then selected from the regions obtained through the region division. For each key region, feature extraction is performed using a feature extraction model to obtain the corresponding regional feature vector. The dimension of the regional feature vectors corresponding to each key region is the same, such as 512 dimensions. If a background image includes 3 key regions, then there will be a (3, 512) dimension regional feature vector.

[0060] When performing global feature extraction on a background image and obtaining the corresponding global feature vector, a feature extraction model can be used to extract features and obtain the processed image features. The obtained features are global features; that is, the feature extraction model can output a global feature vector. The panoramic feature vector and the region feature vector have the same dimensions, such as 512 dimensions.

[0061] Optionally, when dividing the background image into regions, obtaining key regions corresponding to the background image, and determining the corresponding region feature vectors based on the key regions, the process includes:

[0062] The background image is divided into a preset number of grids according to a preset division rule;

[0063] Based on the amount of background information, target grids that meet the requirements are selected from the preset number of grids, and the target grids are determined to be the key areas.

[0064] Regional features are extracted from the identified key regions to obtain the corresponding regional feature vectors.

[0065] For cases where the feature vector of a region is known, when dividing a background image into regions, the background image can be divided into a preset number of grids according to a preset division rule. Each grid corresponds to a region, and the background image can be evenly divided to obtain grids of the same size. For example, drawing a 3*3 uniform grid on the background image divides the background image into 9 regions. See [link to relevant documentation]. Figure 3 As shown.

[0066] After determining a preset number of grids, target grids that meet the requirements can be selected from these grids based on the amount of background information corresponding to each grid. A target grid is one whose corresponding background information content accounts for a proportion greater than a preset ratio. For example, the ratio of the background information content of the grid to the background image is greater than a preset ratio, or the proportion of background pixels in the grid is greater than a preset ratio. Each selected target grid is a corresponding key region; that is, there is a one-to-one correspondence between target grids and key regions.

[0067] The following example illustrates the process of filtering target grids and determining key regions. After dividing the area into nine regions (grids), regions with little background information are ignored, while regions with more background information are retained. Background information is determined by calculating the proportion of background pixels in the entire region; a higher proportion indicates more background information, and a lower proportion indicates less. By filtering out regions with less background information, the target grid can be obtained, thereby determining the key regions.

[0068] After identifying one or more key regions, feature extraction models can be used to extract regional features for each key region, thereby obtaining the regional feature vector corresponding to the current key region and thus obtaining one or more regional feature vectors.

[0069] The above implementation scheme achieves the acquisition of key regions with relatively rich information by dividing the background image into regions and selecting target grids from a preset number of grids based on the amount of background information. Feature extraction is then performed on the key regions to obtain regional feature vectors. By performing global feature extraction on the background image, global feature vectors representing the global features of the background image can be obtained, which facilitates subsequent processing based on the global feature vectors.

[0070] The following describes the process of image clustering and determining multiple image subsets. The step of clustering the first image set based on the region feature vectors to determine multiple image subsets includes:

[0071] A first image is determined in the first image set. Based on the region feature vector, the first image is matched with other images in the first image set that are different from the first image to obtain a second image whose background similarity with the first image meets the second preset condition. The first image is any image in the first image set.

[0072] Based on the first image and at least a portion of the second image, a subset of images is determined, and the determined subset of images is removed from the first image set to update the first image set and obtain the target image set;

[0073] The target image set is used as the first image set. The steps of determining a first image in the first image set and obtaining a second image whose background similarity with the first image meets the second preset condition are repeated until clustering is completed and the multiple image subsets are determined.

[0074] After obtaining the feature vector set corresponding to the background image, including the global feature vector and at least one regional feature vector, the first image set can be clustered based on the regional feature vector corresponding to the background image to determine multiple clusters (multiple image subsets). Images in the same cluster mean that the backgrounds are similar.

[0075] The specific process of clustering is described below:

[0076] For the first image set, an image is randomly selected from the first image set and designated as the first image. Since the images in the first image set correspond one-to-one with the images in the second image set, the region feature vector corresponding to the first image and the region feature vectors corresponding to other images in the first image set (images that are different from the first image) can be obtained. The number of region feature vectors corresponding to different images in the first image set can be different.

[0077] Based on the region feature vector, the first image is matched with other images in the first image set that are different from the first image, so as to obtain a second image whose background similarity with the first image meets the second preset condition through the matching of region feature vectors; the second image whose background similarity with the first image meets the second preset condition can be that the background similarity with the first image is greater than a set threshold.

[0078] After obtaining a second image whose background similarity to the first image meets the second preset condition based on the regional feature vector, an image subset (such as image subset 1) is determined based on at least some of the second images related to the first image and the first image, and image subset 1 is removed from the first image set to update the first image set and obtain the target image set.

[0079] After obtaining the target image set, the target image set is used as the first image set (which can be understood as the updated first image set). For the current first image set, a first image is determined again. Based on the region feature vector, the current first image is matched with other images in the current first image set that are different from the current first image to obtain a second image whose background similarity with the current first image meets the second preset condition. Based on the current first image and at least some of the second images in all the second images associated with the current first image, an image subset (e.g., image subset 2) is determined. Image subset 2 is removed from the current first image set to update the current first image set.

[0080] Then, for the updated first image set, return to the steps of determining a first image, obtaining a second image whose background similarity to the first image meets the second preset condition, and repeat the process of determining an image subset and updating the current first image set, and so on, until no new image subset can be formed, at which point the clustering stops, and the clustering process is completed.

[0081] Multiple image subsets obtained through clustering correspond to multiple clusters, and images within the same cluster imply background similarity. By performing image clustering on the first image set, background similarity in the overall dataset can be identified.

[0082] Optionally, the vector dimensions corresponding to each of the region feature vectors are the same; when matching the first image with other images in the first image set that are different from the first image based on the region feature vectors to obtain a second image whose background similarity to the first image meets the second preset condition, the process includes:

[0083] The first matrix is ​​determined based on the feature vectors of the N regions corresponding to the first image;

[0084] For each image in the first image set that is different from the first image, a second matrix is ​​determined based on the M region feature vectors corresponding to the image, where M and N are both integers greater than or equal to 1, and the value of M can be different for different images;

[0085] For each second matrix, a target matrix is ​​determined based on the result of multiplying the first matrix and the transpose of the second matrix, and the image corresponding to the current second matrix is ​​detected as the second image based on the target matrix;

[0086] Based on the detection results, a second image in the first image set whose background similarity to the first image meets the second preset condition is determined.

[0087] When obtaining a second image associated with a first image based on regional feature vectors, a first matrix corresponding to the first image can be determined based on the N regional feature vectors corresponding to the first image. For each image in the set of first images that is different from the first image, a second matrix can be determined based on the M regional feature vectors corresponding to the current image. The number of regional feature vectors corresponding to different images can be different, so the value of M can be different.

[0088] For each second matrix (each image in the first image set that is distinct from the first image), the target matrix can be determined based on the result of multiplying the first matrix with the transpose of the second matrix. Then, based on the target matrix corresponding to the current image, it can be detected whether the current image is a second image associated with the first image, and the detection result can be obtained.

[0089] Since each second matrix can correspond to a detection result, based on the detection results corresponding to multiple second matrices, a second image whose background similarity to the first image satisfies the second preset condition can be determined from the first image set.

[0090] For the target matrix, the number of its corresponding elements is determined based on M and N. Each element in the target matrix represents the matching degree between a key region in the first image and a key region in the current image (the image corresponding to the target matrix that is different from the first image). The matching degree here can be understood as similarity.

[0091] The process of determining the target matrix is ​​illustrated below with an example. The first image (Image 1) corresponds to 3 region feature vectors, each with 512 dimensions. Therefore, the matrix A corresponding to the first image is 3 rows and 512 columns. Image 2 corresponds to 4 region feature vectors, each with 512 dimensions. Therefore, the matrix B corresponding to Image 2 is 4 rows and 512 columns. Since the condition for matrix multiplication is that the number of columns in the first matrix equals the number of rows in the second matrix, matrix B is transposed. The resulting transposed matrix Bt is... TThe matrix has 521 rows and 4 columns, which satisfies the conditions for matrix multiplication. Then, we perform A*B... T After the operation, the target matrix is ​​obtained. The target matrix is ​​a 3x4 matrix with 12 elements. Each element represents the matching degree between a key region of the first image and a key region of the second image.

[0092] It should be noted that the process of matrix multiplication can be regarded as calculating the similarity distance (cosine similarity) between each pair of feature vectors in the region, and the larger the similarity distance, the greater the similarity between the vectors.

[0093] In the above implementation process, when determining the second image associated with the first image, the background similarity between images is obtained based on the matrix multiplication operation. Images that meet the conditions are selected based on the background similarity to determine the second image, thus realizing the selection of the second image based on the similarity between regional feature vectors.

[0094] The following describes the detection process based on the target matrix. When detecting whether the image corresponding to the current second matrix is ​​the second image based on the target matrix, the process includes:

[0095] Each element in the target matrix is ​​compared with a preset matching threshold.

[0096] If the number of elements greater than the preset matching threshold is greater than the first threshold, the image corresponding to the current second matrix is ​​determined to be the second image.

[0097] After determining the target matrix through matrix operations, each element in the target matrix is ​​compared with a preset matching threshold. Since the elements in the target matrix represent the matching degree (similarity) of key regions of two images, the matching status of the key regions of the two images can be obtained by comparing each element of the target matrix with the preset matching threshold.

[0098] If the number of elements is determined to be greater than the preset matching threshold by comparison, the determined number of elements is compared with the first threshold. If the determined number of elements is greater than the first threshold, it is determined that the current image corresponding to the target matrix has a high background similarity with the first image, and the current image is determined to be the second image.

[0099] In this embodiment, the preset matching threshold is a pre-defined threshold. If the value is greater than this threshold, the key region is considered to have matched successfully. Successful matching here can be understood as the background similarity of the key region meeting the requirements. By counting the number of elements greater than this threshold, the number of successfully matched key regions can be determined. Correspondingly, the first threshold is also a pre-defined extreme value. For example, if the first threshold is 2, then if the number of successfully matched elements is greater than 2, the backgrounds of the two images are considered similar.

[0100] The above implementation scheme compares each element in the target matrix with a preset matching threshold, determines the number of elements greater than the preset matching threshold, and when the determined number of elements is greater than a first threshold, determines that the image corresponding to the target matrix is ​​a second image with a background similar to the first image, thereby determining the background similarity of the image based on the background similarity of the key regions of the image.

[0101] The process of determining an image subset based on a first image and an associated second image is described below. Determining an image subset based on the first image and at least a portion of the second image includes:

[0102] When the number of corresponding second images is greater than the number of first images, the first number of second images with the highest background similarity are selected based on the order of background similarity from largest to smallest, and the image subset is determined based on the selected second images and the first images.

[0103] If the number corresponding to the second image is less than or equal to the first number, the image subset is determined based on the second image and the first image.

[0104] After determining the associated second images based on the first image, the number of determined second images is compared with a first number. If the number of determined second images is greater than the first number, based on background similarity as a parameter, the top-ranked first number of second images are selected from the determined second images in descending order of background similarity to obtain second images with high background similarity to the first image. Then, based on the selected second images and the first image, an image subset is determined. If the number of second images is less than or equal to the first number, the image subset can be directly determined based on the second image and the first image.

[0105] For example, if the first number is 20, and the number of determined second images is greater than 20, 20 second images with high background similarity to the first image can be selected according to background similarity (which can be key region background similarity or image background similarity); if the number of determined second images is no more than 20, a cluster can be directly determined.

[0106] The above implementation scheme, by comparing the determined number of second images with the first number, and based on the comparison, adopts an appropriate strategy to determine the image subset including the first and second images, which can ensure that the number of images in the image subset meets the requirements.

[0107] The following describes the process of filtering images in an image subset based on global feature vectors. As an optional embodiment, the image filtering of the image subset based on the global feature vectors corresponding to each image in the image subset includes:

[0108] For each image in the image subset, based on the global feature vector, the similarity distance between the current image and other images in the image subset that are different from the current image is calculated, so as to match the current image with other images pairwise, and obtain the background matching degree between the current image and other images based on the similarity distance;

[0109] Based on the background matching degree, target images are selected and filtered from the image subset, wherein the background matching degree between the target image and any image in the image subset is less than a second threshold.

[0110] After dividing the first image set into multiple image subsets through clustering, the global feature vector can be used as a parameter for each image subset to filter the image subset and obtain more accurate background similarity clusters.

[0111] When filtering images from a subset, for each image in the subset, based on the global feature vector, the similarity distance (background similarity distance) between the current image and other images in the subset that are distinct from the current image needs to be calculated. This allows for one-to-one matching of the current image with other images, determining the background matching degree (similarity) between the current image and other images. The similarity distance and background matching degree can be positively correlated. When determining the background matching degree based on the similarity distance, a mapping relationship can be used, or the similarity distance can be directly used as the background matching degree.

[0112] After calculating the similarity distance between each image and other images and determining the background matching degree for each image in the image subset, target images can be filtered out and removed based on the background matching degree. The background matching degree between the target image and any image in the image subset is less than the second threshold, thereby removing target images with low background similarity to other images in the image subset and obtaining a relatively accurate image subset.

[0113] The above implementation scheme obtains more accurate background similarity clusters by calculating the similarity distance between images in a subset of images based on global feature vectors, determining the background matching degree of the images based on the similarity distance, and filtering target images based on the background matching degree.

[0114] The image recognition method provided in this application is described below through an overall implementation process. (See attached document for details.) Figure 4 As shown, it includes:

[0115] Step 401: Determine the first image set to be reviewed in the target application scenario, which includes multiple face authentication images.

[0116] Step 402: For each face authentication image in the first image set, perform face masking processing on the face authentication image based on the human body masking algorithm to obtain the background image, so as to obtain the second image set.

[0117] Step 403: For the background image in the second image set, divide the background image into a preset number of grids according to the preset division rules. Based on the amount of background information, select the target grid that meets the requirements from the preset number of grids, determine the target grid as the key region, extract the region features of the determined key region, and obtain the region feature vector corresponding to the key region.

[0118] Step 404: For the background image in the second image set, perform global feature extraction on the background image to obtain the global feature vector corresponding to the background image.

[0119] Step 405: Determine a first image in the first image set; based on the region feature vector, match the first image with other images in the first image set that are different from the first image to obtain a second image whose background similarity with the first image meets the second preset condition; determine an image subset based on the first image and at least some of the second images; and remove the determined image subset from the first image set to update the first image set.

[0120] Step 406: For the updated first image set, repeat step 405 until no new image subset can be formed, in order to obtain multiple image subsets.

[0121] Step 407: For each image subset, perform image filtering based on the global feature vector to obtain a more accurate target image subset.

[0122] The above implementation process involves obtaining a background image by performing image matting on the original image, extracting features from the background image to obtain global and regional feature vectors, performing image clustering based on the regional feature vectors, and filtering the clustered image subsets based on the global feature vectors. This process can obtain a relatively accurate set of background-similar images to achieve accurate image recognition.

[0123] The implementation scheme of this application removes human figures from images using a matting algorithm, thereby highlighting the background information of the image and avoiding confusion with the foreground information; it selects key regions for feature extraction using a region segmentation method, so that the extracted features have rich background information, thus having strong recognition and discrimination; and it uses a lazy clustering algorithm to form initial clusters of background similarity based on the regional features of the image, and uses global image features for threshold matching within the initial clusters to eliminate falsely identified background similar images in the clusters, thereby obtaining more accurate background similarity clusters.

[0124] The above is the overall implementation scheme of the image recognition method provided in this application. By performing image matting processing on the face authentication images in the first image set to obtain a second image set including multiple background images, feature extraction is performed on the background images in the second image set to obtain the feature vector set corresponding to the background images, and based on the regional feature vectors in the feature vector set, the first image set is clustered to determine multiple image subsets. For each image subset, image filtering is performed based on the global feature vectors in the feature vector set to obtain the target image subset. This method can identify background similarity situations in the overall data set without requiring a large amount of manual labor, and achieves automated push for background similarity recognition.

[0125] By identifying background similarities, violations related to similar backgrounds in different scenarios can be quickly resolved, saving manpower costs while improving processing efficiency. Furthermore, by recognizing the full amount of image data, the problem of missed violations caused by manual sampling can be avoided, enabling intelligent and automated image background similarity recognition and pushing of recognition results.

[0126] Furthermore, by processing the face authentication image to be reviewed using a human body matting algorithm to remove the human figure, a background image can be obtained, which can then be used to extract key information that highlights the background. By dividing the background image into regions and filtering out target grids based on the amount of background information, key regions with relatively rich information can be obtained, and feature extraction can be performed on these key regions to obtain regional feature vectors. By performing global feature extraction on the background image, a global feature vector representing the global features of the background image can be obtained, which can then be used for subsequent processing based on the global feature vector.

[0127] By comparing each element in the target matrix with a preset matching threshold, the number of elements greater than the preset matching threshold is determined. When the determined number of elements is greater than a first threshold, the image corresponding to the target matrix is ​​determined to be a second image with a background similar to the first image. This achieves the determination of image background similarity based on the background similarity of key regions of the image. By comparing the number of determined second images with a first number, an appropriate strategy is adopted to determine an image subset including the first and second images, which can ensure that the number of images in the image subset meets the requirements.

[0128] By calculating the similarity distance between images in a subset of images based on global feature vectors, determining the background matching degree of the images based on the similarity distance, and filtering target images based on the background matching degree, a more accurate background similarity cluster can be obtained.

[0129] This application provides an image recognition device, see [link to relevant documentation]. Figure 5 As shown, it includes:

[0130] The processing and acquisition module 501 is used to perform image cutout processing on multiple face authentication images in the first image set to obtain a second image set including multiple background images;

[0131] The acquisition module 502 is used to extract features from the background image in the second image set and obtain a feature vector set corresponding to the background image. The feature vector set includes a global feature vector and at least one regional feature vector.

[0132] The clustering determination module 503 is used to cluster the first image set according to the region feature vector to determine multiple image subsets;

[0133] The filtering acquisition module 504 is used to perform image filtering on each image subset based on the global feature vector corresponding to each image in the image subset, so as to obtain the target image subsets corresponding to the multiple image subsets respectively, and the background similarity of images corresponding to different images belonging to the same target image subset satisfies the first preset condition.

[0134] Optionally, the processing and acquisition module includes:

[0135] The first acquisition submodule is used to acquire the first image set, which includes the multiple face authentication images to be reviewed in the target application scenario;

[0136] The processing and acquisition submodule is used to perform background image processing on each face authentication image based on the human body matting algorithm to obtain the background image corresponding to the face authentication image.

[0137] Optionally, the acquisition module includes:

[0138] The first processing submodule is used to divide the background image into regions, obtain the key regions corresponding to the background image, and determine the corresponding region feature vectors based on the key regions, with each key region corresponding to a region feature vector.

[0139] The second acquisition submodule is used to perform global feature extraction on the background image and obtain the global feature vector corresponding to the background image.

[0140] Optionally, the first processing submodule includes:

[0141] A partitioning unit is used to divide the background image into a preset number of grids according to a preset partitioning rule;

[0142] The first determining unit is used to filter out target grids that meet the requirements from the preset number of grids based on the amount of background information, and determine the target grids as the key areas;

[0143] The acquisition unit is used to extract regional features from the determined key regions and obtain the regional feature vectors corresponding to the key regions.

[0144] Optionally, the clustering determination module includes:

[0145] The second processing submodule is used to determine a first image in the first image set, and based on the region feature vector, match the first image with other images in the first image set that are different from the first image to obtain a second image whose background similarity with the first image meets a second preset condition. The first image is any image in the first image set.

[0146] The third processing submodule is used to determine an image subset based on the first image and at least a portion of the second image, and to remove the determined image subset from the first image set in order to update the first image set and obtain the target image set.

[0147] The fourth processing submodule is used to take the target image set as the first image set, return to the second processing submodule for further processing, and have the second processing module perform the following: determine a first image in the first image set, obtain a second image whose background similarity with the first image meets the second preset condition, until clustering is completed and the multiple image subsets are determined.

[0148] Optionally, the vector dimensions corresponding to the feature vectors of each region are the same; the second processing submodule includes:

[0149] The second determining unit is used to determine the first matrix based on the N region feature vectors corresponding to the first image;

[0150] The third determining unit is used to determine a second matrix for each image in the first image set that is different from the first image, based on the M region feature vectors corresponding to the image, where M and N are both integers greater than or equal to 1, and the value of M corresponding to different images may be different.

[0151] The processing unit is configured to determine a target matrix for each second matrix based on the result of multiplying the first matrix and the transpose of the second matrix, and to detect whether the image corresponding to the current second matrix is ​​the second image based on the target matrix;

[0152] The fourth determining unit is used to determine, based on the detection results, a second image in the first image set whose background similarity to the first image meets the second preset condition.

[0153] Optionally, the processing unit includes:

[0154] The comparison subunit is used to compare each element in the target matrix with a preset matching threshold.

[0155] A sub-unit is defined to determine the image corresponding to the current second matrix as the second image when the number of elements greater than the preset matching threshold is greater than the first threshold.

[0156] Optionally, the third processing submodule includes:

[0157] The filtering and determining unit is used to filter out the first number of second images that are ranked first based on the background similarity from largest to smallest when the number of second images corresponding to the second images is greater than the first number, and to determine the image subset based on the filtered second images and the first images;

[0158] The fifth determining unit is configured to determine the image subset based on the second image and the first image when the number corresponding to the second image is less than or equal to the first number.

[0159] Optionally, the filtering acquisition module includes:

[0160] The fifth processing submodule is used to calculate the similarity distance between the current image and other images in the image subset that are different from the current image, based on the global feature vector, for each image in the image subset, so as to match the current image with other images pairwise, and obtain the background matching degree between the current image and other images based on the similarity distance;

[0161] The filtering submodule is used to filter target images in the image subset based on the background matching degree, wherein the background matching degree between the target image and any image in the image subset is less than a second threshold.

[0162] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0163] This application also provides an electronic device, including: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the various processes of the above-described image recognition method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0164] For example, Figure 6 A schematic diagram of the physical structure of an electronic device is shown. (For example...) Figure 6 As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communication interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 can call logical instructions in the memory 630, and the processor 610 is used to perform the following steps: performing image matting processing on multiple face authentication images in a first image set to obtain a second image set including multiple background images; performing feature extraction on the background images in the second image set to obtain a set of feature vectors corresponding to the background images, the set of feature vectors including global feature vectors and at least one region feature vector; clustering the first image set according to the region feature vectors to determine multiple image subsets; for each image subset, performing image filtering on the image subset based on the global feature vectors corresponding to each image in the image subset to obtain target image subsets corresponding to the multiple image subsets respectively, and the background similarity of images corresponding to different images belonging to the same target image subset satisfies a first preset condition. The processor 610 can also execute other schemes in the embodiments of this application, which will not be further described here.

[0165] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0166] This application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described image recognition method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0167] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0168] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0169] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0170] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0171] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0172] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0173] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0174] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0175] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0176] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An image recognition method, characterized in that, include: Multiple face authentication images in the first image set are processed to remove background images, resulting in a second image set including multiple background images; Feature extraction is performed on the background image in the second image set to obtain a feature vector set corresponding to the background image. The feature vector set includes a global feature vector and at least one region feature vector. Based on the region feature vectors, the first image set is clustered to determine multiple image subsets; For each image subset, image filtering is performed on the image subset based on the global feature vector corresponding to each image in the image subset to obtain the target image subsets corresponding to the multiple image subsets respectively. The background similarity of images corresponding to different images belonging to the same target image subset satisfies the first preset condition. The step of extracting features from the background images in the second image set to obtain the feature vector set corresponding to the background images includes: The background image is divided into a preset number of grids according to a preset division rule; Based on the amount of background information, target grids that meet the requirements are selected from the preset number of grids, and the target grids are determined to be key areas. Regional features are extracted from the key regions to obtain the corresponding regional feature vectors, with each key region corresponding to a regional feature vector. Global feature extraction is performed on the background image to obtain the global feature vector corresponding to the background image.

2. The method according to claim 1, characterized in that, The step of performing image cutout processing on multiple face authentication images in the first image set to obtain a second image set including multiple background images includes: Obtain the first image set, which includes the multiple face authentication images to be reviewed in the target application scenario; For each face authentication image, a human body matting algorithm is used to perform matting processing on the face authentication image to obtain the background image corresponding to the face authentication image.

3. The method according to claim 1, characterized in that, The step of clustering the first image set based on the region feature vector to determine multiple image subsets includes: A first image is determined in the first image set. Based on the region feature vector, the first image is matched with other images in the first image set that are different from the first image to obtain a second image whose background similarity with the first image meets the second preset condition. The first image is any image in the first image set. Based on the first image and at least a portion of the second image, a subset of images is determined, and the determined subset of images is removed from the first image set to update the first image set and obtain the target image set; The target image set is used as the first image set. The steps of determining a first image in the first image set and obtaining a second image whose background similarity with the first image meets the second preset condition are repeated until clustering is completed and the multiple image subsets are determined.

4. The method according to claim 3, characterized in that, The vector dimensions corresponding to the feature vectors of each region are the same. The step of matching the first image with other images in the first image set that are different from the first image based on the region feature vector to obtain a second image whose background similarity to the first image meets a second preset condition includes: The first matrix is ​​determined based on the feature vectors of the N regions corresponding to the first image; For each image in the first image set that is different from the first image, a second matrix is ​​determined based on the M region feature vectors corresponding to the image, where M and N are both integers greater than or equal to 1, and the value of M can be different for different images; For each second matrix, a target matrix is ​​determined based on the result of multiplying the first matrix and the transpose of the second matrix, and the image corresponding to the current second matrix is ​​detected as the second image based on the target matrix; Based on the detection results, a second image in the first image set whose background similarity to the first image meets the second preset condition is determined.

5. The method according to claim 4, characterized in that, The step of detecting whether the image corresponding to the current second matrix is ​​the second image based on the target matrix includes: Each element in the target matrix is ​​compared with a preset matching threshold. If the number of elements greater than the preset matching threshold is greater than the first threshold, the image corresponding to the current second matrix is ​​determined to be the second image.

6. The method according to claim 3, characterized in that, Determining a subset of images based on the first image and at least a portion of the second image includes: When the number of corresponding second images is greater than the number of first images, the first number of second images with the highest background similarity are selected based on the order of background similarity from largest to smallest, and the image subset is determined based on the selected second images and the first images. If the number corresponding to the second image is less than or equal to the first number, the image subset is determined based on the second image and the first image.

7. The method according to claim 1, characterized in that, The step of filtering the image subset based on the global feature vectors corresponding to each image in the image subset includes: For each image in the image subset, based on the global feature vector, the similarity distance between the current image and other images in the image subset that are different from the current image is calculated, so as to match the current image with other images pairwise, and obtain the background matching degree between the current image and other images based on the similarity distance; Based on the background matching degree, target images are selected and filtered from the image subset, wherein the background matching degree between the target image and any image in the image subset is less than a second threshold.

8. An image recognition device, characterized in that, include: The processing and acquisition module is used to perform image cutout processing on multiple face authentication images in the first image set to obtain a second image set including multiple background images; The acquisition module is used to extract features from the background image in the second image set and acquire a feature vector set corresponding to the background image. The feature vector set includes a global feature vector and at least one region feature vector. The clustering determination module is used to cluster the first image set according to the region feature vector to determine multiple image subsets; The filtering acquisition module is used to perform image filtering on each image subset based on the global feature vector corresponding to each image in the image subset, so as to obtain the target image subsets corresponding to the multiple image subsets respectively, and the background similarity of images corresponding to different images belonging to the same target image subset satisfies the first preset condition. The acquisition module includes: The first processing submodule is used to divide the background image into regions, obtain the key regions corresponding to the background image, and determine the corresponding region feature vectors based on the key regions, with each key region corresponding to a region feature vector. The second acquisition submodule is used to perform global feature extraction on the background image and obtain the global feature vector corresponding to the background image; The first processing submodule includes: A partitioning unit is used to divide the background image into a preset number of grids according to a preset partitioning rule; The first determining unit is used to filter out target grids that meet the requirements from the preset number of grids based on the amount of background information, and determine the target grids as the key areas; The acquisition unit is used to extract regional features from the determined key regions and obtain the regional feature vectors corresponding to the key regions.

9. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the image recognition method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the steps of the image recognition method as described in any one of claims 1 to 7.