A file creation method, device, electronic device and computer-readable storage medium

By obtaining and utilizing the feature space of the picture, determining the target picture and selecting the file-building picture, the problem of low file-building quality in the existing technology of classifiers in selecting file-building pictures is solved, and a higher file-building quality is achieved.

CN114140642BActive Publication Date: 2025-06-10SHENZHEN XUMI YUNTU SPACE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111361233.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-17
Publication Date
2025-06-10
Estimated Expiration
2041-11-17

AI Technical Summary

Technical Problem

In the prior art, the use of classifiers to select file-building pictures leads to poor file-building quality and lacks effective solutions.

Method used

By obtaining the first feature space and the second feature space of the collected pictures, the space within the first category and the space between the second category are described, the target picture is determined and the filed picture is selected.

Benefits of technology

The accuracy of selection of file-building pictures has been improved, and the quality of file-building is improved, avoiding the problem of low file-building quality caused by classifier methods in the existing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114140642B_ABST
    Figure CN114140642B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of image recognition technology, and provides a method, an apparatus, an electronic device, and a computer-readable storage medium for creating a file. The method includes: respectively obtaining a first feature space and a second feature space corresponding to the collected images, where the first feature space is used to describe the space of the image within a first category, the second feature space is used to describe the space between the image and a second category, the second category is a category other than the first category, and the first category is the category corresponding to the image; determining a target image from the collected images according to the first feature space and the second feature space; and determining a file-creation image according to the target image. By means of the present disclosure, the problem in the prior art that the quality of file creation is poor due to the use of a classifier to select the file-creation image is solved, and the technical effect of improving the quality of file creation is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image recognition technology, and in particular to a filing method, device, electronic device and computer-readable storage medium. Background Art

[0002] In the existing face archiving, for the selection of face images for archiving, the first solution is to rely on the detection score to sort and select the one with the highest detection score; the second solution is to rely on multiple classification or regression models (occlusion classification, front and side face classification, brightness classification, pose regression) and select the one with the highest score in multiple models.

[0003] These technical solutions are actually some manually designed classifiers, and their scores are related to manually designed supervision signals (such as detection frame signals, occlusion signals, etc.), but have no direct relationship with the face space. They do not directly model the quality based on the distribution of the face space. Therefore, when the accuracy of face image selection is low, the quality of face archiving will be poor. In response to this technical problem, the relevant technologies have not proposed an effective solution. Summary of the invention

[0004] In view of this, the embodiments of the present disclosure provide a filing method, device, electronic device and computer-readable storage medium to solve the problem in the prior art of simply using a classifier to select filing pictures resulting in poor filing quality.

[0005] According to a first aspect of an embodiment of the present disclosure, a method for creating an archive is provided, comprising: respectively acquiring a first feature space and a second feature space corresponding to a collected picture, wherein the first feature space is used to describe the space of the picture within a first category, and the second feature space is used to describe the space between the picture and a second category, the second category is a category other than the first category, and the first category is a category corresponding to the picture; determining a target picture from the collected pictures according to the first feature space and the second feature space; and determining an archive picture according to the target picture.

[0006] According to a second aspect of an embodiment of the present disclosure, a filing device is provided, comprising: an acquisition module, for respectively acquiring a first feature space and a second feature space corresponding to a collected picture, wherein the first feature space is used to describe a space of the picture within a first category, and the second feature space is used to describe a space between the picture and a second category, the second category being a category other than the first category, and the first category being a category corresponding to the picture; a first determination module, for determining a target picture from the collected pictures according to the first feature space and the second feature space; and a second determination module, for determining a filing picture according to the target picture.

[0007] In a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where the processor implements the steps of the above method when executing the computer program.

[0008] In a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing a computer program, where the computer program implements the steps of the above method when executed by a processor.

[0009] The beneficial effects of the embodiments of the present disclosure compared with the prior art are as follows: by respectively obtaining a first feature space and a second feature space corresponding to the collected pictures, where the first feature space is used to describe the space of the picture within the first category, the second feature space is used to describe the space between the picture and the second category, the second category is a category other than the first category, and the first category is the category corresponding to the picture; determining a target picture from the collected pictures according to the first feature space and the second feature space; and determining an archived picture according to the target picture. When selecting an archived picture in the embodiments of the present disclosure, it is based on the picture space estimation, rather than simply using a classifier as in the prior art, thereby solving the problem of poor archiving quality caused by using a classifier to select archived pictures in the prior art and achieving the technical effect of improving the archiving quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0011] Figure 1 is a schematic diagram of the application scenario of the embodiments of the present disclosure;

[0012] Figure 2 is a schematic flowchart of an archiving method provided by the embodiments of the present disclosure;

[0013] Figure 3 is a schematic flowchart of another archiving method provided by the embodiments of the present disclosure;

[0014] Figure 4 is a schematic structural diagram of an archiving device provided by the embodiments of the present disclosure;

[0015] Figure 5 is a schematic structural diagram of an electronic device provided by the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] In the following description, specific details such as specific system architectures and technologies are presented for illustration rather than limitation in order to provide a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from obstructing the description of the present disclosure.

[0017] A file creation method and apparatus according to an embodiment of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0018] Figure 1 It is a schematic diagram of an application scenario of an embodiment of the present disclosure. The application scenario may include terminal devices 101, 102, and 103, a server 104, and a network 105.

[0019] The terminal devices 101, 102, and 103 may be hardware or software. When the terminal devices 101, 102, and 103 are hardware, they may be various electronic devices with a display screen and supporting communication with the server 104, including but not limited to smartphones, tablet computers, laptop portable computers, and desktop computers, etc.; when the terminal devices 101, 102, and 103 are software, they may be installed in the electronic devices as described above. The terminal devices 101, 102, and 103 may be implemented as multiple software or software modules, or may also be implemented as a single software or software module, and the embodiments of the present disclosure do not limit this. Further, various applications may be installed on the terminal devices 101, 102, and 103, such as data processing applications, instant messaging tools, social platform software, search applications, shopping applications, etc.

[0020] The server 104 may be a server providing various services. For example, it is a background server that receives requests sent by terminal devices establishing communication connections with it. The background server may receive and analyze requests sent by terminal devices and generate processing results. The server 104 may be a single server, or may be a server cluster composed of several servers, or may also be a cloud computing service center, and the embodiments of the present disclosure do not limit this.

[0021] It should be noted that the server 104 may be hardware or software. When the server 104 is hardware, it may be various electronic devices providing various services for the terminal devices 101, 102, and 103. When the server 104 is software, it may be multiple software or software modules providing various services for the terminal devices 101, 102, and 103, or may also be a single software or software module providing various services for the terminal devices 101, 102, and 103, and the embodiments of the present disclosure do not limit this.

[0022] The network 105 can be a wired network connected by coaxial cables, twisted pairs, and optical fibers, or a wireless network that can interconnect various communication devices without wiring. For example, Bluetooth, Near Field Communication (NFC), Infrared, etc. The embodiments of the present disclosure do not limit this.

[0023] Users can establish a communication connection with the server 104 via the network 105 through the terminal devices 101, 102, and 103 to receive or send information, etc. It should be noted that the specific types, quantities, and combinations of the terminal devices 101, 102, and 103, the server 104, and the network 105 can be adjusted according to the actual needs of the application scenario. The embodiments of the present disclosure do not limit this.

[0024] Figure 2 It is a schematic flowchart of a file creation method provided by the embodiments of the present disclosure. Figure 2 The file creation method can be executed by Figure 1 the terminal device or the server of Figure 2 As shown in

[0025] S201, respectively obtain the first feature space and the second feature space corresponding to the collected pictures. The first feature space is used to describe the space of the picture within the first category, and the second feature space is used to describe the space between the picture and the second category. The second category is a category other than the first category, and the first category is the category corresponding to the picture.

[0026] It should be noted that in the embodiments of the present disclosure, the above pictures include but are not limited to face pictures, animal pictures, building pictures, etc. Here, no limitations are made.

[0027] The above collected pictures can include one or multiple, and can be set according to specific file creation requirements. Here, no limitations are made.

[0028] The parameters used in the above first feature space and the second feature space include but are not limited to distance, and can also be other parameters that can describe space. Here, no limitations are made. For example, taking distance as an example, the above first feature space includes the first distance between the picture and the dense points in the space within the first category, and the second distance between the picture and the class center of the first category; the second feature space includes the third distance and the fourth distance between the picture and the second category. The third distance is the nearest neighbor distance, and the fourth distance is the distance to the nearest neighboring edge.

[0029] The above second category can include one category or multiple categories. Here, no limitations are made.

[0030] S202. Determine a target picture from the collected pictures according to the first feature space and the second feature space.

[0031] S203. Determine a picture for filing according to the target picture.

[0032] Through the above steps S201 - S203, when selecting a picture for filing, according to the picture space estimation, rather than simply using the classifier method in the prior art, the problem of poor filing quality caused by using the classifier method to select a picture for filing in the prior art is solved, and the technical effect of improving the filing quality is achieved.

[0033] Next, specific examples are combined to illustrate the embodiments of the present disclosure.

[0034] This example provides a face filing method based on face space estimation, as Figure 3 shown, including the following steps:

[0035] Step S301. Build a neural network N and complete the training on the data set V1.

[0036] Step S302. The neural network N extracts features on V2, performs clustering, calculates the center, dense points, and divergence of each class, and calculates the inter-class distance.

[0037] Step S303. Connect shared parameters and two main branches on the basis of the neural network N to predict the intra-class and inter-class distances and predict the divergence.

[0038] Step S304. Make the mean square error loss between the predicted value and the true value of the neural network N, and perform backpropagation to guide the network optimization.

[0039] Step S305. When filing, use a quality-based similarity formula to determine the class attribution of the picture.

[0040] Step S306. After determining the class attribution, calculate the filing index of the picture to determine whether the picture is used as a picture for filing.

[0041] It can be seen that this example directly estimates the intra-class distribution and inter-class distribution based on the face space, fundamentally defines the face quality, and based on the quality score estimated from the face space, it can be used to select a picture for filing or adjust the recognition similarity score to ensure the accuracy of the recognition task.

[0042] In some embodiments, the above first distance and the above second distance can be determined in the following manner:

[0043] Step S401. Complete the training of the built neural network on the first training set to obtain the trained neural network.

[0044] Step S402: Set the trained neural network to extract the features of the collected pictures on the second training set and perform clustering operations to obtain the in-class spatial dense points of each class of the collected pictures.

[0045] Optionally, the above clustering algorithm includes but is not limited to: the density-based clustering algorithm DBSCAN (Density-Based Spatial Clustering of Applications with Noise).

[0046] It should be noted that the above first training set and the second training set are two training sets of the same scale.

[0047] Step S403: Obtain the distances between the collected pictures and all the spatial dense points in the first class respectively to obtain a distance set.

[0048] Step S404: Set the minimum distance in the distance set as the first distance, and set the average value of all the distances in the distance set as the second distance.

[0049] Through the above steps S401 - S404, it is possible to predict the distance to the "nearest" spatial dense point of the picture and predict its distance to the "class center".

[0050] Next, specific examples are combined to illustrate the embodiments of the present disclosure.

[0051] Theoretically, if a face picture has good quality, it should be close to the in-class spatial dense points; otherwise, it is far from the in-class spatial dense points. Therefore, when inputting a face picture, we hope that the picture can predict its distance to the in-class spatial dense points.

[0052] Within a face class, there are multiple spatial dense points. When inputting a face picture, it is necessary to (1) predict the distance to its "nearest" spatial dense point. (2) It is necessary to predict its distance to the "class center". The so-called "class center" refers to the center of all sample points.

[0053] The technical solution of this example is as follows:

[0054] There are two training sets V1 and V2 of equal size. First, a deep neural network N for face recognition is trained on the training set V1, which can extract facial features. After training, we use the neural network N to extract features from the training set V2 and perform clustering operations using the DBSAN algorithm. After the clustering operation is completed, multiple dense clusters will appear in each class of the training set V2. We call the multiple centers corresponding to these multiple dense clusters intra-class dense points. There are not too many intra-class dense points, generally between 1 and 5. The reason why multiple dense points are formed within a class is often due to local aggregation caused by differences in age, posture, lighting, etc.

[0055] After clustering, we can calculate the distance between all samples of V2 and their "nearest" dense points. Since each sample has a corresponding category, we only need to calculate the distance between the sample and all the dense points in its class:

[0056] The minimum distance is the distance to the "nearest dense point".

[0057] Take the average distance, which is the distance from the "cluster center".

[0058] Optionally, the disclosed embodiment may also build a network to predict the “intra-class nearest dense point distance” and the “class center distance” respectively.

[0059] Optionally, using the neural network N trained with V1, several (generally 4) shared parameter matrices are subsequently added to the neural network N, and the output of the shared module is M. Then two main branches are designed. The first main branch calculates the intra-class distance, and the second main branch will be introduced later and is used to calculate the inter-class distance. On the first main branch, it contains 2 parameter matrices. The dimension of the first parameter matrix W01 is (512, 512), which represents the same-dimensional transformation, followed by relu activation; the dimension of the second parameter matrix W02 is (512, 128), which represents the transformation of dimensional compression, filtering invalid information, followed by tanh activation, and the output is set to a. Then two sub-branches are connected. The first sub-branch has two parameter matrices, namely W11 (128, 64), relu activation, W12 (64, 1), sigmoid operation, which is used to predict the distance of the nearest dense point; the second sub-branch has two parameter matrices, namely W21 (128, 128), swish activation, W22 (128, 1), sigmoid operation, which is used to predict the distance of the center point. The formula is as follows:

[0060] a=tanh(W 02 relu(W 01 M))

[0061] d n =sigmoid(W 12 relu(W 11 a))

[0062] d c =sigmoid(W 22 swish(W 21 a))

[0063] In addition, the loss function can also be calculated using squared error loss:

[0064] L d =(D n -d n ) 2 +(D c -d c ) 2

[0065] In this way, the network can accurately predict the distance to the nearest dense point within the class and the distance to the class center.

[0066] In an optional embodiment, the third distance and the fourth distance may be determined in the following manner:

[0067] Step S501, setting the average value of the distances between the image and the class centers of the nearest K1 classes as the third distance, where K1 is an integer greater than 0;

[0068] Step S502: Set the average value of the closest distances between the image and the nearest dense points of K1 classes as the fourth distance.

[0069] Specifically, faces are unevenly distributed in the feature space. Some faces have large inter-class distances with other faces, while some faces have small inter-class distances with other faces. We hope to predict the inter-class nearest neighbor distance. For faces of the same class, the larger the inter-class distance, the better the quality. So we propose a simple algorithm for "inter-class distance": the average of the closest distances to the dense points of the nearest k classes is called the "nearest neighbor edge distance"; the average of the distances to the class centers of the nearest k classes is called the "nearest neighbor distance".

[0070] Alternatively, the third distance and the fourth distance are also determined by the following method:

[0071] Step S601, calculating the average distance between the class center corresponding to the image and all dense points in the class to obtain the intra-class divergence;

[0072] Step S602, setting the distance between the image and the class centers of the nearest K2 classes multiplied by the intra-class divergences and calculating the average value as the third distance, where K2 is an integer greater than 0;

[0073] Step S603: Set the closest distance between the image and the nearest dense points of the K2 classes multiplied by the intra-class divergence and calculate the average value, which is used as the fourth distance.

[0074] Specifically, in order to estimate the "inter-class distance" more accurately, we propose a more complex algorithm. We must first calculate another statistic within the class: the intra-class divergence. The intra-class divergence is defined as the average distance between the class center and all the dense points in the class. The larger the intra-class divergence, the higher its spatial importance; the smaller the intra-class divergence, the lower its spatial importance. We propose a complex algorithm for the "inter-class distance": the distance to the nearest k class dense points is multiplied by their divergence to find the average, which is called the "nearest neighbor edge distance"; the distance to the class center of the nearest k classes is multiplied by their divergence to find the average, which is called the "nearest neighbor distance". In this algorithm, the divergence is equivalent to a weight to effectively distinguish the importance of neighboring nodes. We take k as 10, that is, for a certain sample x, we calculate the weighted (weight is the divergence) average of the distances to the 10 other class centers closest to it as the "nearest neighbor distance"; at the same time, we calculate the weighted (weight is the divergence) average of the distances to the 10 other class dense points closest to it as the "nearest neighbor edge distance". By traversing all samples, we can get the "nearest neighbor distance" E of all samples. c and the “nearest neighbor edge distance” E n .

[0075] Alternatively, you can build a network to predict, the network is the same as the one described above. There are several shared parameter matrices, the output is M, and two main branches. One main branch predicts the intra-class distance. The other main branch predicts the "inter-class distance". First, it contains 2 parameter matrices. The dimension of the first parameter matrix W31 is (512, 512), followed by relu activation; the dimension of the second parameter matrix W32 is (512, 256), followed by tanh activation, and the output is b. Then there are two sub-branches. The first sub-branch has two parameter matrices, namely W41 (256, 128), mish activation, W42 (128, 1), sigmoid operation, which is used to predict the nearest neighbor edge distance; the second sub-branch has two parameter matrices, namely W51 (256, 128), swish activation, W52 (128, 1), sigmoid operation, which is used to predict the nearest neighbor distance. The formula is as follows:

[0076] b = tanh(W 32 relu(W 31 M))

[0077] e n =sigmoid(W 42 mish(W 41 b))

[0078] e c =sigmoid(W 52 swish(W 51 b))

[0079] You can also calculate the loss function using squared error loss:

[0080] L e =(E n -e n ) 2 +(E c -e c ) 2

[0081] In this way, the network can accurately predict the distance to the nearest neighbor between classes and the distance to the nearest neighbor edge between classes.

[0082] It can be seen that the branch for predicting intra-class distance and the branch for predicting inter-class distance have roughly similar structures.

[0083] This embodiment proposes that, because the intra-class distance and the inter-class distance have a certain correlation, a calculation module combining two main branches is designed to calculate the divergence. The first main branch W02 outputs f1, the second main branch W41 outputs f2, and the second main branch W51 outputs f3. The three vectors are concatenated, connected to a (384, 128) matrix and swish activation, a (128, 1) matrix and a sigmoid function, and a value, i.e., the divergence value, is obtained.

[0084] The divergence value calculated by the network and the true divergence value are calculated using squared difference loss:

[0085] L s =(Ss) 2

[0086] The final loss is the sum of three losses:

[0087] L=L d +L e +L s

[0088] In an optional implementation, determining the target image from the acquired images according to the first feature space and the second feature space includes: calculating a target parameter y corresponding to the image by the following formula:

[0089]

[0090] When the target parameter is less than a first preset threshold, the picture is determined as a target picture.

[0091] Alternatively, determining the target image from the collected images according to the first feature space and the second feature space includes:

[0092] The target parameter y corresponding to the image is calculated using the following formula:

[0093]

[0094] When the target parameter is less than a second preset threshold, the picture is determined as a target picture.

[0095] By determining the target image through the above method, the accuracy of file creation can be further improved.

[0096] The present embodiment is described below with reference to specific examples.

[0097] Face archiving, that is, building a profile for a person, requires selecting a number of the most representative and highest quality face images from the many snapshots of this person for archiving. If the archiving images are well selected, the recall rate and precision rate of recognition can be greatly improved.

[0098] If there is no limit on the number of algorithm models and algorithm running time on the device where the algorithm is deployed, we will pre-install several quality models. For example, we will train an image blur recognition model, a facial occlusion recognition model, and a facial posture model. A face image will first pass through the above three models to obtain blur scores, occlusion scores, and posture scores, which represent whether the image is blurry, whether occlusion occurs, and whether the posture is too large. Obviously, poor quality images will be filtered out.

[0099] For the remaining images, we will enter the spatial estimation neural network proposed in this paper. Of course, if the algorithm deployment device has computing power and memory limitations, we do not need the above multi-mass model and directly use only the network model proposed in this paper. The single model in this paper can achieve better results than the multi-mass model and output richer spatial estimates.

[0100] For a picture, through the network of this paper, we can get d n , d c , e n , e c and intra-class divergence s,d n The smaller the e n The larger it is, the more suitable it is for archiving images. We propose the archiving index y as follows:

[0101]

[0102] For multiple snapshots of the same person, take the picture with the smallest y as the file creation picture.

[0103] Alternatively, the archiving indicator can be changed to:

[0104]

[0105] Adding exponents to these distances can be used as weight values; the shared parameter matrix can be 2-4;

[0106] Only predict d n and e n , or just predict d n and d c , that is, predicting a subset of this technical solution and deleting some branches in the network structure all fall within the scope of this technology.

[0107] In an optional implementation, determining the archiving picture according to the target picture includes:

[0108] Step S701, calculating the similarity between the target image and the existing image;

[0109] Step S702: Choose whether to replace the existing picture according to the similarity.

[0110] Through the above steps S701 to S702, a picture with better quality can be selected as a file creation picture.

[0111] The present embodiment is described below with reference to specific examples.

[0112] Optionally, if the person already has a profile picture p, and several new pictures are captured. If the similarity between a face picture p1 and the profile picture p in the new captured picture is less than 0.5, but d c <0.02, then the image can also be used as the archive image. Of course, if the y value of a face image in the new captured image is smaller than that of the archived image, it can be replaced.

[0113] The above can also adjust the score of face recognition.

[0114] Given two face images, the new face recognition similarity score is:

[0115]

[0116] It can be seen that the smaller the file building index y is, the higher the similarity score in the above formula is.

[0117] We use this new similarity to calculate whether the faces belong to the same class.

[0118] In addition, if the quality is too poor, such as the intra-class distance is too large and the inter-class distance is too small, we can discard the image to prevent image attacks.

[0119] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present application, which will not be described one by one here.

[0120] The following are embodiments of the device disclosed herein, which can be used to execute the method embodiments disclosed herein. For details not disclosed in the device embodiments disclosed herein, please refer to the method embodiments disclosed herein.

[0121] Figure 4 Schematic diagram of a file creation device provided by an embodiment of the present disclosure. Figure 4 As shown, the archiving device includes:

[0122] The acquisition module 401 is configured to respectively acquire a first feature space and a second feature space corresponding to the collected picture, wherein the first feature space is used to describe the space of the picture within the first category, and the second feature space is used to describe the space between the picture and the second category, the second category is a category other than the first category, and the first category is the category corresponding to the picture.

[0123] It should be noted that in the embodiments of the present disclosure, the above-mentioned pictures include but are not limited to pictures of human faces, pictures of animals, pictures of buildings, etc., and no limitation is made here.

[0124] The above-mentioned collected pictures may include one or more pictures, and may be set according to specific file creation requirements. No limitation is made here.

[0125] The parameters used in the first feature space and the second feature space include but are not limited to distance, and may also be other parameters that can describe space, and no limitation is made here. For example, taking distance as an example, the first feature space includes a first distance between the image and a spatially dense point in the first class, and a second distance between the image and the class center of the first class; the second feature space includes a third distance and a fourth distance between the image and the second class, the third distance being the nearest neighbor distance, and the fourth distance being the distance of the nearest neighbor edge.

[0126] The second category mentioned above may include one category or multiple categories, and no limitation is made here.

[0127] The first determination module 402 is configured to determine a target image from the collected images according to the first feature space and the second feature space.

[0128] The second determination module 403 is configured to determine the file creation picture according to the target picture.

[0129] Through the above module, when selecting archive pictures, it is estimated based on the picture space, rather than simply using the classifier method in the prior art, thereby solving the problem of poor archive quality caused by using the classifier method to select archive pictures in the prior art, and achieving the technical effect of improving the archive quality.

[0130] The embodiments of the present disclosure are described below with reference to specific examples.

[0131] This example provides a face profile creation method based on face space estimation. Figure 3 As shown, the following steps are included:

[0132] Step S301, building a neural network N, and training it on the data set V1;

[0133] Step S302, the neural network N extracts features from V2, clusters, calculates the center, dense points, and divergence of each class, and calculates the distance between classes;

[0134] Step S303, based on the neural network N, the shared parameters and two main branches are connected to predict the intra-class and inter-class distances, and the divergence;

[0135] Step S304, the squared difference loss is performed between the predicted value and the true value of the neural network N, and back propagation is used to guide network optimization;

[0136] Step S305, when creating a file, a quality-based similarity formula is used to determine the category of the image;

[0137] Step S306, after determining the category affiliation, we calculate the image's archiving index to determine whether the image should be used as an archiving image.

[0138] It can be seen that this example estimates the intra-class distribution and inter-class distribution directly based on the face space, which fundamentally defines the face quality. The quality score estimated from the face space can be used to select archive pictures, or to adjust the recognition similarity score to ensure the accuracy of the recognition task.

[0139] In some embodiments, the first distance and the second distance are determined in the following manner: the constructed neural network is trained on the first training set to obtain a trained neural network; the trained neural network is set to extract features of the collected images on the second training set and perform clustering operations to obtain spatial dense points within each category of the collected images; the distances between the collected images and all spatial dense points in the first category are obtained respectively to obtain a distance set: the minimum distance in the distance set is set as the first distance, and the average value of all distances in the distance set is set as the second distance.

[0140] Optionally, the above clustering algorithm includes but is not limited to: a density-based clustering algorithm DBSCAN (Density-Based Spatial Clustering of Applications with Noise).

[0141] It should be noted that the first training set and the second training set are two training sets of equal size.

[0142] By determining the method, the distance of the "nearest" spatially dense point of the image and its distance to the "class center" can be predicted.

[0143] The embodiments of the present disclosure are described below with reference to specific examples.

[0144] Theoretically, if a face image is of good quality, it should be close to the spatially dense points within the class; otherwise, it should be far away from the spatially dense points within the class. Therefore, when inputting a face image, we hope that the image can predict its distance from the spatially dense points within the class.

[0145] There are multiple spatially dense points in a face class. When inputting a face image, we need to (1) predict the distance to its “nearest” spatially dense point. (2) We need to predict its distance to the “class center”. The so-called “class center” refers to the center of all sample points.

[0146] The technical solution for this example is as follows:

[0147] There are two training sets V1 and V2 of equal size. First, a deep neural network N for face recognition is trained on the training set V1, which can extract facial features. After training, we use the neural network N to extract features from the training set V2 and perform clustering operations using the DBSAN algorithm. After the clustering operation is completed, multiple dense clusters will appear in each class of the training set V2. We call the multiple centers corresponding to these multiple dense clusters intra-class dense points. There are not too many intra-class dense points, generally between 1 and 5. The reason why multiple dense points are formed within a class is often due to local aggregation caused by differences in age, posture, lighting, etc.

[0148] After clustering, we can calculate the distance between all samples of V2 and their "nearest" dense points. Since each sample has a corresponding category, we only need to calculate the distance between the sample and all the dense points in its class:

[0149] The minimum distance is the distance to the "nearest dense point".

[0150] Take the average distance, which is the distance from the "cluster center".

[0151] Optionally, the disclosed embodiment may also build a network to predict the “intra-class nearest dense point distance” and the “class center distance” respectively.

[0152] Optionally, using the neural network N trained with V1, several (generally 4) shared parameter matrices are subsequently added to the neural network N, and the output of the shared module is M. Then two main branches are designed. The first main branch calculates the intra-class distance, and the second main branch will be introduced later and is used to calculate the inter-class distance. On the first main branch, it contains 2 parameter matrices. The dimension of the first parameter matrix W01 is (512, 512), which represents the same-dimensional transformation, followed by relu activation; the dimension of the second parameter matrix W02 is (512, 128), which represents the transformation of dimensional compression, filtering invalid information, followed by tanh activation, and the output is set to a. Then two sub-branches are connected. The first sub-branch has two parameter matrices, namely W11 (128, 64), relu activation, W12 (64, 1), sigmoid operation, which is used to predict the distance of the nearest dense point; the second sub-branch has two parameter matrices, namely W21 (128, 128), swish activation, W22 (128, 1), sigmoid operation, which is used to predict the distance of the center point. The formula is as follows:

[0153] a=tanh(W 02 relu(W 01 M))

[0154] d n =sigmoid(W 12 relu(W 11 a))

[0155] d c =sigmoid(W 22 swish(W 21 a))

[0156] In addition, the loss function can also be calculated using squared error loss:

[0157] L d =(D n -d n ) 2 +(D c -d c ) 2

[0158] In this way, the network can accurately predict the distance to the nearest dense point within the class and the distance to the class center.

[0159] In an optional embodiment, the third distance and the fourth distance are determined in the following manner: the average value of the distances between the image and the class centers of the nearest K1 classes is set as the third distance, where K1 is an integer greater than 0; the average value of the distances between the image and the dense points of the nearest K1 classes is set as the fourth distance.

[0160] Specifically, faces are unevenly distributed in the feature space. Some faces have large inter-class distances with other faces, while some faces have small inter-class distances with other faces. We hope to predict the inter-class nearest neighbor distance. For faces of the same class, the larger the inter-class distance, the better the quality. So we propose a simple algorithm for "inter-class distance": the average of the closest distances to the dense points of the nearest k classes is called the "nearest neighbor edge distance"; the average of the distances to the class centers of the nearest k classes is called the "nearest neighbor distance".

[0161] Alternatively, the third distance and the fourth distance are also determined by the following method:

[0162] Step S601, calculating the average distance between the class center corresponding to the image and all dense points in the class to obtain the intra-class divergence;

[0163] Step S602, setting the distance between the image and the class centers of the nearest K2 classes multiplied by the intra-class divergences and calculating the average value as the third distance, where K2 is an integer greater than 0;

[0164] Step S603: Set the closest distance between the image and the nearest dense points of the K2 classes multiplied by the intra-class divergence and calculate the average value, which is used as the fourth distance.

[0165] Specifically, in order to estimate the "inter-class distance" more accurately, we propose a more complex algorithm. We must first calculate another statistic within the class: the intra-class divergence. The intra-class divergence is defined as the average distance between the class center and all the dense points in the class. The larger the intra-class divergence, the higher its spatial importance; the smaller the intra-class divergence, the lower its spatial importance. We propose a complex algorithm for the "inter-class distance": the distance to the nearest k class dense points is multiplied by their divergence to find the average, which is called the "nearest neighbor edge distance"; the distance to the class center of the nearest k classes is multiplied by their divergence to find the average, which is called the "nearest neighbor distance". In this algorithm, the divergence is equivalent to a weight to effectively distinguish the importance of neighboring nodes. We take k as 10, that is, for a certain sample x, we calculate the weighted (weight is the divergence) average of the distances to the 10 other class centers closest to it as the "nearest neighbor distance"; at the same time, we calculate the weighted (weight is the divergence) average of the distances to the 10 other class dense points closest to it as the "nearest neighbor edge distance". By traversing all samples, we can get the "nearest neighbor distance" E of all samples. c and the “nearest neighbor edge distance” E n .

[0166] Alternatively, you can build a network to predict, the network is the same as the one described above. There are several shared parameter matrices, the output is M, and two main branches. One main branch predicts the intra-class distance. The other main branch predicts the "inter-class distance". First, it contains 2 parameter matrices. The dimension of the first parameter matrix W31 is (512, 512), followed by relu activation; the dimension of the second parameter matrix W32 is (512, 256), followed by tanh activation, and the output is b. Then there are two sub-branches. The first sub-branch has two parameter matrices, namely W41 (256, 128), mish activation, W42 (128, 1), sigmoid operation, which is used to predict the nearest neighbor edge distance; the second sub-branch has two parameter matrices, namely W51 (256, 128), swish activation, W52 (128, 1), sigmoid operation, which is used to predict the nearest neighbor distance. The formula is as follows:

[0167] b = tanh(W 32 relu(W 31 M))

[0168] e n =sigmoid(W 42 mish(W 41 b))

[0169] e c =sigmoid(W 52 swish(W 51 b))

[0170] You can also calculate the loss function using squared error loss:

[0171] L e =(E n -e n ) 2 +(E c -e c ) 2

[0172] In this way, the network can accurately predict the distance to the nearest neighbor between classes and the distance to the nearest neighbor edge between classes.

[0173] It can be seen that the branch for predicting intra-class distance and the branch for predicting inter-class distance have roughly similar structures.

[0174] This embodiment proposes that, because the intra-class distance and the inter-class distance have a certain correlation, a calculation module combining two main branches is designed to calculate the divergence. The first main branch W02 outputs f1, the second main branch W41 outputs f2, and the second main branch W51 outputs f3. The three vectors are concatenated, connected to a (384, 128) matrix and swish activation, a (128, 1) matrix and a sigmoid function, and a value, i.e., the divergence value, is obtained.

[0175] The divergence value calculated by the network and the true divergence value are calculated using squared difference loss:

[0176] L s =(Ss) 2

[0177] The final loss is the sum of three losses:

[0178] L=L d +L e +L s

[0179] In an optional embodiment, the first determination module 402 is further configured to calculate the target parameter y corresponding to the image by using the following formula:

[0180]

[0181] When the target parameter is less than a first preset threshold, the picture is determined as a target picture.

[0182] Alternatively, the first determining module 402 is further configured to calculate the target parameter y corresponding to the image by using the following formula:

[0183]

[0184] When the target parameter is less than a second preset threshold, the picture is determined as a target picture.

[0185] By determining the target image through the first determination module 402, the accuracy of file creation can be further improved.

[0186] The present embodiment is described below with reference to specific examples.

[0187] Face archiving, that is, building a profile for a person, requires selecting a number of the most representative and highest quality face images from the many snapshots of this person for archiving. If the archiving images are well selected, the recall rate and precision rate of recognition can be greatly improved.

[0188] If there is no limit on the number of algorithm models and algorithm running time on the device where the algorithm is deployed, we will pre-install several quality models. For example, we will train an image blur recognition model, a facial occlusion recognition model, and a facial posture model. A face image will first pass through the above three models to obtain blur scores, occlusion scores, and posture scores, which represent whether the image is blurry, whether occlusion occurs, and whether the posture is too large. Obviously, poor quality images will be filtered out.

[0189] For the remaining images, we will enter the spatial estimation neural network proposed in this paper. Of course, if the algorithm deployment device has computing power and memory limitations, we do not need the above multi-mass model and directly use only the network model proposed in this paper. The single model in this paper can achieve better results than the multi-mass model and output richer spatial estimates.

[0190] For a picture, through the network of this paper, we can get d n , d c , e n , e c and intra-class divergence s,d n The smaller the e n The larger it is, the more suitable it is for archiving images. We propose the archiving index y as follows:

[0191]

[0192] For multiple snapshots of the same person, take the picture with the smallest y as the file creation picture.

[0193] Alternatively, the archiving indicator can be changed to:

[0194]

[0195] Adding exponents to these distances can be used as weight values; the shared parameter matrix can be 2-4;

[0196] Only predict d n and e n , or just predict d n and d c , that is, predicting a subset of this technical solution and deleting some branches in the network structure all fall within the scope of this technology.

[0197] In an optional embodiment, the second determination module 403 is further configured to calculate the similarity between the target image and the existing image; and select whether to replace the existing image according to the similarity.

[0198] Through the above-mentioned second determination module 403, a picture with better quality can be selected as the file creation picture.

[0199] The present embodiment is described below with reference to specific examples.

[0200] Optionally, if the person already has a profile picture p, and several new pictures are captured. If the similarity between a face picture p1 and the profile picture p in the new captured picture is less than 0.5, but d c <0.02, then the image can also be used as the archive image. Of course, if the y value of a face image in the new captured image is smaller than that of the archived image, it can be replaced.

[0201] The above can also adjust the score of face recognition.

[0202] Given two face images, the new face recognition similarity score is:

[0203]

[0204] It can be seen that the smaller the file building index y is, the higher the similarity score in the above formula is.

[0205] We use this new similarity to calculate whether the faces belong to the same class.

[0206] In addition, if the quality is too poor, such as the intra-class distance is too large and the inter-class distance is too small, we can discard the image to prevent image attacks.

[0207] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present disclosure.

[0208] Figure 5 Schematic diagram of an electronic device 5 provided in an embodiment of the present disclosure. Figure 5 As shown, the electronic device 5 of this embodiment includes: a processor 501, a memory 502, and a computer program 503 stored in the memory 502 and executable on the processor 501. When the processor 501 executes the computer program 503, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 501 executes the computer program 503, the functions of the modules / units in the above-mentioned device embodiments are implemented.

[0209] Exemplarily, the computer program 503 may be divided into one or more modules / units, which are stored in the memory 502 and executed by the processor 501 to complete the present disclosure. The one or more modules / units may be a series of computer program instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program 503 in the electronic device 5.

[0210] The electronic device 5 may be a desktop computer, a notebook, a PDA, a cloud server, or other electronic device. The electronic device 5 may include, but is not limited to, a processor 501 and a memory 502. Those skilled in the art will appreciate that Figure 5 It is only an example of the electronic device 5 and does not constitute a limitation of the electronic device 5. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.

[0211] The processor 501 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc.

[0212] The memory 502 may be an internal storage unit of the electronic device 5, for example, a hard disk or memory of the electronic device 5. The memory 502 may also be an external storage device of the electronic device 5, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 5. Further, the memory 502 may also include both an internal storage unit of the electronic device 5 and an external storage device. The memory 502 is used to store computer programs and other programs and data required by the electronic device. The memory 502 may also be used to temporarily store data that has been output or is to be output.

[0213] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0214] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0215] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this disclosure.

[0216] In the embodiments provided in the present disclosure, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. Multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection of devices or units, which may be electrical, mechanical or other forms.

[0217] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0218] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0219] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present disclosure implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. The computer program may include computer program code, and the computer program code may be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electric carrier signals and telecommunication signals.

[0220] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them. Although the present disclosure has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be included in the protection scope of the present disclosure.

Claims

1. A method for creating a file, characterized in that, it includes: respectively obtaining a first feature space and a second feature space corresponding to the collected pictures, wherein the first feature space is used to describe the space of the pictures within the first category, the second feature space is used to describe the space between the pictures and the second category, the second category is a category other than the first category, and the first category is the category corresponding to the pictures; determining target pictures from the collected pictures according to the first feature space and the second feature space; determining file-creation pictures according to the target pictures; the first feature space includes a first distance between the picture and the dense points in the space within the first category, and a second distance between the picture and the class center of the first category; the second feature space includes a third distance and a fourth distance between the picture and the second category, the third distance is the nearest neighbor distance, and the fourth distance is the distance to the nearest adjacent edge; determining file-creation pictures according to the target pictures includes: calculating the similarity between the target pictures and the existing pictures; selecting whether to replace the existing pictures according to the similarity.

2. The method according to claim 1, characterized in that, the first distance and the second distance are determined by the following method: training the established neural network on a first training set to obtain a trained neural network; setting the trained neural network to extract the features of the collected pictures on a second training set and performing a clustering operation to obtain the dense points in the space within each category of the collected pictures; respectively obtaining the distances between the collected pictures and all the dense points in the space within the first category to obtain a distance set: setting the minimum distance in the distance set as the first distance, and setting the average value of all the distances in the distance set as the second distance.

3. The method according to claim 1, characterized in that, The first distance and the second distance are determined by the following formula: Among them, the parameter matrix has a dimension of (512, 512); Parameter matrix has a dimension of (512, 128); M is the output after adding a shared parameter matrix to the trained neural network; Parameter matrix has a dimension of (128, 64); Parameter matrix has a dimension of (64, 1); Parameter matrix has a dimension of (128, 128); Parameter matrix has a dimension of (128, 1).

4. The method according to claim 1, characterized in that, the third distance and the fourth distance are determined by the following method: setting the average value of the distances between the picture and the class centers of the nearest K1 classes as the third distance, where K1 is an integer greater than 0; setting the average value of the nearest distances between the picture and the dense points of the nearest K1 classes as the fourth distance.

5. The method according to claim 1, characterized in that, the third distance and the fourth distance are determined by the following method: calculating the average distance between the class center corresponding to the picture and all the dense points within the class to obtain the within-class divergence; setting the average value obtained by multiplying the distances between the picture and the class centers of the nearest K2 classes by the within-class divergence as the third distance, where K2 is an integer greater than 0; setting the average value obtained by multiplying the nearest distances between the picture and the dense points of the nearest K2 classes by the within-class divergence as the fourth distance.

6. The method according to claim 3, characterized in that, The third distance and the fourth distance are determined by the following method: Among them, the parameter matrix has a dimension of (512, 512); the dimension of the parameter matrix W32 is (512, 256); M is the output after adding a shared parameter matrix to the trained neural network; The dimension of the parameter matrix W41 is (256, 128); The dimension of the parameter matrix W42 is (128, 1); The dimension of the parameter matrix W51 is (256, 128); The dimension of the parameter matrix W52 is (128, 1).

7. The method according to claim 6, wherein, determining a target picture from the acquired pictures according to the first feature space and the second feature space includes: calculating a target parameter y corresponding to the picture by the following formula: when the target parameter is less than a first preset threshold, determining the picture as the target picture.

8. The method according to claim 6, wherein, determining a target picture from the acquired pictures according to the first feature space and the second feature space includes: calculating a target parameter y corresponding to the picture by the following formula: when the target parameter is less than a second preset threshold, determining the picture as the target picture.

9. An archiving device, wherein, comprising: an acquisition module, configured to respectively acquire a first feature space and a second feature space corresponding to the acquired pictures, wherein the first feature space is used to describe the space of the pictures within the first category, the second feature space is used to describe the space between the pictures and the second category, the second category is a category other than the first category, and the first category is the category corresponding to the pictures; the first feature space includes a first distance between the picture and the dense points in the space within the first category, and a second distance between the picture and the class center of the first category; the second feature space includes a third distance and a fourth distance between the picture and the second category, the third distance is the nearest neighbor distance, and the fourth distance is the distance to the nearest neighboring edge; a first determination module, configured to determine a target picture from the acquired pictures according to the first feature space and the second feature space; a second determination module, configured to determine an archived picture according to the target picture; the second determination module is specifically configured to: calculate the similarity between the target picture and the existing pictures; and select whether to replace the existing pictures according to the similarity.

10. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, when the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

11. A computer-readable storage medium storing a computer program, wherein, when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Picture clustering management method and system, equipment and medium

    CN111598012A