A method, apparatus, and electronic device for processing image files
By filtering, classifying and merging the similarity matrix in the face image archive, the problem of archive noise covering the archive subject is solved, and the integrity and accuracy of the face image archive is improved.
Patent Information
- Application Number
- CN202111419049.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-26
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-11-26
AI Technical Summary
When processing face image files, the prior art has the problem that archive noise covers the archive subject and causes the face image to be lost, destroying the integrity of the image file.
By filtering out the face image sets containing at least two target objects, using the feature extraction network to generate a similarity matrix, combining preset network model classification and merging similarity values, reclassifying and merging archival noise to the face image set with the same target object.
Ensure the integrity of the face image in the face image archive and the accuracy of filtering out the face images corresponding to the same target object, improving the overall integrity and accuracy of the image archive.
Smart Images

Figure CN114155579B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of face clustering, and particularly to a method, apparatus, and electronic device for processing image files. Background Art
[0002] With the development of image recognition technology, in order to classify a large number of face images, a face image clustering algorithm has been introduced. The face image clustering algorithm can classify the face images of the same person into one file through an online mode and an offline mode. The online mode is to put the unclassified face image file into the already formed face image file, or put the unclassified face images into the already formed face image set. The offline mode is to cluster the unclassified face images to form at least one new face image set. There are multiple face image sets in each face image file, and there is a centroid in each face image set. The centroid is obtained by calculating all the face features in the face image set.
[0003] When classifying face images based on the face image clustering algorithm, as the face image file continues to increase, face images of other people appear in the face image set, resulting in the generation of incorrect face image files. In the incorrect face image file, the face image with the same face and the largest number in the file is called the file main body, and the other face images are called file noises. There can be multiple file noises in the same face image file.
[0004] Currently, in order to process file noises, the method adopted is to count the number value of the face images corresponding to each face in the face image file. If the number value of the face images corresponding to the file noise in the face image file is greater than the number value of the face images corresponding to the file main body in the face image file, then the file noise with the largest number of face images in the face image file is used to cover the file main body, which will cause some face images to be lost and damage the integrity of the face images in the face image file. Summary of the Invention
[0005] This application provides a method, apparatus, and electronic device for processing image files. By reclassifying all the face images in the file noise, the face images of the same target object are classified into a second target image set, and the second target image set is merged into the face image set with the same target object, which improves the accuracy of screening out the face images corresponding to the same target object and further ensures the integrity of the face images in the face image file.
[0006] In a first aspect, this application provides a method for processing an image file, the method including:
[0007] Screen out a first target image set from the face image file to be processed, where the first target image set contains face images of at least two target objects;
[0008] Input the first target image set into the first preset network model to obtain the file subject and file noise corresponding to the first target image set;
[0009] Input all the face images in the file noise into the second preset network model for classification to obtain at least one second target image set, where the target objects in each second target image set are different;
[0010] Merge the second target image set with the face image set having the same target object in the face image file.
[0011] In a possible design, screening out the first target image set from the face image file to be processed includes:
[0012] Input each face image set in the face image file into the feature extraction network model to obtain the face image features of each face image in the face image set;
[0013] Calculate the similarity values between every two face image features in each face image set based on the face image features, generate a similarity matrix according to the similarity values, and obtain multiple similarity matrices;
[0014] Screen out at least one similarity matrix from the multiple similarity matrices according to the preset rule, and obtain the first target image set corresponding to the at least one similarity matrix.
[0015] In a possible design, screening out at least one similarity matrix from the multiple similarity matrices according to the preset rule, and obtaining the first target image set corresponding to the at least one similarity matrix includes:
[0016] Obtain the lowest similarity value in each similarity matrix, and detect whether the lowest similarity value is lower than the first preset threshold;
[0017] When it is determined that the lowest similarity value is lower than the first preset threshold, take the face image set as the first target image set.
[0018] In a possible design, obtaining the lowest similarity value in each similarity matrix includes:
[0019] Traverse the similarity values in each similarity matrix, and sort the similarity values based on the magnitude of the similarity values;
[0020] Screen out the minimum similarity value of the similarity matrix based on the sorting of the similarity values.
[0021] In a possible design, at least one similarity matrix is selected from multiple similarity matrices according to a preset rule, and a first target image set corresponding to the at least one similarity matrix is obtained. It further includes:
[0022] Obtain the average similarity value of each row in each similarity matrix, and detect whether the average similarity value is lower than a second preset threshold;
[0023] When it is determined that the average similarity value is lower than the second preset threshold, use the face image set as the first target image set.
[0024] In a possible design, obtaining the average similarity value of each row in each similarity matrix includes:
[0025] Obtain the number of similarity values of each row and / or each column in each similarity matrix;
[0026] Extract the similarity values of each row in the order from top to bottom, and calculate the sum of the similarity values of each row;
[0027] Divide the sum of the similarity values of each row by the number of similarity values to obtain the average similarity value of each row in each similarity matrix.
[0028] In a possible design, inputting the first target image set into a first preset network model to obtain the file subject and file noise corresponding to the first target image set includes:
[0029] Input the first target image set into the first preset network model to obtain a first probability value and a second probability value corresponding to each first target image in the first target image set;
[0030] Detect whether the first probability value is greater than the second probability value;
[0031] If so, use the first target image as the face image in the file subject;
[0032] If not, use the first target image as the face image in the file noise.
[0033] In a possible design, inputting all the face images in the file noise into a second preset network model for classification to obtain at least one second target image set includes:
[0034] Input all the face images in the file noise into the second preset network model to obtain the face image features of each face image in the file noise;
[0035] Based on the face image features of all the face images in the file noise, screen out the face images corresponding to each target object in the file noise;
[0036] Generate a second target image set based on the face images corresponding to the same target object, and obtain at least one second target image set.
[0037] In one possible design, based on the face image features of all face images in the file noise, filter out the face images corresponding to each target object in the file noise, including:
[0038] Calculate the similarity values between the face image features of all face images in the file noise pairwise;
[0039] Detect whether the similarity value exceeds a third preset threshold;
[0040] When it is determined that the similarity value exceeds the third preset threshold, regard the two face image features corresponding to the similarity value as the face image features corresponding to the same target object;
[0041] Obtain the face images corresponding to the target object according to the face image features corresponding to the same target object.
[0042] In one possible design, merge the second target image set with the face image set in the face image file that has the same target object, including:
[0043] Calculate the similarity values between each second target image set and the face image set in the face image file;
[0044] Sort the face image sets in the face image file based on the similarity values, and determine the second target image set and the face image set corresponding to the maximum similarity value;
[0045] Merge the second target image set with the face image set.
[0046] In a second aspect, the present application provides an image file processing device, and the device includes:
[0047] A screening module, configured to screen out a first target image set in a face image file to be processed;
[0048] An input module, configured to input the first target image set into a first preset network model to obtain the file main body and file noise corresponding to the first target image set;
[0049] A classification module, configured to classify all face images in the file noise by inputting them into a second preset network model to obtain at least one second target image set;
[0050] A merging module, configured to merge the second target image set with the face image set in the face image file that has the same target object.
[0051] In a possible design, the screening module is specifically configured to input each face image set in the face image file into a feature extraction network model to obtain face image features of each face image in the face image set, calculate similarity values between every two of the face image features in each face image set based on the face image features, generate a similarity matrix according to the similarity values, obtain multiple similarity matrices, screen out at least one similarity matrix from the multiple similarity matrices according to a preset rule, and obtain a first target image set corresponding to the at least one similarity matrix.
[0052] In a possible design, the screening module is further configured to obtain the lowest similarity value in each similarity matrix, detect whether the lowest similarity value is lower than a first preset threshold, and when it is determined that the lowest similarity value is lower than the first preset threshold, use the face image set as the first target image set.
[0053] In a possible design, the screening module is further configured to traverse the similarity values in each similarity matrix, sort the similarity values based on the magnitudes of the similarity values, and screen out the minimum similarity value of the similarity matrix based on the sorting of the similarity values.
[0054] In a possible design, the screening module is further configured to obtain the average similarity value of each row in each similarity matrix, detect whether the average similarity value is lower than a second preset threshold, and when it is determined that the average similarity value is lower than the second preset threshold, use the face image set as the first target image set.
[0055] In a possible design, the screening module is further configured to obtain the number of similarity values in each row and / or each column of each similarity matrix, extract the similarity values of each row in the order from top to bottom, calculate the sum of the similarity values of each row, and divide the sum of the similarity values of each row by the number of similarity values to obtain the average similarity value of each row in each similarity matrix.
[0056] In a possible design, the input module is specifically configured to input the first target image set into a first preset network model to obtain a first probability value and a second probability value corresponding to each first target image in the first target image set, detect whether the first probability value is greater than the second probability value, and if so, use the first target image as the face image in the file main body, and if not, use the first target image as the face image in the file noise.
[0057] In a possible design, the classification module is specifically configured to input all face images in the file noise into a second preset network model to obtain face image features of each face image in the file noise. Based on the face image features of all face images in the file noise, the face images corresponding to each target object in the file noise are screened out, and a second target image set is generated according to the face images corresponding to the same target object, and at least one second target image set is obtained.
[0058] In a possible design, the classification module is further configured to calculate similarity values between pairwise face image features of all face images in the file noise, detect whether the similarity values exceed a third preset threshold, and when it is determined that the similarity values exceed the third preset threshold, use the two face image features corresponding to the similarity values as face image features corresponding to the same target object, and obtain the face images corresponding to the target object according to the face image features corresponding to the same target object.
[0059] In a possible design, the merging module is specifically configured to calculate similarity values between each second target image set and the face image set in the face image file, sort the face image set in the face image file based on the similarity values, determine the second target image set and the face image set corresponding to the maximum similarity value, and merge the second target image set with the face image set.
[0060] In a third aspect, the present application provides an electronic device, including:
[0061] A memory for storing a computer program;
[0062] A processor, when executing the computer program stored on the memory, implements the steps of the above-mentioned method for processing an image file.
[0063] In a fourth aspect, a computer-readable storage medium stores a computer program therein, and when the computer program is executed by a processor, the steps of the above-mentioned method for processing an image file are implemented.
[0064] For the various aspects in the above first aspect to fourth aspect and the possible technical effects that each aspect may achieve, please refer to the technical effects that can be achieved by the above-mentioned first aspect or various possible solutions in the first aspect, and details will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 It is a flowchart of the steps of a method for processing an image file provided by the present application;
[0066] Figure 2Schematic structural diagram of a processing device for image files provided by this application;
[0067] Figure 3 Schematic structural diagram of an electronic device provided by this application. Detailed implementation manners
[0068] In order to make the objectives, technical solutions, and advantages of this application clearer, the following will further describe this application in detail with reference to the accompanying drawings. The specific operation methods in the method embodiments can also be applied to the device embodiments or system embodiments. It should be noted that in the description of this application, "a plurality of" is understood as "at least two". "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The connection between A and B can represent: the direct connection between A and B and the connection between A and B through C. In addition, in the description of this application, words such as "first" and "second" are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order.
[0069] In the prior art, to process the file noise in the face image file, the number of face images corresponding to the same target object in the face image file is calculated, and the face image set with the largest number of face images corresponding to the same target object in the file noise is selected. When the number of face images in the face image set in the file noise exceeds the number of face images in the face image set in the file main body, the file noise with the largest number of face images in the face image file is used to overwrite the file main body, which will cause some face images to be lost and damage the integrity of the face images in the face image file.
[0070] To solve the above problems, the embodiments of this application provide a method for processing image files to reclassify the file noise in the image files and merge it into other face image sets with the same target object, ensuring the integrity of the face images in the face image file and the accuracy of screening out the face images corresponding to the same target object. Among them, the methods and devices in the embodiments of this application are based on the same technical concept. Since the principles of the problems solved by the methods and devices are similar, the embodiments of the device and the method can be referred to each other, and the repeated parts will not be described again.
[0071] The following will describe the embodiments of this application in detail with reference to the accompanying drawings.
[0072] Refer to Figure 1 , this application provides a method for processing image files, which can ensure the integrity of the face images in the face image file. The implementation process of this method is as follows:
[0073] Step S1: Screen out the first target image set from the face image files to be processed.
[0074] In the embodiments of the present application, the face image files contain multiple face image sets. When a face image set contains face images of at least two target objects, it means there is file noise in this face image set. At this time, it is necessary to screen out the first target image set from the multiple face image sets. The first target image set contains face images corresponding to at least two target objects. The specific screening method is as follows:
[0075] First, input each face image set in the face image files into the feature extraction network, and use the feature extraction network to obtain the face image features of each face image in each face image set. After obtaining the face image features of each face image in each face image set, in order to better screen out the face image sets containing file noise, it is necessary to calculate the similarity values between every two face image features in each face image set. After obtaining the similarity values between every two face image features in each face image set, a similarity matrix is formed based on the similarity values. The similarity matrix formed based on the similarity values is shown in Table 1:
[0076]
[0077]
[0078] Table 1
[0079] In Table 1 above, the recorded similarity matrix is composed of the similarity values between every two of 5 face image features. It should be noted that there is only one face image feature for one target object, and the number of face image features is not only 5 in practical applications. For other face image features that generate a similarity matrix according to the similarity values between face image features, reference can be made to Table 1, and no further elaboration will be made here.
[0080] Furthermore, it should be noted that since the number of face image sets in the face image files is at least one, therefore, the number of similarity matrix values generated according to the face image features is at least one, and the number of similarity matrices is equal to the number of face image sets in the face image files.
[0081] After obtaining the similarity matrix, in order to screen out the first target image set from the face image sets, the following method can be used for screening:
[0082] Method 1: To screen out the first target image set from the face image set, the method adopted is to extract the lowest similarity value in the similarity matrix, and detect whether the lowest similarity value is lower than the first preset threshold. If the lowest similarity value is lower than the first preset threshold, it means that there are other face images in the face image set corresponding to the current similarity matrix, and the face image set corresponding to the similarity matrix is used as the first target image set. If the lowest similarity value is higher than the first preset threshold, it means that there are no face images of other target objects in the face image set corresponding to the current similarity matrix, and the face image set corresponding to the current similarity matrix is ignored.
[0083] Method 2: To screen out the first target image set from the face image set, the method adopted is to calculate the average similarity value of each row in each similarity matrix. After obtaining the average similarity value, detect whether each average similarity value is lower than the second preset threshold. If in the similarity matrix, the average similarity value of one row exceeds the second preset threshold, it means that there are face images of other target objects in the face image set corresponding to the current similarity matrix, and the face image set corresponding to the similarity matrix is used as the first target image set. If in the similarity matrix, the average similarity value of each row is lower than the second preset threshold, it means that there are no face images of other target objects in the face image set corresponding to the similarity matrix, and the face image set corresponding to the current similarity matrix is ignored.
[0084] For example: The similarity matrix is shown in Table 1. The specific process of calculating the average similarity value of each row in the similarity matrix in Table 1 is as follows:
[0085] In Table 1, the number of rows and columns of the similarity matrix is equal, both are 5. Calculate the sum of the similarity values of each row. The sum of the similarity values of the first row is 4.31, and the average similarity value corresponding to the similarity values of the first row is 0.862; the sum of the similarity values of the second row is 4.35, and the average similarity value corresponding to the similarity values of the second row is 0.87; the sum of the similarity values of the third row is 4.34, and the average similarity value corresponding to the similarity values of the third row is 0.868; the sum of the similarity values of the fourth row is 4.43, and the average similarity value corresponding to the similarity values of the fourth row is 0.886; the sum of the similarity values of the fifth row is 4.53, and the average similarity value corresponding to the similarity values of the fifth row is 0.906.
[0086] If the second preset threshold is 0.899, and the average similarity value of the fifth row is 0.906, 0.906 > 0.899. Therefore, the face image set corresponding to the similarity matrix contains at least two face images of target objects.
[0087] Further, it should be noted that the method for screening out the first target image set from the face image set can be any one of Method 1 and Method 2, or a combination of Method 1 and Method 2, which will not be elaborated here.
[0088] Through the above method, screening out the first target image set from the face image archive, using Method 1 and / or Method 2 can ensure that the face image set in the face image archive containing at least two target objects is screened out, and this face image set is used as the first target image set, so that the first target image set can be accurately screened out, improving the accuracy of screening out the first target image set.
[0089] Step S2: Input the first target image set into the first preset network model to obtain the archive subject and archive noise corresponding to the first target image set.
[0090] After obtaining the first target image set, the face images in the first target image set contain at least two target objects. In order to divide the face images in the first target image set into archive subjects and archive noise, it is necessary to input the first target image set into the first preset network model. In the embodiment of the present application, the first preset network model is a graph convolutional classification model. The number of graph convolutional layers in the graph convolutional classification model is adjusted according to the actual situation. However, the number of output channels of the last graph convolutional layer is 2. The graph convolutional classification model will calculate the first probability value and the second probability value corresponding to each first target image in the first target image set. The first probability value represents the probability value that the first target image is a face image in the archive subject, and the second probability value represents the probability value that the first target image is a face image in the archive noise.
[0091] After inputting the first target image set into the first preset network model, obtaining the first probability value and the second probability value of each first target image, after determining the first probability value and the second probability value of each first target image, detect whether the first probability value exceeds the second probability value. If the first probability value exceeds the second probability value, it means that the probability that the first target image is a face image in the archive subject is relatively high, then this first target image is used as the face image in the archive main image. If the first probability value is lower than the second probability value, it means that the probability that the first target image is a face image in the archive noise is relatively high, then this first target image is used as the face image in the archive noise.
[0092] Through the above method, the face images in the first target image set are divided into archive subjects and archive noise, ensuring that the face images of the archive subjects are the face images corresponding to the same target object, and putting the other face images in the first target image set into the archive noise, which is beneficial to the correct classification of face images.
[0093] Step S3: Input all the face images in the file noise into a second preset network model for classification to obtain at least one second target image set.
[0094] After obtaining the file noise in the first target image set, input the file noise into the second preset network model. In the embodiments of the present application, the second preset network model can be a density clustering model, a spectral clustering model, or a hierarchical clustering model. Since the density clustering model, the spectral clustering model, and the hierarchical clustering model are well-known technologies in the art, they will not be elaborated here.
[0095] After inputting the file noise into the second preset network model, in order to distinguish the face images corresponding to each target object, the second preset network model needs to collect the face image features of each face image. After obtaining the face image features of each face image in the file noise, calculate the similarity values between the face image features pairwise, and detect whether the similarity value exceeds a third preset threshold. If the similarity value exceeds the third preset threshold, then regard the two face image features corresponding to the similarity value as the face image features corresponding to the same target object, thereby determining the face images corresponding to the same target object. If the similarity value is lower than the third preset threshold, then the two face image features corresponding to the similarity value are not the face image features corresponding to the same target object, thereby determining that each of the face image features corresponds to a target object.
[0096] Based on the method described above, screen out the face images corresponding to each target object, and generate a second target image set according to the face images corresponding to each target object. Further, it should be noted that since the target objects corresponding to the face images in the file noise are at least one, the second target image set corresponding to the file noise is at least one.
[0097] Through the above method, input the file noise into the second preset network model, classify the face images in the file noise through the second preset network model, and classify the face images corresponding to the same target object into a second target image set, which is beneficial to merging the second target image set into other face images.
[0098] Step S4: Merge the second target image set with the face image set in the face image file that has the same target object.
[0099] After obtaining at least one second target image set, in order to ensure the integrity of the face images in the face image archive, the filtered archive noise cannot be deleted, so it is necessary to retain the second target image set generated based on the archive noise. The face images in the second target image set are the face images corresponding to the same target object. It is necessary to calculate the similarity value between the second target image set and the face image set in the face image archive, and then sort the obtained similarity values in descending order, and screen out the face image set corresponding to the maximum similarity value. Finally, merge the second target image with the face image set corresponding to the maximum similarity value.
[0100] For example: the face image set in the face image archive is A, B, C, D, and the second target image set is X. Calculate the similarity values between X and A, B, C, D respectively. The similarity values between X and A, B, C, D are 0.66, 0.87, 0.97, 0.69 in turn. Sort the similarity values in descending order, 0.97>0.87>0.69>0.66. The face image set corresponding to the maximum similarity value is C, and merge C and X.
[0101] It should be further noted that since there is at least one second target image set, therefore, the process of merging other second target image sets with other face image sets refers to the above example. Since the process of merging the second target image set with other face image sets is the same, it will not be elaborated here too much.
[0102] After merging the second target image set into the face image set in the face image archive, there is only one target object corresponding to the face image in each face image set, and the centroid of the current face image set is calculated.
[0103] This application embodiment only exemplifies a processing method for archive noise in one face image archive. When there are multiple face image archives, the processing methods for the archive noise corresponding to multiple face image archives are the same as those in this application embodiment. The specific process of processing archive noise can refer to the embodiments of this application. In order to avoid a large amount of repetitive content, it will not be elaborated here too much.
[0104] Through the method described above, screen out the face image set containing at least two target objects in the face image archive as the first target image set, then screen out the archive noise in the first target image set and classify it according to the face images corresponding to the same target object to obtain at least one second target image set. Finally, merge the second target image set into the face image set in the face image archive according to the similarity value, reclassify and merge the face images in the face image archive, ensuring the integrity of the face images in the face image archive and the accuracy of screening out the face images corresponding to the same target object.
[0105] Based on the same inventive concept, an image file processing device is further provided in an embodiment of the present application. The image file processing device is used to implement the functions of an image file processing method. Refer to Figure 2 , the device includes:
[0106] A screening module 201, configured to screen out a first target image set from the face image files to be processed;
[0107] An input module 202, configured to input the first target image set into a first preset network model to obtain a file subject and file noise corresponding to the first target image set;
[0108] A classification module 203, configured to input all face images in the file noise into a second preset network model for classification to obtain at least one second target image set;
[0109] A merging module 204, configured to merge the second target image set with the face image set having the same target object in the face image file.
[0110] In a possible design, the screening module 201 is specifically configured to input each face image set in the face image file into a feature extraction network model, obtain face image features of each face image in the face image set, calculate similarity values between any two of the face image features in each face image set based on the face image features, generate a similarity matrix according to the similarity values, obtain multiple similarity matrices, screen out at least one similarity matrix from the multiple similarity matrices according to a preset rule, and obtain the first target image set corresponding to the at least one similarity matrix.
[0111] In a possible design, the screening module 201 is further configured to obtain the lowest similarity value in each similarity matrix, detect whether the lowest similarity value is lower than a first preset threshold, and when it is determined that the lowest similarity value is lower than the first preset threshold, use the face image set as the first target image set.
[0112] In a possible design, the screening module 201 is further configured to traverse the similarity values in each similarity matrix, sort the similarity values based on the magnitudes of the similarity values, and screen out the minimum similarity value of the similarity matrix based on the sorting of the similarity values.
[0113] In a possible design, the screening module 201 is further configured to obtain the average similarity value of each row in each similarity matrix, detect whether the average similarity value is lower than a second preset threshold, and when it is determined that the average similarity value is lower than the second preset threshold, use the face image set as the first target image set.
[0114] In a possible design, the screening module 201 is further configured to obtain the number of similarity values in each row and / or each column of each similarity matrix, extract the similarity values of each row in the top-down order, calculate the sum of the similarity values of each row, and divide the sum of the similarity values of each row by the number of similarity values to obtain the average similarity value of each row in each similarity matrix.
[0115] In a possible design, the input module 202 is specifically configured to input the first target image set into a first preset network model, obtain a first probability value and a second probability value corresponding to each first target image in the first target image set, detect whether the first probability value is greater than the second probability value, and if so, use the first target image as the face image in the file body, and if not, use the first target image as the face image in the file noise.
[0116] In a possible design, the classification module 203 is specifically configured to input all the images in the file noise into a second preset network model, obtain the face image features of each face image in the file noise, screen out the face images corresponding to each target object in the file noise based on the face image features of all the face images in the file noise, generate a second target image set according to the face images corresponding to the same target object, and obtain at least one second target image set.
[0117] In a possible design, the classification module 203 is further configured to calculate the similarity values between the face image features of all the face images in the file noise pairwise, detect whether the similarity values exceed a third preset threshold, and when it is determined that the similarity values exceed the third preset threshold, use the two face image features corresponding to the similarity values as the face image features corresponding to the same target object, and obtain the face images corresponding to the target object according to the face image features corresponding to the same target object.
[0118] In a possible design, the merging module 204 is specifically configured to calculate the similarity values between each second target image set and the face image set in the face image file, sort the face image set in the face image file based on the similarity values, determine the second target image set and the face image set corresponding to the maximum similarity value, and merge the second target image set and the face image set.
[0119] Based on the same inventive concept, an electronic device is further provided in an embodiment of the present application. The electronic device can implement the functions of the foregoing image file processing device. Refer to Figure 3 , the electronic device includes:
[0120] At least one processor 301, and a memory 302 connected to the at least one processor 301. In the embodiments of the present application, the specific connection medium between the processor 301 and the memory 302 is not limited. Figure 3 In the example, the processor 301 and the memory 302 are connected through a bus 300. The bus 300 is Figure 3 represented by a thick line in the figure. The connection manners between other components are only for illustrative purposes and are not limiting. The bus 300 can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 3 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus. Alternatively, the processor 301 can also be called a controller, and there is no limit to the name.
[0121] In the embodiments of the present application, the memory 302 stores instructions executable by the at least one processor 301. By executing the instructions stored in the memory 302, the at least one processor 301 can execute a method for processing an image file described above. The processor 301 can implement Figure 2 the functions of each module in the device shown.
[0122] Among them, the processor 301 is the control center of the device, and can connect various parts of the entire control device through various interfaces and lines. By running or executing the instructions stored in the memory 302 and calling the data stored in the memory 302, various functions of the device and process data, so as to monitor the device as a whole.
[0123] In a possible design, the processor 301 may include one or more processing units. The processor 301 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above modem processor may not be integrated into the processor 301. In some embodiments, the processor 301 and the memory 302 can be implemented on the same chip, and in some embodiments, they can also be separately implemented on independent chips.
[0124] The processor 301 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of a method for processing an image file disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0125] The memory 302, being a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 302 can include at least one type of storage medium. For example, it can include flash memory, hard disks, multimedia cards, card-type memories, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memories, magnetic disks, optical disks, and so on. The memory 302 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 302 in the embodiments of the present application can also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.
[0126] By programming the design of the processor 301, the code corresponding to the method for processing an image file introduced in the foregoing embodiments can be solidified into the chip, so that the chip can execute Figure 1 the processing steps of an image file in the embodiments shown. How to program the design of the processor 301 is a well-known technology to those skilled in the art and will not be elaborated here.
[0127] Based on the same inventive concept, the embodiments of the present application also provide a storage medium storing computer instructions, which, when run on a computer, cause the computer to execute a method for processing an image file discussed above.
[0128] In some possible implementation manners, various aspects of a method for processing an image file provided in the present application can also be implemented in the form of a program product, which includes program code. When the program product runs on a device, the program code is used to cause the control device to execute the steps in a method for processing an image file according to various exemplary embodiments of the present application described above in this specification.
[0129] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0130] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0131] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0132] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0133] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. A method for processing an image file, characterized in that, Including: Screening out a first target image set from the face image files to be processed, where the first target image set contains face images of at least two target objects; Inputting the first target image set into a first preset network model to obtain the file main body and file noise corresponding to the first target image set; where the file main body is the face image with the same face and the largest number in the first target image set, and the file noise is the face images in the first target image set except the file main body; Inputting all the face images in the file noise into a second preset network model to obtain the face image features of each face image in the file noise; based on the face image features of all the face images in the file noise, screening out the face images corresponding to each target object in the file noise; generating a second target image set according to the face images corresponding to the same target object, and obtaining at least one second target image set, where the target objects in each second target image set are different; Merging the second target image set with the face image set in the face image file that has the same target object.
2. The method according to claim 1, wherein Screening out a first target image set from the face image files to be processed includes: Inputting each face image set in the face image file into a feature extraction network model to obtain the face image features of each face image in the face image set; Calculating the similarity values between every two face image features in each face image set based on the face image features, generating a similarity matrix according to the similarity values, and obtaining multiple similarity matrices; Screening out at least one similarity matrix from the multiple similarity matrices according to a preset rule, and obtaining the first target image set corresponding to the at least one similarity matrix.
3. The method according to claim 2, wherein Screening out at least one similarity matrix from the multiple similarity matrices according to a preset rule, and obtaining the first target image set corresponding to the at least one similarity matrix includes: Obtaining the lowest similarity value in each similarity matrix, and detecting whether the lowest similarity value is lower than a first preset threshold; When it is determined that the lowest similarity value is lower than the first preset threshold, taking the face image set as the first target image set.
4. The method according to claim 3, characterized in that, Obtaining the lowest similarity value in each similarity matrix includes: Traversing the similarity values in each similarity matrix, and sorting the similarity values based on the size of the similarity values; Screening out the minimum similarity value of the similarity matrix based on the sorting of the similarity values.
5. The method according to claim 2, wherein Screening out at least one similarity matrix from the multiple similarity matrices according to a preset rule, and obtaining the first target image set corresponding to the at least one similarity matrix further includes: Obtaining the average similarity value of each row in each similarity matrix, and detecting whether the average similarity value is lower than a second preset threshold; When it is determined that the average similarity value is lower than the second preset threshold, taking the face image set as the first target image set.
6. The method according to claim 5, wherein Obtaining the average similarity value of each row in each similarity matrix includes: Obtaining the number of similarity values in each row and / or each column of each similarity matrix; Extract the similarity values of each line in the order from top to bottom, and calculate the sum of the similarity values of each line; Divide the sum of the similarity values of each line by the number of the similarity values to obtain the average similarity value of each line in each similarity matrix.
7. The method according to claim 1, wherein Input the first target image set into the first preset network model to obtain the file subject and file noise corresponding to the first target image set, including: Input the first target image set into the first preset network model to obtain the first probability value and the second probability value corresponding to each first target image in the first target image set; Detect whether the first probability value is greater than the second probability value; If so, use the first target image as the face image in the file subject; If not, use the first target image as the face image in the file noise.
8. The method according to claim 1, wherein Based on the face image features of all the face images in the file noise, screen out the face images corresponding to each target object in the file noise, including: Calculate the similarity values between the face image features of all the face images in the file noise pairwise; Detect whether the similarity value exceeds a third preset threshold; When it is determined that the similarity value exceeds the third preset threshold, use the two face image features corresponding to the similarity value as the face image features corresponding to the same target object; Obtain the face image corresponding to the target object according to the face image features corresponding to the same target object.
9. The method according to claim 1, characterized in that, Merge the second target image set with the face image set in the face image file that has the same target object, including: Calculate the similarity values between each second target image set and the face image set in the face image file; Sort the face image sets in the face image file based on the similarity values, and determine the second target image set and the face image set corresponding to the maximum similarity value; Merge the second target image set with the face image set.
10. An image file processing device, characterized in that, The device includes: A screening module, configured to screen out a first target image set from the face image file to be processed; An input module, configured to input the first target image set into the first preset network model to obtain the file subject and file noise corresponding to the first target image set; wherein, the file subject is the face image with the same face and the largest number in the first target image set, and the file noise is the face image in the first target image set except the file subject; A classification module, configured to input all the face images in the file noise into a second preset network model to obtain the face image features of each face image in the file noise; based on the face image features of all the face images in the file noise, screen out the face images corresponding to each target object in the file noise; generate a second target image set according to the face images corresponding to the same target object, and obtain at least one second target image set; A merging module, configured to merge the second target image set with the face image set in the face image file that has the same target object.
11. The device according to claim 10, characterized in that, The screening module is specifically configured to input each face image set in the face image file into the feature extraction network model, obtain the face image features of each face image in the face image set, calculate the similarity values between every two of the face image features in each face image set based on the face image features, generate a similarity matrix according to the similarity values, obtain multiple similarity matrices, screen out at least one similarity matrix from the multiple similarity matrices according to a preset rule, and obtain a first target image set corresponding to the at least one similarity matrix.
12. The device according to claim 11, characterized in that, The screening module is further configured to obtain the lowest similarity value in each similarity matrix, detect whether the lowest similarity value is lower than a first preset threshold, and when it is determined that the lowest similarity value is lower than the first preset threshold, use the face image set as the first target image set.
13. The device according to claim 11, characterized in that, The screening module is further configured to obtain the average similarity value of each row in each similarity matrix, detect whether the average similarity value is lower than a second preset threshold, and when it is determined that the average similarity value is lower than the second preset threshold, use the face image set as the first target image set.
14. The device according to claim 10, characterized in that, The input module is specifically configured to input the first target image set into a first preset network model, obtain a first probability value and a second probability value corresponding to each first target image in the first target image set, detect whether the first probability value is greater than the second probability value, and if so, use the first target image as the face image in the file main body, and if not, use the first target image as the face image in the file noise.
15. The device according to claim 10, characterized in that, The merging module is specifically configured to calculate the similarity values between each second target image set and the face image set in the face image file, sort the face image set in the face image file based on the similarity values, determine the second target image set and the face image set corresponding to the maximum similarity value, and merge the second target image set and the face image set.
16. An electronic device, characterized in that, Including: A memory for storing a computer program; A processor for implementing the method steps according to any one of claims 1-9 when executing the computer program stored on the memory.
17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the method steps according to any one of claims 1-9 are implemented.
Citation Information
Patent Citations
Archiving method and device
CN109815370A
Clustering method, clustering device and computer readable storage medium
CN113255841A