Lipstick recognition method, device, medium and equipment for ID photo based on convolutional neural network

By constructing a lipstick recognition system based on convolutional neural networks, and using a convolutional neural network with multi-dimensional information fusion to identify lipsticks in ID photos, the existing technology has solved the problem of difficulty in recognition when light changes, environment complexity or partial lip lesions, and achieved higher recognition accuracy and robustness.

CN114155575BActive Publication Date: 2025-05-02GUANGZHOU PRESTIGE TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111328067.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-10
Publication Date
2025-05-02
Estimated Expiration
2041-11-10

AI Technical Summary

Technical Problem

The prior art cannot effectively identify the lipstick in the ID photo when light changes, environment is complex or partial lesions of the lips, and manual screening is time-consuming and labor-intensive and has a high misidentification rate.

Method used

Using a method based on a convolutional neural network, by obtaining the image training set, the RGB spatial image of the lips area is preprocessed and converted into a YCbCr spatial image, and a convolutional neural network is constructed for training, and a lipstick recognition model is obtained, which is used to recognize the image to be recognized.

Benefits of technology

It improves the robustness and accuracy of lipstick recognition, and can effectively identify lipstick in ID photos in the case of changes in light, complex environment or partial lesions of the lips, reducing the rate of misidentification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114155575B_ABST
    Figure CN114155575B_ABST
Patent Text Reader

Abstract

The present invention discloses a lipstick recognition method for ID photos based on a convolutional neural network, comprising: obtaining an image training set, preprocessing the image training set to obtain an RGB space image of the lip area; converting the RGB space image of the lip area into a YCbCr space image, and forming a training sample pair with the RGB space image and the YCbCr space image; constructing a convolutional neural network, using the training sample pair to train the convolutional neural network to obtain a lipstick recognition model; obtaining an image to be recognized, preprocessing it, and performing lipstick recognition on the preprocessed image to be recognized by the trained lipstick recognition model to obtain the category probability corresponding to the image to be recognized, and selecting the category label with the largest probability as the lipstick recognition result of the image to be recognized. The present invention solves the problem that the prior art cannot effectively and stably recognize lipstick in ID photos when the light changes, the environment is complex, and the lips have partial lesions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a method, device, medium and equipment for identifying lipstick in ID photos based on a convolutional neural network. Background Art

[0002] When using ID photo equipment to take ID photos, users must dress up strictly according to the ID photo standards. Various lipsticks seriously affect the quality of ID photo shooting and fail to meet the ID photo standards. The existing technology is to manually screen lipsticks in the background, which is time-consuming and laborious, and the inspectors are easily fatigued. In the case of poor condition, the misrecognition rate is high.

[0003] In the prior art, the patent disclosed in CN201910858868.8 provides a lipstick color recognition method and device, which determines the lipstick color in a facial image through hue, saturation and brightness information. Although this method improves the accuracy of lipstick recognition to a certain extent, it cannot perform effective and stable recognition when the light changes, the environment is complex, or there are partial lesions on the lips. Summary of the invention

[0004] The embodiments of the present invention provide a method, device, medium and equipment for lipstick recognition in ID photos based on a convolutional neural network to solve the problem that the prior art cannot effectively and stably recognize lipstick in ID photos when the light changes, the environment is complex, or there are partial lesions on the lips.

[0005] A lipstick recognition method for ID photos based on a convolutional neural network, the method comprising:

[0006] Obtain an image training set, preprocess each training sample in the image training set, and obtain an RGB space image of a lip area corresponding to the training sample;

[0007] Converting the RGB spatial image of the lip area corresponding to the training sample into a YCbCR spatial image, and forming a training sample pair with the RGB spatial image and the YCbCR spatial image;

[0008] Constructing a convolutional neural network, and using the training sample pairs to train the convolutional neural network to obtain a lipstick recognition model;

[0009] Obtain an image to be recognized, preprocess the image to be recognized, perform lipstick recognition on the preprocessed image to be recognized through the trained lipstick recognition model, obtain the category probability corresponding to the image to be recognized, and select the category label with the largest probability as the lipstick recognition result of the image to be recognized.

[0010] Optionally, acquiring an image training set includes:

[0011] Collect several lipstick-applying images of several users obtained from different angles by ID photo shooting equipment as lipstick-applying positive samples;

[0012] Collect several images of several users without lipstick taken from different angles by ID photo shooting equipment as negative samples of lipstick-applied users.

[0013] Optionally, preprocessing each training sample in the image training set to obtain an RGB space image of a lip area corresponding to the training sample includes:

[0014] For each training sample in the image training set, the SeetaFace face detection algorithm is used to extract the lip contour control points of the face area in the training sample based on the 68-point positioning method;

[0015] Using the PolyLine function to sequentially connect the lip contour control points to form a closed area;

[0016] Filling the closed area with a flooding method to form a lip mask;

[0017] Counting the total pixels of the lip area in the training sample, and creating a three-channel image according to the total pixels;

[0018] The pixels corresponding to the RGB three channels of the lip mask are obtained, the pixels corresponding to the RGB three channels are filled into the three-channel image, and the length and width of the pixel-filled three-channel image are converted into preset pixel values ​​through interpolation transformation to obtain the RGB spatial image of the lip area corresponding to the training sample.

[0019] Optionally, counting the total pixels of the lip region in the training sample and creating a three-channel image according to the total pixels comprises:

[0020] The total number of pixels N in the lip region of the training sample is counted to create a three-channel image with a length and a width of (n, n), where n=floor(N).

[0021] Optionally, the convolutional neural network includes a first group of convolutions, a second group of convolutions, a third group of convolutions, a first fully connected layer, and a second fully connected layer, the first group of convolutions is used to extract features from RGB spatial images, and the second group of convolutions is used to extract features from YCbCR spatial images;

[0022] The first convolution and the second convolution both include four layers, wherein the first convolution layer uses a convolution kernel of size (3,3,3,32) to convolve the input image to obtain a feature map of size (126,126,32), then uses the LeakyRelu activation function for activation, and then uses the maximum pooling convolution to perform a pooling operation to obtain a first feature map of size (63,63,32); the second convolution layer uses a convolution kernel of size (3,3,36,64) to convolve the first feature map to obtain a feature map of size (61,61,64), then uses the LeakyRelu activation function for activation, and then uses the maximum pooling convolution to perform a pooling operation to obtain a first feature map of size (30,30,6 4); the third convolution layer uses a convolution kernel with a size of (3,3,64,128) to convolve the second feature map to obtain a feature map with a size of (28,28,128), then uses the LeakyRelu activation function to activate it, and then uses the maximum pooling convolution to perform a pooling operation to obtain a third feature map with a size of (14,14,128); the fourth convolution layer uses a convolution kernel with a size of (3,3,128,64) to convolve the third feature map to obtain a feature map with a size of (12,12,128), then uses the LeakyRelu activation function to activate it, and then uses the maximum pooling convolution to perform a pooling operation to obtain a fourth feature map with a size of (6,6,64);

[0023] The third group of convolutions is used to concatenate the feature image extracted from the RGB spatial image and the feature image extracted from the YCbCR spatial image to obtain a (6,6,128)-dimensional feature map, and convolve the (6,6,128)-dimensional feature map using a convolution kernel of size (3,3,128,64) to obtain a (4,4,64)-dimensional feature map, and activate and expand it using the LeakyRelu activation function to obtain a (4*4*64)-dimensional feature vector;

[0024] The (4*4*64)-dimensional feature vector is used as the input of the first fully connected layer. After passing through the first fully connected layer, a 1024-dimensional feature vector is obtained, and the LeakyRelu activation function is used for activation. Then, the 2-dimensional feature vector is obtained after passing through the second fully connected layer.

[0025] The 2D feature vector is normalized by SoftMax to obtain the category probability.

[0026] Optionally, the convolutional neural network uses cross entropy as a training loss function.

[0027] Optionally, the adopting the training sample pairs to train the convolutional neural network to obtain a lipstick recognition model comprises:

[0028] Input 128 training sample pairs as a batch into the convolutional neural network, use the SGD optimizer to optimize the cross entropy of each batch, and perform back propagation, and stop iteration when the loss cost of the convolutional neural network drops to a preset accuracy;

[0029] The 64 training sample pairs are input as a batch into the convolutional neural network for testing.

[0030] Optionally, the method further comprises:

[0031] For the image to be identified whose lipstick recognition result is that the lipstick is not applied, extract the average color of the area without lipstick applied;

[0032] Extracting the image texture gradient of the lip area of ​​the image to be recognized, and extracting the Alpha mask of the lip area;

[0033] Performing color correction on the lip region of the image to be recognized;

[0034] Among them, the formula for color correction is:

[0035] g(x)=(1-α)f0(x)+αf1(x)

[0036] Where f0(x) represents the original pixel of the lip area, and f1(x) represents The average color of , g(x) represents the image after color correction, and α represents the fusion ratio of the Alpha mask.

[0037] A lipstick recognition device for ID photos based on a convolutional neural network, the device comprising:

[0038] A preprocessing module, used to obtain an image training set, preprocess each training sample in the image training set, and obtain an RGB space image of a lip area corresponding to the training sample;

[0039] A conversion module, used for converting the RGB spatial image of the lip area corresponding to the training sample into a YCbCR spatial image, and forming a training sample pair with the RGB spatial image and the YCbCR spatial image;

[0040] A training module, used for constructing a convolutional neural network, and using the training sample pairs to train the convolutional neural network to obtain a lipstick recognition model;

[0041] The recognition module is used to obtain an image to be recognized, preprocess the image to be recognized, perform lipstick recognition on the preprocessed image to be recognized through the trained lipstick recognition model, obtain the category probability corresponding to the image to be recognized, and select the category label with the largest probability as the lipstick recognition result of the image to be recognized.

[0042] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for recognizing lipstick in a ID photo based on a convolutional neural network is implemented.

[0043] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for recognizing lipstick in ID photos based on a convolutional neural network as described above is implemented.

[0044] The embodiment of the present invention obtains an image training set, preprocesses each training sample in the image training set, and obtains an RGB space image of the lip area corresponding to the training sample; converts the RGB space image of the lip area corresponding to the training sample into a YCbCR space image, and forms a training sample pair with the RGB space image and the YCbCR space image; constructs a convolutional neural network, uses the training sample pair to train the convolutional neural network, and obtains a lipstick recognition model; obtains an image to be recognized, preprocesses the image to be recognized, and performs lipstick recognition on the preprocessed image to be recognized by the trained lipstick recognition model to obtain the category probability corresponding to the image to be recognized, and selects the category label with the largest probability as the lipstick recognition result of the image to be recognized. The present invention uses a multidimensional information fusion convolutional neural network based on RGB space and YCbCr space to identify whether the ID photo is smeared with lipstick, which effectively improves the robustness and accuracy of lipstick recognition, and can also effectively and stably recognize the lipstick in the ID photo when the light changes, the environment is complex, and the lips have partial lesions. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0046] Figure 1 It is a flow chart of a method for recognizing lipstick in ID photos based on a convolutional neural network provided by an embodiment of the present invention;

[0047] Figure 2 It is a flowchart for implementing the preprocessing in step S101 of the lipstick recognition method for ID photos based on a convolutional neural network provided by an embodiment of the present invention;

[0048] Figure 3 is a schematic diagram of the structure of a convolutional neural network provided by an embodiment of the present invention;

[0049] Figure 4 It is a flowchart for implementing model training in step S103 of a lipstick recognition method for ID photos based on a convolutional neural network provided by an embodiment of the present invention;

[0050] Figure 5 It is a structural schematic diagram of a lipstick recognition device for ID photos based on a convolutional neural network provided by an embodiment of the present invention;

[0051] Figure 6 is a schematic diagram of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION

[0052] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0053] The embodiment of the present invention uses a multi-dimensional information fusion convolutional neural network based on RGB space and YCbCr space to identify whether lipstick is applied on the ID photo, thereby effectively improving the robustness and accuracy of lipstick recognition. When the light changes, the environment is complex, or there are partial lesions on the lips, the lipstick in the ID photo can also be effectively and stably recognized.

[0054] The following is a detailed description of the lipstick recognition method for ID photos based on a convolutional neural network provided in this embodiment. Figure 1 As shown, the lipstick recognition method for ID photos based on convolutional neural network includes:

[0055] In step S101, an image training set is obtained, and each training sample in the image training set is preprocessed to obtain an RGB space image of a lip region corresponding to the training sample.

[0056] Optionally, as a preferred example of the present invention, the step of obtaining an image training set in step S101 includes:

[0057] Collect several lipstick-applying images of several users obtained from different angles by ID photo shooting equipment as lipstick-applying positive samples;

[0058] Collect several images of several users without lipstick taken from different angles by ID photo shooting equipment as negative samples of lipstick-applied users.

[0059] Exemplarily, an embodiment of the present invention collects five photos of lipstick applied at different angles provided by 200 volunteers taken by ID photo equipment, for a total of 1000 images as positive samples of lipstick applied; and collects five photos of no lipstick applied at different angles provided by 200 volunteers taken by ID photo equipment, for a total of 1000 images as negative samples of lipstick applied.

[0060] Optionally, as a preferred example of the present invention, Figure 2 As shown, the step S101 of preprocessing each training sample in the image training set to obtain the RGB space image of the lip area corresponding to the training sample includes:

[0061] In step S201, for each training sample in the image training set, the SeetaFace face detection algorithm is used to extract lip contour control points of the face area in the training sample based on the 68-point positioning method.

[0062] Exemplarily, the embodiment of the present invention uses a 68-point positioning method to extract the 48th to 64th points of the face area as lip contour control points.

[0063] In step S202, the lip contour control points are sequentially connected using a PolyLine function to form a closed area.

[0064] In step S203, the closed area is filled with a flooding method to form a lip mask.

[0065] In step S204, the total pixels of the lip region in the training sample are counted, and a three-channel image is created according to the total pixels.

[0066] Here, for each training sample, the embodiment of the present invention counts the total pixels N in the lip area of ​​the training sample and creates a three-channel image x with a length and width of (n, n) i , where n = floor(N).

[0067] In step S205, the pixels corresponding to the RGB three channels of the lip mask are obtained, the pixels corresponding to the RGB three channels are filled into the three-channel image, and the length and width of the pixel-filled three-channel image are converted into preset pixel values ​​through interpolation transformation to obtain the RGB spatial image of the lip area corresponding to the training sample.

[0068] The preset pixel value is (128, 128). After obtaining a three-channel image x with a length and width of (n, n) iAfter that, the training sample is traversed, and the pixels corresponding to the three channels of R, G, and B in the lip area are extracted according to the lip mask, and the colors of the three channels of R, G, and B are filled into the three-channel image x with a length and width of (n, n) according to the traversal information. i ; Then for the three-channel image x i Perform interpolation transformation and transform its length and width to (128,128).

[0069] Similarly, the same operation is performed on each training sample in the image training set to obtain the RGB space image set X = {x1, ...x i , …x n} and its lipstick identification information Y = {y1, ...y i ,…y n}, where y i In practical applications, we can randomly select the RGB space image set X = {x1, ...x2, ... ... i , …x n Two-thirds of the samples are used for training, and the remaining one-third is used for testing.

[0070] In step S102, the RGB spatial image of the lip area corresponding to the training sample is converted into a YCbCr spatial image, and the RGB spatial image and the YCbCr spatial image form a training sample pair.

[0071] Since the color of light makeup lipstick is closer to that of no lipstick, in view of this, the embodiment of the present invention needs to synchronously convert the image features of the training samples into the YCbCr space to obtain a YCbCr space image, and then form the RGB space image and the YCbCr space image into a training sample pair, which are used together as the input of the convolutional neural network to enhance the convolutional neural network's ability to distinguish lipstick.

[0072] Among them, the conversion formula from RGB space image to YCbCr space image is as follows:

[0073] Y=0.299R+0.587G+0.114B

[0074] Cb=0.564(BY)

[0075] Cr=0.713(RY)

[0076] In the above formula, R represents the red component in the RGB color model, G represents the green component in the RGB color model, and B represents the blue component in the RGB color model. Y represents the brightness in the YCbCr color model, Cb represents the concentration offset component of blue, and Cr represents the concentration offset component of red.

[0077] In step S103, a convolutional neural network is constructed, and the convolutional neural network is trained using the training sample pairs to obtain a lipstick recognition model.

[0078] Figure 3 A schematic diagram of the structure of a convolutional neural network provided for an embodiment of the present invention. Here, the convolutional neural network includes a first group of convolutions, a second group of convolutions, a third group of convolutions, a first fully connected layer, and a second fully connected layer. The first group of convolutions is used to extract features from RGB spatial images, and the second group of convolutions is used to extract features from YCbCr spatial images. The embodiment of the present invention uses four layers of convolution to perform convolution operations on the RGB spatial image and the YCbCr spatial image in the input training sample pair, specifically:

[0079] The first convolution and the second convolution both include four layers, wherein the first convolution layer uses a convolution kernel of size (3,3,3,32) to convolve the input image to obtain a feature map of size (126,126,32), then uses the LeakyRelu activation function for activation, and then uses the maximum pooling convolution to perform a pooling operation to obtain a first feature map of size (63,63,32); the second convolution layer uses a convolution kernel of size (3,3,36,64) to convolve the first feature map to obtain a feature map of size (61,61,64), then uses the LeakyRelu activation function for activation, and then uses the maximum pooling convolution to perform a pooling operation to obtain a first feature map of size (30,30,6 4); the third convolution layer uses a convolution kernel with a size of (3,3,64,128) to convolve the second feature map to obtain a feature map with a size of (28,28,128), then activates it with the LeakyRelu activation function, and then uses the maximum pooling convolution to perform a pooling operation to obtain a third feature map with a size of (14,14,128); the fourth convolution layer uses a convolution kernel with a size of (3,3,128,64) to convolve the third feature map to obtain a feature map with a size of (12,12,128), then uses the LeakyRelu activation function to activate it, and then uses the maximum pooling convolution to perform a pooling operation to obtain a fourth feature map with a size of (6,6,64).

[0080] It should be understood that the first group of convolutions and the second group of convolutions have the same structure, and the same convolution method is used to perform convolution operations on the input RGB spatial image and YCbCR spatial image, respectively, to obtain feature maps with sizes of (6, 6, 64).

[0081] The third group of convolutions is used to concatenate the feature image extracted from the RGB spatial image and the feature image extracted from the YCbCr spatial image to obtain a (6,6,128)-dimensional feature map, and convolve the (6,6,128)-dimensional feature map using a convolution kernel of size (3,3,128,64) to obtain a (4,4,64)-dimensional feature map, and activate and expand it using the LeakyRelu activation function to obtain a (4*4*64)-dimensional feature vector;

[0082] The (4*4*64)-dimensional feature vector is used as the input of the first fully connected layer. After passing through the first fully connected layer, a 1024-dimensional feature vector is obtained, and the LeakyRelu activation function is used for activation. Then, the 2-dimensional feature vector is obtained after passing through the second fully connected layer.

[0083] The 2D feature vector is normalized by SoftMax and the category probability is finally obtained.

[0084] Optionally, as a preferred example of the present invention, the convolutional neural network uses cross entropy as a training loss function. The calculation method is as follows:

[0085] H(p,q)=-∑(p(x)logq(x)).

[0086] Among them, H(p,q) represents the output loss value of the convolutional neural network for sample x, p(x) is the category vector value corresponding to sample x, q(x) is the predicted category vector value of sample x, and log(·) represents the logarithmic function.

[0087] After the convolutional neural network is constructed, the convolutional neural network is trained. Optionally, as a preferred embodiment of the present invention, Figure 4 As shown, the convolutional neural network is trained using the training sample pairs to obtain a lipstick recognition model, which includes:

[0088] In step S401, 128 training sample pairs are input into the convolutional neural network as a batch, the cross entropy of each batch is optimized using the SGD optimizer, and back propagation is performed, and the iteration is stopped when the loss cost of the convolutional neural network drops to a preset accuracy.

[0089] In step S402, 64 training sample pairs are input as a batch into the convolutional neural network for testing.

[0090] Here, the embodiment of the present invention uses 128 training sample pairs per batch to iteratively train the convolutional neural network, and after stopping the iteration, uses 64 training sample pairs per batch to test the convolutional neural network, and observes the accuracy and recall rate of the test set. The finally trained convolutional neural network is recorded as a lipstick recognition model.

[0091] In step S104, an image to be recognized is obtained, the image to be recognized is preprocessed, lipstick recognition is performed on the preprocessed image to be recognized using the trained lipstick recognition model to obtain the category probability corresponding to the image to be recognized, and the category label with the largest probability is selected as the lipstick recognition result of the image to be recognized.

[0092] The preprocessing of the image to be recognized is the same as that in the above embodiment. For details, please refer to the description of the above embodiment and will not be repeated here. The preprocessed image to be recognized is input into the lipstick recognition model to obtain the category probability corresponding to the image to be recognized output by the lipstick recognition model. As mentioned above, the category probability is a 2-dimensional feature vector, that is, a prediction vector The prediction vector Including the probability of wearing lipstick and the probability of not wearing lipstick The embodiment of the present invention calculates the category label δ corresponding to the image to be identified according to the category probability, where

[0093] In summary, the present invention uses a multidimensional information fusion convolutional neural network based on RGB spatial images and YCbCr spatial images to identify whether the ID photo is painted with lipstick, which effectively improves the robustness and accuracy of lipstick recognition. When the light changes, the environment is complex, and there are partial lesions on the lips, the lipstick in the ID photo can also be effectively and stably recognized.

[0094] Optionally, as a preferred example of the present invention, for an image to be identified whose lipstick identification information is that lipstick has been applied, if the recognition result output by the lipstick recognition model is that lipstick is not applied, this embodiment may further correct the image to be identified, and the method further includes:

[0095] For the image to be identified whose lipstick recognition result is that the lipstick is not applied, extract the average color of the area without lipstick applied;

[0096] Extracting the image texture gradient of the lip area of ​​the image to be recognized, and extracting the Alpha mask of the lip area;

[0097] Performing color correction on the lip region of the image to be recognized;

[0098] Among them, the formula for color correction is:

[0099] g(x)=(1-α)f0(x)+αf1(x)

[0100] Where f0(x) represents the original pixel of the lip area, and f1(x) represents The average color of , g(x) represents the image after color correction, and α represents the fusion ratio of the Alpha mask.

[0101] Here, the average color of the unapplied lipstick area is extracted As the color input for lipstick color correction, the lip area in the image to be identified is color corrected, and the extracted image texture gradient is used to keep its texture features unchanged when correcting the lip color, thereby automatically correcting the lip color of users who do not meet the standards.

[0102] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.

[0103] In one embodiment, the present invention further provides a lipstick identification device for ID photos based on a convolutional neural network, and the lipstick identification device for ID photos based on a convolutional neural network corresponds one-to-one to the lipstick identification method for ID photos based on a convolutional neural network in the above embodiment. Figure 5 As shown, the lipstick recognition device for ID photos based on convolutional neural network includes a preprocessing module 51, a conversion module 52, a training module 53, and a recognition module 54. The detailed description of each functional module is as follows:

[0104] A preprocessing module 51 is used to obtain an image training set, preprocess each training sample in the image training set, and obtain an RGB space image of the lip area corresponding to the training sample;

[0105] A conversion module 52, configured to convert the RGB spatial image of the lip region corresponding to the training sample into a YCbCR spatial image, and to form a training sample pair with the RGB spatial image and the YCbCR spatial image;

[0106] A training module 53 is used to construct a convolutional neural network, and train the convolutional neural network using the training sample pairs to obtain a lipstick recognition model;

[0107] The recognition module 54 is used to obtain the image to be recognized, preprocess the image to be recognized, perform lipstick recognition on the preprocessed image to be recognized through the trained lipstick recognition model, obtain the category probability corresponding to the image to be recognized, and select the category label with the largest probability as the lipstick recognition result of the image to be recognized.

[0108] Optionally, the preprocessing module 51 includes:

[0109] A sample collection unit, used to collect a number of lipstick-applying images of a number of users obtained from different angles by a ID photo shooting device as lipstick-applying positive samples;

[0110] Collect several images of several users without lipstick taken from different angles by ID photo shooting equipment as negative samples of lipstick-applied users.

[0111] Optionally, the preprocessing module 51 includes:

[0112] An extraction unit is used to extract lip contour control points of the face area in the training sample based on a 68-point positioning method using a SeetaFace face detection algorithm for each training sample in the image training set;

[0113] A fitting unit, used for sequentially connecting the lip contour control points using a PolyLine function to form a closed area;

[0114] A mask forming unit, used for filling the closed area with a flooding method to form a lip mask;

[0115] An image creation unit, used for counting the total pixels of the lip area in the training sample, and creating a three-channel image according to the total pixels;

[0116] The image generation unit is used to obtain pixels corresponding to the RGB three channels of the lip mask, fill the pixels corresponding to the RGB three channels into the three-channel image, and convert the length and width of the pixel-filled three-channel image into preset pixel values ​​through interpolation transformation to obtain the RGB spatial image of the lip area corresponding to the training sample.

[0117] Optionally, the image creation unit is specifically used for:

[0118] The total number of pixels N in the lip region of the training sample is counted to create a three-channel image with a length and a width of (n, n), where n=floor(N).

[0119] Optionally, the convolutional neural network includes a first group of convolutions, a second group of convolutions, a third group of convolutions, a first fully connected layer, and a second fully connected layer, the first group of convolutions is used to extract features from RGB spatial images, and the second group of convolutions is used to extract features from YCbCR spatial images;

[0120] The first convolution and the second convolution both include four layers, wherein the first convolution layer uses a convolution kernel of size (3,3,3,32) to convolve the input image to obtain a feature map of size (126,126,32), then uses the LeakyRelu activation function for activation, and then uses the maximum pooling convolution to perform a pooling operation to obtain a first feature map of size (63,63,32); the second convolution layer uses a convolution kernel of size (3,3,36,64) to convolve the first feature map to obtain a feature map of size (61,61,64), then uses the LeakyRelu activation function for activation, and then uses the maximum pooling convolution to perform a pooling operation to obtain a first feature map of size (30,30,6 4); the third convolution layer uses a convolution kernel with a size of (3,3,64,128) to convolve the second feature map to obtain a feature map with a size of (28,28,128), then uses the LeakyRelu activation function to activate it, and then uses the maximum pooling convolution to perform a pooling operation to obtain a third feature map with a size of (14,14,128); the fourth convolution layer uses a convolution kernel with a size of (3,3,128,64) to convolve the third feature map to obtain a feature map with a size of (12,12,128), then uses the LeakyRelu activation function to activate it, and then uses the maximum pooling convolution to perform a pooling operation to obtain a fourth feature map with a size of (6,6,64);

[0121] The third group of convolutions is used to concatenate the feature image extracted from the RGB spatial image and the feature image extracted from the YCbCR spatial image to obtain a (6,6,128)-dimensional feature map, and convolve the (6,6,128)-dimensional feature map using a convolution kernel of size (3,3,128,64) to obtain a (4,4,64)-dimensional feature map, and activate and expand it using the LeakyRelu activation function to obtain a (4*4*64)-dimensional feature vector;

[0122] The (4*4*64)-dimensional feature vector is used as the input of the first fully connected layer. After passing through the first fully connected layer, a 1024-dimensional feature vector is obtained, and the LeakyRelu activation function is used for activation. Then, the 2-dimensional feature vector is obtained after passing through the second fully connected layer.

[0123] The 2D feature vector is normalized by SoftMax to obtain the category probability.

[0124] Optionally, the convolutional neural network uses cross entropy as a training loss function.

[0125] Optionally, the training module 53 includes:

[0126] A training unit, used for inputting 128 training sample pairs as a batch into the convolutional neural network, optimizing the cross entropy of each batch using an SGD optimizer, and performing back propagation, and stopping iteration when the loss cost of the convolutional neural network drops to a preset accuracy;

[0127] The testing unit is used to input 64 training sample pairs as a batch into the convolutional neural network for testing.

[0128] Optionally, the device further comprises:

[0129] The average color extraction unit is used to extract the average color of the unapplied lipstick area for the image to be identified whose lipstick identification result is that the lipstick is not applied;

[0130] A mask extraction unit, used to extract the image texture gradient of the lip area of ​​the image to be recognized, and extract the Alpha mask of the lip area;

[0131] A color correction unit, used for performing color correction on the lip region of the image to be recognized;

[0132] Among them, the formula for color correction is:

[0133] g(x)=(1-α)f0(x)+αf1(x)

[0134] Where f0(x) represents the original pixel of the lip area, and f1(x) represents The average color of , g(x) represents the image after color correction, and α represents the fusion ratio of the Alpha mask.

[0135] For the specific definition of the ID photo lipstick recognition device based on convolutional neural network, please refer to the definition of the ID photo lipstick recognition method based on convolutional neural network above, which will not be repeated here. Each module in the above-mentioned ID photo lipstick recognition device based on convolutional neural network can be implemented in whole or in part by software, hardware and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0136] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 6As shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a lipstick recognition method for a ID photo based on a convolutional neural network is implemented.

[0137] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program:

[0138] Obtain an image training set, preprocess each training sample in the image training set, and obtain an RGB space image of a lip area corresponding to the training sample;

[0139] Converting the RGB spatial image of the lip area corresponding to the training sample into a YCbCR spatial image, and forming a training sample pair with the RGB spatial image and the YCbCR spatial image;

[0140] Constructing a convolutional neural network, and using the training sample pairs to train the convolutional neural network to obtain a lipstick recognition model;

[0141] Obtain an image to be recognized, preprocess the image to be recognized, perform lipstick recognition on the preprocessed image to be recognized through the trained lipstick recognition model, obtain the category probability corresponding to the image to be recognized, and select the category label with the largest probability as the lipstick recognition result of the image to be recognized.

[0142] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0143] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0144] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A lipstick recognition method for ID photos based on convolutional neural network, characterized in that: include: Obtain an image training set, preprocess each training sample in the image training set, and obtain an RGB space image of a lip area corresponding to the training sample; Converting the RGB spatial image of the lip area corresponding to the training sample into a YCbCR spatial image, and forming a training sample pair with the RGB spatial image and the YCbCR spatial image; Constructing a convolutional neural network, and using the training sample pairs to train the convolutional neural network to obtain a lipstick recognition model; Acquire an image to be recognized, preprocess the image to be recognized, perform lipstick recognition on the preprocessed image to be recognized using the trained lipstick recognition model, obtain the category probability corresponding to the image to be recognized, and select the category label with the largest probability as the lipstick recognition result of the image to be recognized; The convolutional neural network includes a first group of convolutions, a second group of convolutions, a third group of convolutions, a first fully connected layer, and a second fully connected layer, the first group of convolutions is used to extract features from RGB spatial images, and the second group of convolutions is used to extract features from YCbCR spatial images; The first convolution and the second convolution both include four layers, wherein the first convolution layer uses a convolution kernel of size (3,3,3,32) to convolve the input image to obtain a feature map of size (126,126,32), then uses the LeakyRelu activation function for activation, and then uses the maximum pooling convolution to perform a pooling operation to obtain a first feature map of size (63,63,32); the second convolution layer uses a convolution kernel of size (3,3,36,64) to convolve the first feature map to obtain a feature map of size (61,61,64), then uses the LeakyRelu activation function for activation, and then uses the maximum pooling convolution to perform a pooling operation to obtain a first feature map of size (30,30,6 4); the third convolution layer uses a convolution kernel of size (3,3,64,128) to convolve the second feature map to obtain a feature map of size (28,28,128), then uses the LeakyRelu activation function to activate it, and then uses the maximum pooling convolution to perform pooling operation to obtain a third feature map of size (14,14,128); the fourth convolution layer uses a convolution kernel of size (3,3,128,64) to convolve the third feature map to obtain a feature map of size (12,12,128), then uses the LeakyRelu activation function to activate it, and then uses the maximum pooling convolution to perform pooling operation to obtain a fourth feature map of size (6,6,64); The third group of convolutions is used to concatenate the feature image extracted from the RGB spatial image and the feature image extracted from the YCbCR spatial image to obtain a (6,6,128)-dimensional feature map, and convolve the (6,6,128)-dimensional feature map using a convolution kernel of size (3,3,128,64) to obtain a (4,4,64)-dimensional feature map, and activate and expand it using the LeakyRelu activation function to obtain a (4*4*64)-dimensional feature vector; The (4*4*64)-dimensional feature vector is used as the input of the first fully connected layer. After passing through the first fully connected layer, a 1024-dimensional feature vector is obtained, and the LeakyRelu activation function is used for activation. Then, the 2-dimensional feature vector is obtained after passing through the second fully connected layer. The 2D feature vector is normalized by SoftMax to obtain the category probability.

2. The lipstick recognition method for ID photos based on convolutional neural network as claimed in claim 1, characterized in that: The acquiring of the image training set comprises: Collect several lipstick-applying images of several users obtained from different angles by ID photo shooting equipment as lipstick-applying positive samples; Collect several images of several users without lipstick taken from different angles by ID photo shooting equipment as negative samples of lipstick-applied users.

3. The lipstick recognition method for ID photos based on convolutional neural network as claimed in claim 1, characterized in that: The preprocessing of each training sample in the image training set to obtain an RGB spatial image of the lip area corresponding to the training sample comprises: For each training sample in the image training set, the SeetaFace face detection algorithm is used to extract the lip contour control points of the face area in the training sample based on the 68-point positioning method; Using the PolyLine function to sequentially connect the lip contour control points to form a closed area; Filling the closed area with a flooding method to form a lip mask; Counting the total pixels of the lip area in the training sample, and creating a three-channel image according to the total pixels; The pixels corresponding to the RGB three channels of the lip mask are obtained, the pixels corresponding to the RGB three channels are filled into the three-channel image, and the length and width of the pixel-filled three-channel image are converted into preset pixel values ​​through interpolation transformation to obtain the RGB spatial image of the lip area corresponding to the training sample.

4. The lipstick recognition method for ID photos based on convolutional neural network as claimed in claim 3, characterized in that: The counting of total pixels in the lip region of the training sample and creating a three-channel image according to the total pixels comprises: The total number of pixels N in the lip area of ​​the training sample is counted to create a three-channel image with a length and width of (n, n), where .

5. The method for identifying lipstick in ID photos based on convolutional neural network as claimed in claim 1, characterized in that: The convolutional neural network uses cross entropy as a training loss function.

6. The method for identifying lipstick in ID photos based on convolutional neural network as claimed in claim 1, characterized in that: The step of training the convolutional neural network using the training sample pairs to obtain a lipstick recognition model comprises: Input 128 training sample pairs as a batch into the convolutional neural network, use the SGD optimizer to optimize the cross entropy of each batch, and perform back propagation, and stop iteration when the loss cost of the convolutional neural network drops to a preset accuracy; The 64 training sample pairs are input as a batch into the convolutional neural network for testing.

7. The method for identifying lipstick in ID photos based on convolutional neural network as claimed in claim 1, characterized in that: The method further comprises: For the image to be identified whose lipstick recognition result is that the lipstick is not applied, extract the average color of the area without lipstick applied; Extracting the image texture gradient of the lip area of ​​the image to be recognized, and extracting the Alpha mask of the lip area; Performing color correction on the lip region of the image to be recognized; Among them, the formula for color correction is: ; in Represents the original pixels of the lip area, express The average color of represents the color-corrected image, Indicates the blending ratio of the Alpha mask.

8. A lipstick recognition device for ID photos based on convolutional neural network, characterized in that: include: A preprocessing module, used to obtain an image training set, preprocess each training sample in the image training set, and obtain an RGB space image of a lip area corresponding to the training sample; A conversion module, used for converting the RGB spatial image of the lip area corresponding to the training sample into a YCbCR spatial image, and forming a training sample pair with the RGB spatial image and the YCbCR spatial image; A training module, used for constructing a convolutional neural network, and using the training sample pairs to train the convolutional neural network to obtain a lipstick recognition model; A recognition module is used to obtain an image to be recognized, preprocess the image to be recognized, perform lipstick recognition on the preprocessed image to be recognized using the trained lipstick recognition model, obtain the category probability corresponding to the image to be recognized, and select the category label with the largest probability as the lipstick recognition result of the image to be recognized; The convolutional neural network includes a first group of convolutions, a second group of convolutions, a third group of convolutions, a first fully connected layer, and a second fully connected layer, the first group of convolutions is used to extract features from RGB spatial images, and the second group of convolutions is used to extract features from YCbCR spatial images; The first convolution and the second convolution both include four layers, wherein the first convolution layer uses a convolution kernel of size (3,3,3,32) to convolve the input image to obtain a feature map of size (126,126,32), then uses the LeakyRelu activation function for activation, and then uses the maximum pooling convolution to perform a pooling operation to obtain a first feature map of size (63,63,32); the second convolution layer uses a convolution kernel of size (3,3,36,64) to convolve the first feature map to obtain a feature map of size (61,61,64), then uses the LeakyRelu activation function for activation, and then uses the maximum pooling convolution to perform a pooling operation to obtain a first feature map of size (30,30,6 4); the third convolution layer uses a convolution kernel of size (3,3,64,128) to convolve the second feature map to obtain a feature map of size (28,28,128), then uses the LeakyRelu activation function to activate it, and then uses the maximum pooling convolution to perform pooling operation to obtain a third feature map of size (14,14,128); the fourth convolution layer uses a convolution kernel of size (3,3,128,64) to convolve the third feature map to obtain a feature map of size (12,12,128), then uses the LeakyRelu activation function to activate it, and then uses the maximum pooling convolution to perform pooling operation to obtain a fourth feature map of size (6,6,64); The third group of convolutions is used to concatenate the feature image extracted from the RGB spatial image and the feature image extracted from the YCbCR spatial image to obtain a (6,6,128)-dimensional feature map, and convolve the (6,6,128)-dimensional feature map using a convolution kernel of size (3,3,128,64) to obtain a (4,4,64)-dimensional feature map, and activate and expand it using the LeakyRelu activation function to obtain a (4*4*64)-dimensional feature vector; The (4*4*64)-dimensional feature vector is used as the input of the first fully connected layer. After passing through the first fully connected layer, a 1024-dimensional feature vector is obtained, and the LeakyRelu activation function is used for activation. Then, the 2-dimensional feature vector is obtained after passing through the second fully connected layer. The 2D feature vector is normalized by SoftMax to obtain the category probability.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for lipstick recognition in ID photos based on a convolutional neural network is implemented as described in any one of claims 1 to 7.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for recognizing lipstick in ID photos based on a convolutional neural network is implemented as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Automatic identification photo shooting method and device

    CN110536044A

  • Color number identification method and device, computer equipment and storage medium

    CN109934092A

  • A method for identifying lipstick number

    CN110348530A