Method, device and equipment for generating standard identification photo of any region and storage medium

By automatically obtaining the standard parameters of the ID license in the target area and making corresponding adjustments, the problems of low efficiency and poor accuracy of ID license generation in the prior art are solved, and efficient and accurate generation of standard ID licenses in any area are achieved.

CN120013828APending Publication Date: 2025-05-16BEIJING THUNDERSTONE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510081247.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing ID photo generation system can only support fixed sizes and formats. Users need to manually adjust the size, head position and background color of the photo, which is inefficient and prone to errors. Users need to choose the ID photo standards for the target area themselves.

Method used

Provide a method for generating standard ID photos in any area. By obtaining the target area entered by the user, it automatically obtains the standard parameters of the ID photos, adjusts the size, head proportion and head posture of the photo, adjusts the background, and finally corrects the facial expression to generate the final standard ID photos.

Benefits of technology

It improves the efficiency and accuracy of standard certificate licenses in any area, reduces the user's professional requirements, avoids the error of manual adjustment, and ensures the accuracy of the generated certificate licenses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013828A_ABST
    Figure CN120013828A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and provides a method for generating a standard identification photo in any region, and the method comprises the steps: obtaining an identification photo standard parameter of a target region according to the target region inputted by a user; adjusting the size, the head ratio and the head posture of the to-be-processed photo according to the identification photo standard parameters of the target area to obtain a first standard identification photo of the to-be-processed photo corresponding to the target area; adjusting the background of the first standard identification photo according to the identification photo standard parameters of the target area to obtain a second standard identification photo of the to-be-processed photo corresponding to the target area; correcting the facial expression of the face corresponding to the face area of the second standard identification photo; and outputting the image after facial expression correction as a final standard identification photo. According to the technical scheme, the generation efficiency and accuracy of the standard identification photo in any area can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device and storage medium for generating a standard ID photo of an arbitrary area. Background Art

[0002] With the deepening of economic globalization, it is becoming more and more common for people to travel to other countries for the purpose of studying, living, working, etc. Whether it is daily life or study and work needs, there are always scenarios in some countries or regions where ID photos are needed. These ID photos are photos used on various certificates (such as identity cards, passports, driver's licenses, student cards, etc.). Since the ID photos of various regions (including countries or regions) may not be unified, people often need to process the photos into standard ID photos of the target area through some means according to the requirements of standard ID photos of the target area after obtaining them.

[0003] Generally speaking, most existing ID photo generation systems only support fixed sizes and formats, and users need to manually adjust the photo size, head position, background color, etc. according to the standards of different countries. However, this manual adjustment of photos is not only inefficient and prone to errors, but also requires users to select the ID photo standard of the target area, which is a big challenge for some users. Summary of the invention

[0004] The present application provides a method, device and storage medium for generating a standard ID photo for any area, which can improve the generation efficiency and accuracy of the standard ID photo for any area.

[0005] On the one hand, the present application provides a method for generating a standard ID photo of any area, the method comprising:

[0006] According to the target area input by the user, obtain the standard parameters of the ID photo of the target area;

[0007] Adjust the size, head proportion and head posture of the photo to be processed according to the standard parameters of the ID photo of the target area, so as to obtain a first standard ID photo of the photo to be processed corresponding to the target area;

[0008] Adjusting the background of the first standard ID photo according to the standard ID photo parameters of the target area to obtain a second standard ID photo of the to-be-processed photo corresponding to the target area;

[0009] Correcting the facial expression of the face corresponding to the face area of ​​the second standard ID photo;

[0010] Output the image after facial expression correction as the final standard ID photo.

[0011] On the other hand, the present application provides a device for generating a standard ID photo of any area, the device comprising:

[0012] An acquisition module, used to acquire the ID photo standard parameters of the target area according to the target area input by the user;

[0013] A first adjustment module is used to adjust the size, head proportion and head posture of the photo to be processed according to the standard parameters of the ID photo of the target area, so as to obtain a first standard ID photo of the photo to be processed corresponding to the target area;

[0014] A second adjustment module, configured to adjust the background of the first standard ID photo according to the ID photo standard parameters of the target area, so as to obtain a second standard ID photo of the to-be-processed photo corresponding to the target area;

[0015] A correction module, used to correct the facial expression of the face corresponding to the face area of ​​the second standard ID photo;

[0016] The output module is used to output the image after facial expression correction as the final standard ID photo.

[0017] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the technical solution of the method for generating a standard ID photo for any area as described above are implemented.

[0018] In a fourth aspect, the present application provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of the technical solution of the method for generating a standard ID photo for any area as described above.

[0019] From the technical solution provided by the present application, it can be seen that, on the one hand, compared with the challenge of the prior art requiring users to input standard parameters of ID photos of the target area, the technical solution of the present application only requires users to input the target area to obtain the standard parameters of ID photos of the target area, which greatly reduces the professional requirements for users and greatly facilitates users; on the other hand, based on the standard parameters of ID photos of the target area, the first standard ID photo, the second standard ID photo and even the final standard ID photo of the target area corresponding to the photo to be processed can be obtained, which can be automatically completed based on a computer program (for example, a large artificial intelligence model) without manual participation or intervention, which not only improves the efficiency of generating standard ID photos of any area, but also minimizes errors or mistakes, so that the accuracy of generating standard ID photos of any area is guaranteed. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0021] Figure 1 It is a flowchart of a method for generating a standard ID photo of any area provided in an embodiment of the present application;

[0022] Figure 2 It is a structural schematic diagram of a device for generating a standard ID photo of any area provided in an embodiment of the present application;

[0023] Figure 3 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0024] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0025] In this specification, adjectives such as first and second may be used only to distinguish one element or action from another element or action, without necessarily requiring or implying any actual such relationship or order. Where circumstances permit, reference to an element or component or step (etc.) should not be interpreted as being limited to only one of the elements, components, or steps, but may be one or more of the elements, components, or steps, etc.

[0026] In this specification, for the convenience of description, the sizes of various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0027] ID photos refer to photos used on various types of certificates (such as identity cards, passports, driver's licenses, student cards, etc.). Since ID photos in different regions (including countries or regions) may not be unified, people often need to process the photos into standard ID photos of the target area through some means according to the requirements of the standard ID photos of the target area after obtaining them. Generally speaking, most existing ID photo generation systems only support fixed sizes and formats, and users need to manually adjust the size, head position, and background color of the photos according to the standards of different countries. However, this manual adjustment of photos is not only inefficient and prone to errors, but also requires users to select the ID photo standard of the target area by themselves, which is a considerable challenge for some users.

[0028] In view of the above problems in the prior art, this application proposes a method for generating a standard ID photo in any area, and its flow chart is shown in the attached figure. Figure 1 As shown, it mainly includes steps S101 to S105, which are described in detail as follows:

[0029] Step S101: according to the target area input by the user, obtain the standard parameters of the ID photo of the target area.

[0030] Here, the target area refers to the standard ID photo or the country or region to which the ID photo standard belongs that the photo to be processed is to be processed. Selecting the target area means that the photo to be processed needs to be processed into an ID photo that meets the ID photo standard of the target area. Unlike the prior art that requires users to select the ID photo standard of the target area, the technical solution of this application only requires users to input the target area, such as the United States. The system extracts the ID photo standard parameter set of the target area based on the preset database, including the ID photo size (i.e., the width and height of the ID photo), the head proportion (i.e., the proportion of the head height in the ID photo to the total height of the image), the background color or texture parameters, and the facial angle requirements (i.e., the range of the head tilt angle in the ID photo), etc.

[0031] Step S102: adjusting the size, head proportion and head posture of the photo to be processed according to the standard parameters of the ID photo of the target area, to obtain a first standard ID photo of the photo to be processed corresponding to the target area.

[0032] As an embodiment of the present application, the size, head proportion and head posture of the photo to be processed are adjusted according to the standard parameters of the ID photo of the target area, and obtaining the first standard ID photo of the photo to be processed corresponding to the target area can be achieved by steps S1021 to S1025, which are described in detail as follows:

[0033] Step S1021: According to the height and width of the ID photo of the target area, the size of the photo to be processed is adjusted to the height and width of the ID photo of the target area.

[0034] Step S1022: According to the head proportion of the ID photo in the target area, the height of the head in the height-adjusted photo to be processed is adjusted to the height of the head in the ID photo in the target area.

[0035] After step S1021, the height and width of the photo to be processed have been adjusted to the height and width of the ID photo in the target area. Then, according to the head proportion of the ID photo in the target area, the height of the head in the photo to be processed with the adjusted height is adjusted to the height of the head in the ID photo in the target area, that is, the height of the photo to be processed with the height adjusted to the height of the ID photo in the target area is multiplied by the head proportion of the ID photo in the target area. The height of the head obtained at this time is the height of the head in the ID photo in the target area. Since the height of the photo to be processed has been adjusted to the height of the ID photo in the target area, and the height of the head therein is adjusted to the height of the head in the ID photo in the target area, this means that the head proportion of the photo to be processed is equal to the head proportion of the ID photo in the target area, which has met the requirements of the target area for the head proportion of the ID photo.

[0036] Step S1023: locate the face area of ​​the photo to be processed.

[0037] From the perspective of image processing, the face area is actually a closed geometric figure, that is, the face area has a boundary or edge. Therefore, an edge detection algorithm, such as an edge detection algorithm based on the Sobel operator, the Canny operator or other operators, can be used to detect the face area in the photo to be processed, thereby locating the face area.

[0038] Step S1024: extracting facial feature points from the face area to obtain the coordinates of the facial feature points.

[0039] Facial landmarks are key representative points on the face, such as the corners of the mouth, corners of the eyes, tip of the nose, eyebrows, etc. Facial landmarks are relatively stable, unlike the overall appearance of the face, which may be greatly affected by factors such as hairstyle and makeup. In the process of generating ID photos, operations based on facial landmarks can capture the essential features of the face and reduce recognition errors caused by external interference. As an embodiment of the present application, facial landmarks are extracted from the face area, and the coordinates of the facial landmarks can be obtained by respectively calculating the gradient values ​​​​I of the face area in the X-axis direction and the Y-axis direction of the two-dimensional coordinate system x and I y ; Use Gaussian function to and I x *I y Filter and obtain matrices A, B and C respectively, where the Gaussian function is recorded as σ is the standard deviation, x and y are the coordinates of the pixel points, And C = G(x,y)*I x *I y; According to matrices A, B and C, calculate the feature point response function R(x,y)=Det(M)-k*(trace(M)) 2 The value R, where Det(M) is the determinant of M, trace(M)=A+B is the trace of M, and k is an empirical parameter; the value R is non-maximum suppressed, the facial feature points in the face area are extracted and their coordinates are recorded.

[0040] Step S1025: According to the coordinates of the facial feature points, the head posture of the face corresponding to the face area is adjusted to a standard head posture of the ID photo that matches the target area.

[0041] To adjust the head posture of the face corresponding to the face area in the second standard ID photo to the standard head posture of the ID photo that meets the target area, the head posture of the face corresponding to the face area must first be calculated. In order to calculate the head posture, the facial feature points of these 2D face areas can be mapped to a standard 3D facial model. This standard 3D facial model usually contains about 60 different facial feature points (such as a simple 3D head mesh), and the 3D coordinates of these points are matched with the facial feature points in the 2D face area. Specifically, the 3D coordinates are converted to 2D coordinates using the known camera intrinsic parameters (such as focal length, principal point position, etc.). This process is usually completed through the following projection matrix:

[0042] P 2D =K*R*P 3D +t

[0043] Among them, P 2D is the 2D coordinate of the facial feature point in the face area, P 3D is the 3D key point corresponding to the facial model, R is the 3D rotation matrix, which indicates the rotation of the face in 3D space, K is the intrinsic parameter matrix of the camera (including focal length, optical center, etc.), and t is the translation vector, which indicates the displacement between the camera and the face. Through optimization algorithms (such as the least squares method), the 3D coordinates and 3D rotation matrix can be inferred from the 2D coordinates in the face area and the standard 3D model.

[0044] After obtaining the 3D model and its mapping with the 2D image, the head posture can be calculated, including the pitch angle θ pitch (Pitch, that is, the angle of rotation along the x-axis, indicating the deviation of the upper and lower head (such as lowering or raising the head)), yaw angle θ yaw (Yaw, the angle of rotation along the y-axis, represents the deviation of the left and right head (for example, turning the head left and right)) and the roll angle θ roll (Roll, that is, the angle of rotation along the z-axis, indicating the angle of the head tilting left and right), where θ pitchis the angle associated with column 1 in the rotation matrix and can be found by:

[0045] θ pitch = atan2(R 32 ,R 33 ), where R 32 and R 33 is an element in the rotation matrix, representing the rotation around the x-axis.

[0046] θ yaw is the angle associated with column 2 in the rotation matrix and can be found by:

[0047] Among them, R 31 is an element in the rotation matrix, representing the rotation around the y-axis.

[0048] θ roll is the angle associated with column 3 in the rotation matrix and can be found by:

[0049] θ roll =atan2(R 21 ,R 11 ), R 21 and R 11 is an element in the rotation matrix, representing the rotation around the z-axis.

[0050] The above θ pitch ,θ yaw and θ roll In the calculation formula, atan2 is a commonly used function used to calculate angles on a two-dimensional plane, especially when coordinate systems are involved. It is an extension of the inverse tangent function (atan) that can handle situations in different quadrants. atan2(y,x) is a function that can return an angle, where y and x are two components in a two-dimensional coordinate system (i.e., the coordinates of a vector). Its function is to calculate the angle between a point and the origin based on the given y and x coordinates, and take into account the quadrant of the angle. Specifically, atan2(y,x) calculates the angle between the point (y, x) and the x-axis. atan2(y,x) returns a radian value (usually in the range of -π to π), which can be converted to an angle value (×180 / π). With θ pitch = atan2(R 32 ,R 33 ) as an example, indicating the use of the element R of the rotation matrix 32 and R 33 To calculate the pitch angle (Pitch), where R 32 is an element in the rotation matrix, representing the projection along the z-axis; R 33Another element in the rotation matrix represents the projection along the x-axis. pitch =atan2(R 32 ,R 33 ), we can get an angle, which represents the pitch angle of the object relative to the horizontal plane.

[0051] Step S103: adjusting the background of the first standard ID photo according to the ID photo standard parameters of the target area, and obtaining a second standard ID photo of the to-be-processed photo corresponding to the target area.

[0052] Although the size, head proportion and head posture are adjusted to adjust the photo to be processed to the first standard ID photo corresponding to the target area, on the one hand, considering that if the background color of the first standard ID photo is complex, or the background has gradients, textures or irregular shapes (such as windows, walls, etc.), the traditional method of background processing (for example, directly replacing the background color) may not be able to effectively segment the characters, and the background may appear very stiff and lack natural transition after replacement; on the other hand, although the background color generated by the traditional method meets the standard, it does not have detailed textures, blurring, gradient effects, etc., which makes the background details not rich enough and easily produces effects that are inconsistent with the actual scene. Therefore, the present application can consider using a generative adversarial network to process the first standard ID photo. Specifically, the background of the first standard ID photo is adjusted according to the standard parameters of the ID photo of the target area, and the second standard ID photo corresponding to the target area of ​​the photo to be processed can be obtained through steps S1031 to S1033, as described in detail as follows:

[0053] Step S1031: Use the background color, background texture, and background blur of the ID photo of the target area to train the generative adversarial network to obtain a trained generative adversarial network.

[0054] Here, the background color, background texture and background blur of the ID photo in the target area are the standard parameters of the ID photo in the target area. The generative adversarial network mainly consists of two parts, namely, the generator and the discriminator. The task of the generator is to generate a background image that meets the standard requirements based on random noise or background information provided by the user (such as color, texture, etc.). The task of the discriminator is to determine whether an image is from a real sample or a pseudo image generated by the generator. The goal of training the generative adversarial network is to improve the generator's generation ability through the feedback of the discriminator, so that the generated background is closer and closer to the background of the ID photo in the target area.

[0055] Step S1032: Generate a background image of the ID photo that conforms to the target area through the objective function set by the trained generative adversarial network.

[0056] In the field of image processing, the Generative Adversarial Network (GAN) consists of two parts: a generator and a discriminator. The task of the generator is to generate a background image that meets the standard requirements based on random noise or background information provided by the user (such as color, texture, etc.). Assuming that the input of the generator is a random noise vector z and its output is a background image G(z), the task of the discriminator is to determine whether an image comes from a real sample or a pseudo image generated by the generator. The input of the discriminator is an image x, and the output is a probability value D(x), indicating whether the image x is a "real" background (that is, the background of the ID photo that meets the target area). The training goal of GAN is to optimize the performance of the generator and the discriminator through adversarial training, that is, to make the generator generate background images that are as realistic as possible, and the discriminator correctly judge whether the background image generated by the generator is a real sample. The objective function of GAN includes the loss function of the generator and the loss function of the discriminator, where the loss function L of the generator G as follows:

[0057] L G =-E z~pz(z) [logD(G(z))]

[0058] Among them, G(z) is the background image output by the generator, D(G(z)) is the predicted value of the background image output by the discriminator, z~p z (z) is the input noise distribution of the generator, E is the expected symbol, which means that for the distribution z~p z The noise z sampled in (z) is averaged over the loss of the generator output.

[0059] The loss function L of the discriminator D as follows:

[0060] L D =-E x~pdata(x) [logD(x)]-E z~pz(z) [1-logD(G(z))]

[0061] Among them, x is the background image from the real data set (standard background sample), that is, the background image of the ID photo of the target area, and E is the expected symbol.

[0062] During the training process of the generative adversarial network, the generator and the discriminator are optimized alternately. The generator deceives the discriminator by generating increasingly realistic background images, while the discriminator challenges the generator by improving its ability to distinguish. The ultimate goal of the training is to minimize the difference between the background image generated by the generator and the background image of the ID photo in the target area, at which point a trained generative adversarial network is obtained.

[0063] Step S1033: synthesize the facial image of the face corresponding to the face area of ​​the first standard ID photo and the background image of the ID photo that matches the target area to obtain a second standard ID photo corresponding to the target area of ​​the photo to be processed.

[0064] On the one hand, during the image synthesis process, especially when the foreground (user facial image) and the background are synthesized, the edges of the characters may have a jagged visual effect, which is caused by the uneven edges or hard transitions during the image segmentation and synthesis process; on the other hand, during the image synthesis process, especially when the foreground (user facial image) is synthesized with a complex background (such as a background with gradients or textures), if the brightness and contrast of the two are too different, the synthesis result will be unnatural. Therefore, in the embodiment of the present application, before synthesis, an anti-aliasing algorithm can be used to smooth the edges of the face area of ​​the first standard ID photo; and the facial image and the background image of the ID photo that meets the target area can be matched by brightness mapping and exposure adjustment. In the above embodiment, the facial image and the background image of the ID photo that meets the target area are matched in illumination by brightness mapping and exposure adjustment. The core lies in adjusting the brightness and contrast of the image so that the transition between the facial image and the background image of the ID photo that meets the target area is smoother. Specifically, the brightness mapping and exposure adjustment process of the facial image are described first. The purpose of brightness mapping is to balance the difference in illumination intensity between the facial image and the background image so that the two images appear to be images under the same illumination conditions after synthesis. It mainly includes steps S1 to S3, which are described in detail as follows:

[0065] Step S1: Extract the brightness information of the facial image and the background image respectively:

[0066] Brightness information is usually extracted from the color information of the image. In the RGB color space, brightness can be calculated by the following formula:

[0067] L=0.299*R+0.587*G+0.114*B

[0068] Among them, L represents the brightness value of the pixel, and R, G, and B are the values ​​of the red, green, and blue channels of the image pixel, respectively.

[0069] Step S2: Calculate the brightness difference between the facial image and the background image.

[0070] In the embodiment of the present application, the brightness difference between the facial image and the background image can be calculated by calculating the statistical characteristics of the brightness of the facial image and the background image. For example, the mean brightness difference ΔL between the facial image and the background image is calculated by the following formula:

[0071]

[0072] Among them, L fg (i) represents the brightness value of the i-th pixel in the facial image, L bg (j) represents the brightness value of the jth pixel in the background image, N and M represent the number of pixels in the facial image and the background image, respectively, and ΔL represents the brightness difference, which measures the inconsistency in brightness between the facial image and the background image.

[0073] Step S3: Adjust the brightness of the facial image.

[0074] According to the calculated brightness difference ΔL between the facial image and the background image, the brightness of the facial image is adjusted. Specifically, the pixel brightness value of the facial image is adjusted so that the overall brightness of the facial image is close to the brightness of the background image:

[0075] L' fg (i) = L fg (i)+ΔL

[0076] Among them, L' fg (i) is the brightness value of the adjusted facial image, ΔL is the brightness difference, and the brightness of the facial image is adjusted to match that of the background image.

[0077] The following describes exposure adjustment. Exposure adjustment is usually used to change the overall brightness or contrast of an image so that the illumination of the facial image and the background image is more consistent. The main method of exposure adjustment is to control the overall brightness by linearly transforming the image pixel values, including steps S'1 to S'3, which are described in detail as follows:

[0078] Step S'1: Calculate the exposure factor.

[0079] The exposure factor is usually expressed as a scalar that represents the change in the exposure of the image. The exposure factor can be calculated by calculating the overall brightness difference between the face image and the background image:

[0080]

[0081] Among them, E factor Represents the exposure factor, which is measured by the ratio of the background image to the facial image. bg (i) represents the brightness of the i-th pixel in the background image, L fg (i) represents the brightness of the i-th pixel in the face image.

[0082] Step S'2: Apply exposure adjustment.

[0083] According to the calculated exposure factor E factor , adjust the brightness of the facial image. Usually, exposure adjustment is to adjust the pixel brightness of the facial image by multiplying the exposure factor:

[0084] L'fg (i) = L fg (i)*E factor

[0085] Among them, L' fg (i) is the adjusted facial image brightness value, E factor It is the exposure factor, which determines the adjustment of the brightness of the facial image.

[0086] Step S'3: Contrast adjustment.

[0087] If you need to adjust the contrast of the facial image, you can use the following formula to nonlinearly adjust the pixel value:

[0088] I' fg =α*(I fg -128)+128

[0089] Among them, I' fg is the adjusted facial image pixel value, I fg is the original facial image pixel value, and α is the contrast adjustment factor, which determines the change in image contrast. If α>1, the facial image contrast is enhanced; if 0<α<1, the facial image contrast is reduced.

[0090] The above embodiment uses an anti-aliasing algorithm to evenly smooth the image and gradually transition the edges, making the joints between the characters and the background more natural and reducing the jagged effect. Through brightness mapping and exposure adjustment, the lighting match between the foreground and background images can be effectively achieved, reducing the visual inconsistency.

[0091] As an embodiment of the present application, synthesizing the facial image of the face corresponding to the face area of ​​the first standard ID photo and the background image of the ID photo that matches the target area to obtain the second standard ID photo corresponding to the target area of ​​the photo to be processed can be achieved by steps S"1 to S"4, as described in detail as follows:

[0092] Step S"1: by estimating the optical flow field between the facial image of the face corresponding to the face area and the background image of the ID photo of the target area, the facial image is transformed to obtain a transformed facial image.

[0093] Specifically, the optical flow algorithm can be used to calculate the facial image I f and background image I b The optical flow field between them represents the movement of image pixels.

[0094] F flow =OpticalFlow(I f ,I b )

[0095] Among them, Fflow is the optical flow field, containing the motion vector for each pixel.

[0096] Then, the optical flow field F flow Application to facial images I f , get the position of the facial image in the background image.

[0097]

[0098] Among them, Warp is an operation that uses the optical flow field to transform the facial image. is the transformed facial image.

[0099] Step S"2: construct a Laplacian pyramid of the background image and the transformed facial image according to the background image and the transformed facial image.

[0100] Specifically, it can be the transformed facial image and background image I b Construct a Gaussian pyramid, where the Gaussian pyramid is a series of images obtained by multiple Gaussian blurring and downsampling of the image:

[0101]

[0102] Among them, l represents the number of layers of the Gaussian pyramid, is the transformed facial image In the Gaussian pyramid representation at level l, is the background image I b Gaussian pyramid representation at level l.

[0103] Then, a Laplacian pyramid is constructed from the Gaussian pyramid. The Laplacian pyramid is obtained by subtracting adjacent layers of the Gaussian pyramid, that is:

[0104]

[0105] Among them, Expand means to perform interpolation and enlargement operation on the facial image or background image. is the facial image I f The Laplacian pyramid representation at the lth level is the difference image of the Gaussian pyramid. is the background image I b The Laplacian pyramid representation at the lth level is the difference image of the Gaussian pyramid.

[0106] Step S'3: Blend the Laplacian pyramid of the background image and the transformed facial image.

[0107] Specifically, the Laplacian pyramids of the facial image and the background image are mixed, and a weight mask W can be used to control the mixing ratio of the facial image and the background image.

[0108]

[0109] Among them, the weight mask W controls the mixing ratio of the facial image and the background image, which is usually set according to the image content or manually.

[0110] Step S"4: reconstructing the image from the mixed Laplacian pyramid to obtain a second standard ID photo corresponding to the target area of ​​the photo to be processed.

[0111] Specifically, the synthetic image is reconstructed from the mixed Laplacian pyramid, and its implementation can be expressed as:

[0112]

[0113] The foreground and background synthesis scheme based on the mixture of optical flow and Laplacian pyramid in the above-mentioned embodiment estimates the movement between the foreground and background through the optical flow field, and uses the Laplacian pyramid to achieve smooth transition of the image, which can achieve natural image synthesis without complex training.

[0114] Step S104: Correcting the facial expression of the face corresponding to the face area of ​​the second standard ID photo.

[0115] Taking into account that the facial expression of the face corresponding to the face area of ​​the second standard ID photo may not meet the standard expression of the ID photo, such as smiling or opening the mouth, etc., it is necessary to correct it to meet the required "neutral" expression (the "neutral" expression is an expression that meets the standard of the ID photo). As an embodiment of the present application, correcting the facial expression of the face corresponding to the face area of ​​the second standard ID photo can be: defining an objective function, which is used to measure the difference between the facial expression of the face corresponding to the face area of ​​the second standard ID photo and the neutral expression; iteratively updating the facial expression of the face corresponding to the face area of ​​the second standard ID photo until the value of the objective function is minimized, and determining that the correction of the corresponding facial expression is completed. In the above embodiment, iteratively updating the facial expression of the face corresponding to the face area of ​​the second standard ID photo until the value of the objective function is minimized, and determining that the correction of the corresponding facial expression is completed can be specifically implemented through steps S1041 and S1042, as described in detail as follows:

[0116] Step S1041: Express the objective function as a weighted sum of content loss and style loss.

[0117] In the embodiment of the present application, the content loss is used to measure the content difference between the facial expression of the face corresponding to the face area of ​​the second standard ID photo and the neutral expression, and the style loss is used to measure the style difference between the facial expression of the face corresponding to the face area of ​​the second standard ID photo and the neutral expression. content (C,G) represents content loss, then:

[0118]

[0119] in, and They are respectively the activation values ​​of the facial expression and the neutral expression of the face corresponding to the face area of ​​the second standard ID photo at a certain layer.

[0120] If you use L style (S,G) represents the style loss, then:

[0121]

[0122] in, and are the Gram matrix features of the facial expression and neutral expression of the face area corresponding to the second standard ID photo, and N and M are the height and width of the image.

[0123] The objective function L can be expressed as:

[0124] L=α*L content (C,G)+β*L style (S,G)

[0125] Among them, α and β are weighted coefficients of content loss and style loss, which can be determined according to the weights of content and style in the image, C is the content feature of the facial expression of the face area corresponding to the second standard ID photo, S is the style feature of the neutral expression image, and G is the generated image.

[0126] Step S1042: Use the gradient descent method to optimize the objective function until the weighted sum is minimized, thereby determining that the correction of the corresponding facial expression is completed.

[0127] The objective function L is optimized by the gradient descent method, that is, the generated image G is iteratively updated until the objective function L is minimized, as follows:

[0128]

[0129] Among them, G (t) is the value of the updated generated image G at the tth iteration, is the gradient of the objective function L with respect to the generated image G, and λ is the learning rate, which is used to control the step size of each iteration.

[0130] Step S105: Output the image after facial expression correction as the final standard ID photo.

[0131] From the above attached Figure 1 It can be seen from the example method for generating a standard ID photo of any area that, on the one hand, compared with the challenge of the prior art requiring the user to input the standard parameters of the ID photo of the target area, the technical solution of the present application only requires the user to input the target area to obtain the standard parameters of the ID photo of the target area, which greatly reduces the professional requirements for the user and greatly facilitates the user; on the other hand, based on the standard parameters of the ID photo of the target area, the first standard ID photo, the second standard ID photo and even the final standard ID photo of the target area corresponding to the photo to be processed can be obtained, which can be automatically completed based on a computer program (for example, an artificial intelligence large model) without manual participation or intervention, which not only improves the efficiency of generating standard ID photos of any area, but also minimizes errors or mistakes, so that the accuracy of generating standard ID photos of any area is guaranteed.

[0132] Please see attached Figure 2 , is a device for generating a standard ID photo of any area provided in an embodiment of the present application, the device may include an acquisition module 201, a first adjustment module 202, a second adjustment module 203, a correction module 204 and an output module 205, which are described in detail as follows:

[0133] The acquisition module 201 is used to acquire the ID photo standard parameters of the target area according to the target area input by the user;

[0134] The first adjustment module 202 is used to adjust the size, head proportion and head posture of the photo to be processed according to the standard parameters of the ID photo of the target area, so as to obtain the first standard ID photo of the photo to be processed corresponding to the target area;

[0135] The second adjustment module 203 is used to adjust the background of the first standard ID photo according to the ID photo standard parameters of the target area to obtain a second standard ID photo corresponding to the target area of ​​the photo to be processed;

[0136] Correction module 204, used to correct the facial expression of the face corresponding to the face area of ​​the second standard ID photo;

[0137] The output module 205 is used to output the image after facial expression correction as the final standard ID photo.

[0138] From the above attached Figure 2It can be seen from the example of the device for generating a standard ID photo of any area that, on the one hand, compared with the challenge of the prior art requiring the user to input the standard parameters of the ID photo of the target area, the technical solution of the present application only requires the user to input the target area to obtain the standard parameters of the ID photo of the target area, which greatly reduces the professional requirements for the user and greatly facilitates the user; on the other hand, according to the standard parameters of the ID photo of the target area, the first standard ID photo, the second standard ID photo and even the final standard ID photo of the target area corresponding to the photo to be processed can be obtained, which can be automatically completed based on a computer program (for example, an artificial intelligence large model) without manual participation or intervention, which not only improves the efficiency of generating standard ID photos of any area, but also minimizes errors or mistakes, so that the accuracy of generating standard ID photos of any area is guaranteed.

[0139] Figure 3 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Figure 3 As shown, the electronic device 3 of this embodiment mainly includes: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30, such as a program for a method for generating a standard ID photo of any region. When the processor 30 executes the computer program 32, the steps in the above-mentioned method for generating a standard ID photo of any region are implemented, such as Figure 1 Alternatively, when the processor 30 executes the computer program 32, the functions of each module / unit in the above-mentioned device embodiments are implemented, for example Figure 2 The functions of the acquisition module 201 are shown.

[0140] Exemplarily, the computer program 32 of the method for generating a standard ID photo of any area mainly includes: obtaining the standard parameters of the ID photo of the target area according to the target area input by the user; adjusting the size, head proportion and head posture of the photo to be processed according to the standard parameters of the ID photo of the target area, and obtaining the first standard ID photo of the photo to be processed corresponding to the target area; adjusting the background of the first standard ID photo according to the standard parameters of the ID photo of the target area, and obtaining the second standard ID photo of the photo to be processed corresponding to the target area; correcting the facial expression of the face corresponding to the face area of ​​the second standard ID photo; outputting the image after facial expression correction as the final standard ID photo. The computer program 32 can be divided into one or more modules / units, one or more modules / units are stored in the memory 31, and executed by the processor 30 to complete the present application. One or more modules / units can be a series of computer program instruction segments that can complete specific functions, and the instruction segments are used to describe the execution process of the computer program 32 in the electronic device 3. For example, the computer program 32 can be divided into the functions of the acquisition module 201 and the first (module in the virtual device), and the specific functions of each module are as follows: the acquisition module 201 is used to obtain the standard parameters of the ID photo of the target area according to the target area input by the user; the first adjustment module 202 is used to adjust the size, head proportion and head posture of the photo to be processed according to the standard parameters of the ID photo of the target area, and obtain the first standard ID photo of the photo to be processed corresponding to the target area; the second adjustment module 203 is used to adjust the background of the first standard ID photo according to the standard parameters of the ID photo of the target area, and obtain the second standard ID photo of the photo to be processed corresponding to the target area; the correction module 204 is used to correct the facial expression of the face corresponding to the face area of ​​the second standard ID photo; the output module 205 is used to output the image after facial expression correction as the final standard ID photo.

[0141] The electronic device 3 may include but is not limited to a processor 30 and a memory 31. Those skilled in the art will appreciate that Figure 3 It is only an example of the electronic device 3 and does not constitute a limitation of the electronic device 3. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.

[0142] The processor 30 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0143] The memory 31 may be an internal storage unit of the electronic device 3, such as a hard disk or memory of the electronic device 3. The memory 31 may also be an external storage device of the electronic device 3, such as a plug-in hard disk, a smart memory card (SmartMedia Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device 3. Further, the memory 31 may also include both an internal storage unit of the electronic device 3 and an external storage device. The memory 31 is used to store computer programs and other programs and data required by the electronic device. The memory 31 may also be used to temporarily store data that has been output or is to be output.

[0144] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0145] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0146] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0147] In the embodiments provided in the present application, it should be understood that the disclosed devices / equipment and methods can be implemented in other ways. For example, the device / equipment embodiments described above are only schematic, for example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0148] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0149] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0150] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program of the method for generating a standard ID photo of any area can be stored in a storage medium. When the computer program is executed by the processor, the steps of each of the above-mentioned method embodiments can be implemented, that is, according to the target area input by the user, the standard parameters of the ID photo of the target area are obtained; according to the standard parameters of the ID photo of the target area, the size, head proportion and head posture of the photo to be processed are adjusted to obtain the first standard ID photo of the photo to be processed corresponding to the target area; according to the standard parameters of the ID photo of the target area, the background of the first standard ID photo is adjusted to obtain the second standard ID photo of the photo to be processed corresponding to the target area; the facial expression of the face corresponding to the face area of ​​the second standard ID photo is corrected; and the image after facial expression correction is output as the final standard ID photo. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. Storage media may include: any entity or device capable of carrying computer program code, recording media, USB flash drives, mobile hard disks, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content contained in the storage medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, storage media do not include electric carrier signals and telecommunication signals.

[0151] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application is described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some of the technical features therein by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application. The specific implementation methods described above further describe the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above description is only the specific implementation method of the present application, and is not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the protection scope of the present invention.

Claims

1. A method for generating a standard ID photo of any area, characterized in that: The method comprises: According to the target area input by the user, obtain the standard parameters of the ID photo of the target area; Adjust the size, head proportion and head posture of the photo to be processed according to the standard parameters of the ID photo of the target area, so as to obtain a first standard ID photo of the photo to be processed corresponding to the target area; Adjusting the background of the first standard ID photo according to the standard ID photo parameters of the target area to obtain a second standard ID photo of the to-be-processed photo corresponding to the target area; Correcting the facial expression of the face corresponding to the face area of ​​the second standard ID photo; Output the image after facial expression correction as the final standard ID photo.

2. The method for generating a standard ID photo of any area as claimed in claim 1, characterized in that: The standard parameters of the ID photo of the target area include the background color, background texture and background blur of the ID photo of the target area, and adjusting the background of the first standard ID photo according to the standard parameters of the ID photo of the target area to obtain the second standard ID photo of the photo to be processed corresponding to the target area includes: Training a generative adversarial network using the background color, background texture, and background blur of the ID photo of the target area to obtain a trained generative adversarial network; Generate a background image of the ID photo that conforms to the target area through the objective function set by the trained generative adversarial network; The facial image of the face corresponding to the face area of ​​the first standard ID photo and the background image are synthesized to obtain a second standard ID photo corresponding to the target area of ​​the photo to be processed.

3. The method for generating a standard ID photo of any area as claimed in claim 2, characterized in that: The synthesizing the facial image of the face corresponding to the face area of ​​the first standard ID photo and the background image to obtain a second standard ID photo corresponding to the target area of ​​the photo to be processed includes: By estimating the optical flow field between the facial image and the background image, a preset transformation operation is performed on the facial image to obtain a transformed facial image; Constructing a Laplacian pyramid of the background image and the transformed facial image according to the background image and the transformed facial image; Mixing the background image and the Laplacian pyramid of the transformed facial image; The image is reconstructed from the mixed Laplacian pyramid to obtain a second standard ID photo corresponding to the target area of ​​the photo to be processed.

4. The method for generating a standard ID photo of any area as claimed in claim 2, characterized in that: Before synthesizing the facial image of the face corresponding to the face area of ​​the first standard ID photo and the background image, the method further includes: Using an anti-aliasing algorithm to smooth the edge of the face area; and The illumination of the facial image and the background image is matched by adjusting brightness and exposure.

5. The method for generating a standard ID photo of any area as claimed in claim 1, characterized in that: The adjusting the size, head proportion and head posture of the photo to be processed according to the standard parameters of the ID photo of the target area to obtain the first standard ID photo of the photo to be processed corresponding to the target area includes: According to the height and width of the ID photo of the target area, the size of the photo to be processed is adjusted to the height and width of the ID photo of the target area; According to the head proportion of the ID photo in the target area, the height of the head in the to-be-processed photo with the adjusted height is adjusted to the height of the head in the ID photo in the target area; Positioning the face area of ​​the photo to be processed; Extracting facial feature points from the facial area to obtain coordinates of the facial feature points; According to the coordinates of the facial feature points, the head posture of the face corresponding to the face area is adjusted to a standard head posture that conforms to the ID photo of the target area.

6. The method for generating a standard ID photo of any area as claimed in claim 1, characterized in that: Correcting the facial expression of the face corresponding to the face area of ​​the second standard ID photo includes: defining an objective function, the objective function being used to measure the difference between a facial expression of a face corresponding to the face region of the second standard ID photo and a neutral expression; Iteratively update the facial expression of the face corresponding to the face area of ​​the second standard ID photo until the value of the objective function is minimized, and determine that the correction of the corresponding facial expression is completed.

7. The method for generating a standard ID photo of any area as claimed in claim 6, characterized in that: The iterative updating of the facial expression of the face corresponding to the face area of ​​the second standard ID photo until the value of the objective function is minimized includes: The objective function is expressed as a weighted sum of a content loss and a style loss, wherein the content loss is used to measure a content difference between a facial expression of a face corresponding to the face region of the second standard ID photo and a neutral expression, and the style loss is used to measure a style difference between a facial expression of a face corresponding to the face region of the second standard ID photo and a neutral expression; The objective function is optimized using a gradient descent method until the weighted sum is minimized.

8. A device for generating a standard ID photo of any area, characterized in that: The device comprises: An acquisition module, used to acquire the ID photo standard parameters of the target area according to the target area input by the user; A first adjustment module is used to adjust the size, head proportion and head posture of the photo to be processed according to the standard parameters of the ID photo of the target area, so as to obtain a first standard ID photo of the photo to be processed corresponding to the target area; A second adjustment module, configured to adjust the background of the first standard ID photo according to the ID photo standard parameters of the target area, so as to obtain a second standard ID photo of the to-be-processed photo corresponding to the target area; A correction module, used to correct the facial expression of the face corresponding to the face area of ​​the second standard ID photo; The output module is used to output the image after facial expression correction as the final standard ID photo.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.