Method for automatically processing portrait photo into standard certificate photo

Through intelligent algorithms and image processing technology, the figures of people in portrait photos are automatically adjusted to make the eyes flush with the shoulders, solving the problem of insufficient automatic adjustment and processing accuracy in the existing technology, and achieving efficient and automated document photo generation.

CN120147111AActive Publication Date: 2025-06-13BEIJING JINSHAJIANG TECH CO LTD

Patent Information

Application Number
CN202510224890.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

The prior art is difficult to automatically adjust the figure of a person, so that the eyes are flush with the shoulders, and the processing accuracy and efficiency are insufficient, so it cannot meet the requirements of standard document photos.

Method used

The image processing is performed using intelligent algorithms, and the face key point detection model and the cutout model are automatically rotated and cropped to align the eyes, and the shoulder key point detection model and local deformation algorithm are used to adjust the shoulder height to align it with the reference value.

Benefits of technology

It realizes the generation of document photos with higher degree of automation and higher accuracy, automatically adjusts the figures, keeps eyes flush with shoulders, improves processing efficiency and reduces manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147111A_ABST
    Figure CN120147111A_ABST
Patent Text Reader

Abstract

The invention discloses a method for automatically processing a portrait photo into a standard certificate photo, belongs to the technical field of computer vision and image processing, and particularly relates to a technology for automatically processing the portrait photo into the certificate photo and a technology capable of automatically adjusting the posture of a figure and enabling the eyes of the figure to be flush with the shoulders. The technology relates to the fields of face detection, key point positioning, image rotation, background synthesis, image segmentation, shoulder key point detection and position adjustment and the like, and is specifically applied to automatic generation of certificate photos. The method for automatically processing the portrait photo into the standard certificate photo is different from traditional manual processing, the portrait photo can be automatically processed into the certificate photo meeting the standard requirement, the figure posture can be automatically adjusted, the eyes of the figure are flush with the shoulders, the complexity of traditional manual processing is avoided, the processing efficiency is improved, and the user experience is improved. Labor is saved, and the processing precision is higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and image processing, and particularly to a method for processing portrait photos into standard ID photos. Background Art

[0002] With the rapid development of information technology, especially the continuous progress of computer vision and artificial intelligence technologies, the automatic generation of ID photos has gradually become a demand. Traditional ID photo shooting usually requires manual operation, demanding that the person maintain a strict pose and standard during shooting. However, in practical applications, users may fail to achieve a perfect pose for various reasons, resulting in the photos taken not meeting the requirements of standard ID photos, such as the misalignment of the eyes and shoulders or non-compliant backgrounds. Therefore, the automatic generation of ID photos that meet the standards has important practical value.

[0003] Currently, some image processing technologies have attempted to generate ID photos through face detection and pose correction, but most of these methods rely on manual adjustment or can only correct photos through static and simple rotation. Moreover, most technologies have not been able to achieve automatic adjustment of the alignment of the shoulders and eyes or cannot fully cope with the differences in the body postures of different users. There is still much room for improvement in the processing accuracy, efficiency, and breadth of application scenarios of the existing technologies.

[0004] Therefore, how to develop an image processing technology with higher automation and better accuracy, which can process portrait photos into ID photos that meet the standard requirements and can automatically adjust the body posture of the person to make the person's eyes level with the shoulders, has become a technical problem to be solved urgently. Based on this demand, the present invention proposes a technical solution for image processing through intelligent algorithms to automatically correct the person's pose and generate ID photos that meet the standard requirements. Summary of the Invention

[0005] To alleviate or solve at least one aspect or at least one point of the above problems, the present invention is proposed. The present invention provides a method for automatically processing portrait photos into standard ID photos, including the following steps:

[0006] Step S1: Input a portrait photo and use a pre-trained face key point detection model to perform face key point detection to locate the coordinates of the two eyes.

[0007] Step S2: Calculate the required rotation direction and angle based on the coordinates of the two eyes to ensure the alignment of the two eyes in the horizontal direction; Step S3: Use a matte extraction model to perform precise matte extraction to generate a portrait image with a transparent background, and rotate the portrait image with a transparent background to obtain a portrait image with a transparent background after the eyes are aligned.

[0008] Step S4: cropping the transparent background portrait image according to the rotation result, and synthesizing the transparent background with the standard background to process it into a standard ID photo;

[0009] Step S5: using a pre-trained human shoulder key point detection model to detect key points of the left and right shoulders of the standard ID photo;

[0010] Step S6: Calculate the difference between the average height of the shoulders and the preset reference height according to the coordinates of the key points of the left and right shoulders, and determine the shoulder that needs to be adjusted;

[0011] Step S7: Using a local deformation algorithm, adjust the shoulder key points that need to be adjusted so that their heights are aligned with the reference values, and obtain a result image in which the eyes and shoulders of the original image are flush.

[0012] Preferably, the facial key point detection model is a model in computer vision, which can identify the structural features of the face in the image and accurately mark a plurality of key point features, wherein the key point features include eyes, nose, ears, corners of the mouth, jaw, shoulder and neck areas below the ears, and key shoulder areas;

[0013] Preferably, the face landmark detection model uses hierarchical feature extraction and feature pyramid structure to ensure detection of objects of various sizes, and refines these predictions using non-maximum suppression to filter out duplicate or low-confidence boxes, thereby achieving more accurate object detection.

[0014] Preferably, the facial key point detection model adopts an improved backbone network and neck architecture, the backbone network adopts a faster and more efficient variant of the cross-stage partial bottleneck structure, and the improved backbone network adds a layer of cross-stage local spatial attention module after the spatial pyramid pooling fast module.

[0015] Preferably, the facial key point detection comprises the following steps:

[0016] Collect and preprocess portrait data;

[0017] The face landmark detection model passes the input image into a convolutional neural network to extract features to perform object detection;

[0018] Self-training of facial key point detection model;

[0019] The facial key point detection model automatically identifies faces in images and annotates their key points.

[0020] Preferably, the step of calculating the required rotation direction and angle according to the coordinates of the eyes comprises the following steps:

[0021] Analyze the coordinates of both eyes in the input portrait photo, calculate the horizontal axis distance and the rotation angle;

[0022] According to the relative relationship between the distance difference of the two eyes on the horizontal axis and the image width, use this difference to judge the horizontal offset degree of the two eyes, and further calculate the rotation angle.

[0023] Preferably, the matting model is used to accurately separate the foreground from the background and generate a high-quality portrait with a transparent background. The matting model includes a shared encoder module, a pyramid pooling module, a semantic context branch module, and a high-resolution detail branch module. The backbone network of the matting model uses a high-resolution network, which is also used as a shared encoder. The output of the last stage of the backbone network is sent to the pyramid pooling module to obtain richer semantic context information.

[0024] Preferably, the semantic context branch module consists of five blocks. Each block includes convolution, batch normalization, an activation function, and a bilinear upsampling module. The output of the semantic context branch is also used for the supervision of the semantic segmentation task, which includes three categories: foreground, background, and transition region.

[0025] Preferably, the input of the high-resolution detail branch is formed by splicing the intermediate features of the shared encoder after an upsampling operation.

[0026] Preferably, the intermediate features extracted by the shared encoder at different levels are first restored to the same resolution through an upsampling operation, and then these feature maps are spliced in the channel dimension to form the initial input of the high-resolution detail branch.

[0027] Preferably, the semantic context feature map of the semantic context branch module and the detail feature map of the high-resolution detail branch module are subjected to a guidance flow process. The guidance flow process is to splice the semantic context feature map from the semantic context branch and the detail feature map of the high-resolution detail branch, and then send it into a module containing convolution, batch normalization, a rectified linear unit activation function, convolution, batch normalization, and an S-shaped activation function to generate a guidance map. Then, the guidance map is subjected to a dot product and addition operation with the detail feature map of the high-resolution detail branch to generate the feature map of the next stage of the high-resolution detail.

[0028] The present invention provides a method for automatically processing a portrait photo into a standard ID photo, an image processing technology with higher automation and better precision, which can automatically process a portrait photo into an ID photo that meets the standard requirements, and can automatically adjust the body posture of the person to make the person's eyes level with the shoulders, eliminating the cumbersome traditional manual processing, improving the processing efficiency, saving labor, and having higher processing precision. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is a flow chart for processing a portrait photo into a standard ID photo.

[0030] Figure 2 It is the backbone network structure diagram of the face key point detection model.

[0031] Figure 3 It is the structure diagram of the spatial pyramid pooling fast module.

[0032] Figure 4 It is the structure diagram of the cross-stage local spatial attention module.

[0033] Figure 5 It is the network architecture of the matting model.

[0034] Figure 6 It is the guiding flow structure. Specific implementation manners

[0035] The following description of the embodiments of the present invention with reference to the accompanying drawings is intended to explain the overall inventive concept of the present invention and should not be construed as a limitation of the present invention. In the present invention, the same reference numerals represent the same or similar components.

[0036] The features described herein can be implemented in different forms and should not be construed as limited to the examples described herein. On the contrary, the examples provided herein are only to illustrate some of the many possible ways of implementing the methods, devices, and / or systems described herein, and many possible ways will be apparent after understanding the disclosure of the present invention.

[0037] The terms used herein are only for describing various examples and will not be used to limit the disclosure. Unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. The terms "comprising", "including", and "having" indicate the presence of the recited features, numbers, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, numbers, operations, components, elements, and / or combinations thereof.

[0038] In order to enable those skilled in the art to use the content of the present invention, the following exemplary embodiments may be given in combination with specific application scenarios, specific system, device, and component parameters, and specific connection manners. However, for those skilled in the art, these embodiments are only examples, and without departing from the spirit and scope of the present invention, the general principles defined here can be applied to other embodiments and application scenarios.

[0039] According to an exemplary embodiment of the present invention: This example provides a method for automatically processing a portrait photo into a standard ID photo, as Figure 1 shown, including the following steps:

[0040] Step S1: Input a portrait photo and use a pre-trained face key-point detection model to perform face key-point detection and locate the coordinates of the two eyes. The face key-point detection model is a model in computer vision that can identify the structural features of a face in an image and accurately mark the coordinates of multiple key points. These key-point features are usually fixed, and these feature points include eyes, nose, ears, corners of the mouth, jaw, the shoulder and neck area below the ears, and key shoulder parts.

[0041] Step S11: Before training the face key-point detection model, it is first necessary to collect and preprocess portrait photos. First, use a labeling tool to label the key parts in the portrait image obtained from the portrait photo. These key parts include eyes, nose, ears, corners of the mouth, jaw, the shoulder and neck area below the ears, and shoulders. Each labeled point should be assigned a unique label to ensure the accuracy and standardization of the labeling information. The label of each marked point not only helps to distinguish different parts but also ensures that the model can accurately learn the spatial relationships and feature information of different parts during the training process. The labeled image data needs to be converted into a text format that can be used for model training together with the corresponding label information. In this way, the model can use this labeled information to automatically identify and infer the positions of each key point in the image. After organizing the labeled data, it is necessary to organize all portrait images and their corresponding label information into a complete training dataset. During the construction of the dataset, to avoid model overfitting and improve its generalization ability, it is necessary to ensure the diversity of the dataset. To comprehensively evaluate the performance of the model, the dataset is usually divided into a training set and a validation set. The training set contains most of the sample data and is used for the training process of the model, while the validation set contains unseen samples and is used to regularly evaluate the performance of the model on new data to ensure that the model does not simply memorize the training data and lose its predictive ability for new samples.

[0042] Step S12: The face key-point detection model extracts features by passing the input image into a convolutional neural network to perform object detection. Specifically, the network predicts the bounding boxes and class probabilities of the objects within these grids. To handle multi-scale detection, Figure 2As shown, exemplarily, the face key point detection model uses hierarchical feature extraction and a feature pyramid structure to ensure the detection of objects of various sizes. Then, non-maximum suppression is used to refine these predictions to filter out duplicate or low-confidence boxes, thereby achieving more accurate object detection. The face key point detection model adopts an improved backbone network and neck architecture, significantly enhancing the feature extraction ability, thus improving the accuracy of object detection and the performance of multi-task processing. The backbone network adopts a faster and more efficient variant of the cross-stage partial bottleneck structure. Exemplarily, this structure achieves accelerated computing through two convolutional layers and a smaller convolutional kernel while maintaining the performance level. To further improve efficiency, exemplarily, the improvement of the backbone network is reflected in adding a cross-stage local spatial attention module after the spatial pyramid pooling fast module, reducing redundant calculations through a depthwise separable method and improving efficiency. The neck is responsible for aggregating features of different resolutions and passing them to the head for prediction. This part usually involves upsampling and concatenating feature maps of different levels to ensure that the model can learn global and local details from information at different scales. The feature pyramid structure upsamples the high-semantic features of the deep layer to the same resolution as the shallow layer features and fuses the upsampled deep layer features with the shallow layer features.

[0043] The cross-stage partial bottleneck structure is a network optimization method aimed at improving the computational efficiency of neural networks. By splitting the computational process of the network into multiple stages and making partial connections between these stages, this structure can transmit information more efficiently, reducing unnecessary calculations while maintaining high performance. The spatial pyramid pooling fast module is used to perform maximum pooling operations at different scales and concatenate the pooling results, thereby capturing features at multiple spatial scales and enhancing the model's adaptability to features of various scales. Its structure diagram is as Figure 3 shown. The cross-stage local spatial attention module enhances the spatial attention in the feature map, helping the model to pay more attention to important regions in the image. This module pools the spatial features, enabling the model to effectively focus on specific regions of interest, thereby improving the accuracy of key point detection. Its structure diagram is as Figure 4 shown.

[0044] Step S13: Improve performance and efficiency during the training process of the face key point detection model. First, by performing various data augmentation operations on the input image, such as rotation, scaling, cropping, and flipping, etc., the diversity of the training data is increased. This not only expands the coverage of the dataset but also effectively improves the generalization ability of the model, reduces the risk of overfitting, and enables the trained model to better adapt to different input images. Second, the present invention adopts a multi-scale training strategy, using images of different scales for training. This strategy enables the model to learn features from small objects to large objects, thereby enhancing the multi-scale detection ability of the model in practical applications. Especially when dealing with face images of different sizes and resolutions, it can maintain good detection accuracy. In addition, in order to enhance the model's learning ability for difficult samples, the present invention also introduces a hard negative sample mining strategy. The hard negative sample mining strategy is an optimization method commonly used in object detection or key point detection tasks, aiming to make the model better learn samples that are difficult to classify or detect by dynamically adjusting the weights of training samples. During the training process, the sample weights are dynamically adjusted, enabling the model to receive more attention when dealing with difficult samples such as complex backgrounds, occluded objects, blurred or small objects. Through this strategy, the model can perform more stably and accurately when facing challenges in real scenarios, further improving the face key point detection accuracy in various environments.

[0045] Step S14: Input the portrait photo into the trained face key point detection model. After being trained, the face key point detection model can automatically identify the face in the image and label its key points. By processing the input image, the model will return a result containing the coordinates of multiple face key points. These coordinates include the eyes, nose, ears, corners of the mouth, jaw, the shoulder and neck area below the ears, and key shoulder parts. Among all the labeled key point coordinates, further screen out the key point coordinates related to both eyes. By extracting the coordinate information of both eyes from the detected key point set, accurate data support is provided for subsequent image processing and optimization steps.

[0046] Step S2: Calculate the required rotation direction and angle based on the coordinates where both eyes are located to ensure the alignment of both eyes in the horizontal direction. Step S21: The present invention analyzes the coordinates of both eyes in the input portrait photo to implement the calculation of the horizontal axis distance and the estimation of the rotation angle. The specific steps are as follows: First, based on the face key point detection model, obtain the coordinates of both eyes in the input portrait photo. Set the left eye coordinates as (x1, y1) and the right eye coordinates as (x2, y2). Calculate the distance difference between both eyes on the horizontal axis to obtain the position difference between both eyes in the horizontal direction. Formula (1) represents its calculation process: Δx = |x2 - x1| (1)

[0047] Δx represents the distance difference; x2 represents the abscissa of the right eye; x1 represents the abscissa of the left eye;

[0048] Step S22: According to the relative relationship between the distance difference between the two eyes on the horizontal axis and the image width, use this difference to judge the horizontal deviation degree of the two eyes, and further calculate the rotation angle. By calculating the ratio of the vertical deviation to the horizontal deviation between the coordinates of the two eyes, use the arctangent function to obtain the rotation angle, which determines the alignment angle required for the two eyes. Then, based on the positive or negative value of the rotation angle, determine the rotation direction. If the angle is positive, the right eye is above, and the rotation direction is clockwise; if the angle is negative, the right eye is below, and the rotation direction is counterclockwise.

[0049] Step S3: Use the matte extraction model to perform precise matte extraction to generate a portrait image with a transparent background, and rotate the portrait image with a transparent background to obtain a portrait image with a transparent background after the eyes are aligned. The matte extraction model is mainly used to accurately separate the foreground from the background and generate a high-quality portrait image with a transparent background. For example Figure 5As shown in the figure, the matting model of the present invention includes a shared encoder module, a pyramid pooling module, a semantic context branch module, and a high-resolution detail branch module. The backbone network of the matting model of the present invention is a high-resolution network, which is also used as a shared encoder. Exemplarily, it consists of five blocks, which are 1 / 2 downsampling, 1 / 4 downsampling, 1 / 8 downsampling, 1 / 16 downsampling, and 1 / 32 downsampling in sequence. At the output of the last stage of the backbone network, one feature map after 1 / 32 downsampling is sent into the pyramid pooling module to obtain richer semantic context information; one feature map after 1 / 4 downsampling is sent into the upsampling module to restore to the same resolution, and then these feature maps are concatenated in the channel dimension to form the initial input of the high-resolution detail branch. A branch design is adopted to clarify the semantic prediction and detail prediction tasks. For the semantic context branch, the feature map obtained by the pyramid pooling module is used as the input of the semantic context. Exemplarily, the semantic context branch consists of five blocks, and each block includes convolution, batch normalization, an activation function, and a bilinear upsampling module, achieving 1 / 16 resolution, 1 / 8 resolution, 1 / 4 resolution, 1 / 2 resolution, and 1 / 1 resolution in sequence. In addition, the output of the semantic context branch is also used for the supervision of the semantic segmentation task, which includes three categories: foreground, background, and transition regions. The input of the high-resolution detail branch is formed by concatenating the intermediate features of the shared encoder after upsampling operations. Specifically, the intermediate features extracted by the shared encoder at different levels are first restored to the same resolution through upsampling operations, and then these feature maps are concatenated in the channel dimension to form the initial input of the high-resolution detail branch. This design can retain more detail information, thereby improving the precise separation ability of the matting model for foreground edges and complex regions. At the same time, a guiding flow is introduced from the intermediate features of 1 / 16 and 1 / 4 resolutions of the semantic context branch to inject semantic information to help the high-resolution details obtain more detailed image details. The guiding flow has a structure as Figure 6 shown. After concatenating the semantic context feature map from the semantic context branch and the detail feature map of the high-resolution detail branch, it is sent into a module containing convolution, batch normalization, a rectified linear unit activation function, and convolution, batch normalization, and an S-shaped activation function to generate a guiding map. Then, the guiding map is subjected to dot multiplication and addition operations with the detail feature map of the high-resolution detail branch to generate the feature map of the next stage of the high-resolution detail. The total cross-entropy loss function is shown in Equation (2), the total loss function is shown in Equation (3), and the overall loss function is shown in Equation (4):

[0050] Ω = Ω f ∪Ω b ∪Ω t

[0051] where L sdenotes the total cross - entropy loss; c denotes the class index; i denotes the pixel index; Ω denotes the set of positions of all pixels in the image; denotes the true label value of the i - th pixel for class c; denotes the predicted probability of the i - th pixel for class c; denotes the natural logarithm of; Ω f denotes belonging to the foreground class; Ω b denotes belonging to the background class; Ω t denotes the set of pixel positions belonging to the third class;

[0052]

[0053] L d denotes the total loss function; Ω t denotes the set of pixels in the target region; i is the pixel index; denotes the transparency loss of the i - th pixel; denotes the gradient loss of the i - th pixel; denotes the predicted transparency value of the i - th pixel, generated by the model; the true transparency value of the i - th pixel; ε denotes a small positive value to avoid the square difference being zero; denotes the gradient of the predicted transparency value of the i - th pixel; denotes the gradient of the true transparency value of the i - th pixel;

[0054]

[0055] L f denotes the overall loss function; i denotes the pixel index; Ω denotes the set of image pixels; denotes the transparency loss of the i - th pixel; denotes the gradient loss of the i - th pixel; denotes the composite loss of the i - th pixel; denotes the predicted composite image value of the i - th pixel; denotes the true composite image value of the i - th pixel; ε denotes a small positive value to avoid the square difference being zero; denotes the predicted transparency value of the i - th pixel; denotes the true foreground pixel value of the i - th pixel; denotes the true background pixel value of the i - th pixel; denotes the probability that a pixel belongs to the background.

[0056] The high-resolution network structure consists of multiple stages, and each stage is processed by network branches with different resolutions. At each stage, the network gradually optimizes and fuses features of different scales through cross-branch information interaction, thereby enhancing the feature expression ability. As the depth of the network increases, it can better capture the detailed information and global semantics in the image. The high-resolution network maintains high-resolution features at each stage, and at the same time generates and fuses low-resolution feature maps to obtain richer semantic information. An effective multi-scale feature fusion strategy is adopted to interactively fuse feature maps with different resolutions. Information is exchanged between low-resolution and high-resolution feature maps, so that each branch can not only learn lower-level detailed information but also incorporate global semantic information from other scales.

[0057] A shared encoder refers to the design in a high-precision natural image matting model where multiple tasks or multiple modules share the same part of the encoding network structure. When sharing the same encoder for upstream and downstream tasks or multiple branches, the network layers of the encoder are jointly used by multiple tasks, thereby reducing the use of computing resources and parameters, and helping to share information between tasks, improving the learning efficiency and generalization ability.

[0058] The pyramid pooling block module performs pooling on the feature map at different scales through pyramid pooling operations, and then splices or fuses the pooled features to obtain multi-scale semantic information. This helps the model better understand and segment objects of different scales in the image. The pyramid pooling operation can effectively expand the receptive field of the neural network, enabling the model to better capture the global and local information in the image. The pyramid pooling block module can be applied to the output feature map of the decoder to provide more context information for the segmentation task. This helps the model more accurately classify pixels into different semantic categories, thereby improving the segmentation accuracy and generalization ability.

[0059] Step S31: The present invention provides a technology based on a deep learning matting model for generating a portrait image with a transparent background.

[0060] The portrait image with a transparent background is a portrait image with a transparent background including an alpha channel. This method accurately segments the input portrait photo, extracts the person region, and generates an image including an alpha channel, where the alpha channel is used to represent the separation region between the portrait and the background, thereby obtaining a portrait image with a transparent background. Specifically, the matting model analyzes the person features in the input image, accurately separates the person contour, and effectively removes the complex background. The finally generated portrait image with a transparent background can not only clearly display the person image but also retain the details of the person, such as hair strands, clothing edges, etc., thereby ensuring that the generated portrait image has a high-quality visual effect and can be easily composited with other backgrounds.

[0061] Step S4: Crop the portrait image with a transparent background according to the rotation result, and synthesize the transparent background with the standard background to process it into a standard ID photo.

[0062] Step S41: Based on the calculated binocular coordinates and rotation angle, first perform a rotation process on the generated portrait image with a transparent background to ensure that the positions of the two eyes in the image are horizontally aligned. This rotation process is based on the coordinate difference between the two eyes and the pre-calculated angle, and precisely rotates the image to the ideal position, thus achieving the required body posture for a standard ID photo. Then, according to the rotation angle of the image and the width-to-height ratio of the image, appropriately crop both sides of the portrait image with a transparent background. The cropping process ensures that the image maintains a reasonable proportion, with the person centered in the image and meeting the size and composition requirements of an ID photo. After cropping, the portrait image with a transparent background will be scaled and synthesized with a preset standard background. This background can be set according to specific requirements, including common standard backgrounds for ID photos such as blue and white. Finally, through the scaling and synthesis of the background, a complete image meeting the standard ID photo format is generated. This step not only ensures the unity and standardization of the image background but also optimizes the presentation of the body posture, guaranteeing that the generated ID photo meets all specification requirements and satisfies the usage needs in practical applications.

[0063] Step S5: Use a pre-trained human shoulder key point detection model to detect the key points of the left and right shoulders of the standard ID photo. The human shoulder key point detection model has the same model structure as the face key point detection model in Step S1, but the training data is different, and it has been specifically labeled and trained for the shoulder part. The human shoulder key point detection model is based on computer vision technology and can extract human body structure features from the input image and accurately locate the key points of the shoulders. These key points include specific parts of the left and right shoulders. Through these key points, the image posture can be further corrected or the symmetry of the shoulder position can be calculated, providing a basis for the standardized processing of the standard ID photo.

[0064] Step S51: Train a human shoulder key point detection model. First, use annotation tools to annotate the key parts in the portrait. There are 12 key points on each side of the face, centered on the face. Each annotation point should be assigned a unique label to ensure the accuracy and standardization of the marked information. The annotated image data needs to be converted into a text format that can be used for model training together with the corresponding label information. After organizing the annotated data, all portrait images and their corresponding label information need to be organized into a complete training dataset. During the construction of the dataset, to avoid model overfitting and improve its generalization ability, the diversity of the dataset must be ensured. To comprehensively evaluate the performance of the model, the dataset is usually divided into a training set and a validation set. The training set contains most of the sample data and is used for the training process of the model, while the validation set contains unseen samples and is used to regularly evaluate the performance of the model on new data to ensure that the model does not simply memorize the training data and lose its predictive ability for new samples.

[0065] Step S52: Input the generated ID photo into the human shoulder key point detection model. After being fully trained, the human shoulder key point detection model can automatically identify the human body in the image and annotate its key points. Through the processing of the input image, the model outputs the results containing the coordinates of multiple human key points, with 12 key points annotated on each side of the face, centered on the face. Among all the annotated key points, further extract and filter out the key point coordinates of the left and right shoulders as the basic data for subsequent processing. Step S6: Calculate the difference between the average height of the shoulders and the preset reference height based on the key point coordinates of the left and right shoulders, and determine the shoulder on the side that needs to be adjusted.

[0066] Step S61: According to the obtained coordinates of the key points of the left and right shoulders, first calculate the average height of the left and right shoulders respectively. Then, set a preset value as the reference height of the shoulders, compare the average height of the left and right shoulders with the reference height respectively, calculate the height difference between each side of the shoulder and the reference height, and take its absolute value. By comparing the magnitudes of the height differences on both sides, the side with the larger difference is determined as the "adjustment side" that needs to be adjusted, and the series of key points on this side are defined as the "key points before adjustment".

[0067] Step S62: Calculate the coordinates of the key points after adjustment one by one according to the relative position relationship between the key points before adjustment and the key points of the other shoulder. Specifically, the abscissa of the key point after adjustment remains the same as that of the key point before adjustment to ensure the accuracy of the horizontal position; while the ordinate of the key point after adjustment is adjusted to be the same as the ordinate of the corresponding key point on the other side to achieve vertical alignment. This adjustment process aims to ensure the symmetry of the shoulders and their related key points on both sides, while optimizing the accuracy and balance of the overall structure, and ensuring the accuracy and consistency of the model in the detection and annotation tasks of human key points.

[0068] Step S7: Using the local deformation algorithm, adjust the shoulder key points that need to be adjusted so that their heights are aligned with the reference values, and obtain the result image with the eyes and shoulders level in the original image. The local deformation algorithm is a technology commonly used in the fields of image processing and computer graphics. By deforming and adjusting specific regions in the image, various functions such as geometric deformation, detail enhancement, and defect repair of the image can be achieved. In the present invention, the local deformation algorithm is used to adjust the positions of the shoulder key points so that they are aligned with the set reference height, thereby realizing the standardized processing of the image.

[0069] Step S71: First, according to the coordinates of the shoulder key points before adjustment and the target key points after adjustment, select the shoulder region in the image as the region to be adjusted, and define the shoulder key points as control points. Control points are used to determine the benchmark and range of image deformation and are crucial parameters in the deformation process. Then, according to the positions and distributions of the control points, construct a deformation grid that covers the entire image. Each cell of this grid is determined by the control points, and its structure and distribution directly affect the smoothness and accuracy of the deformation effect. By analyzing the relationship between the control points and the grid nodes, calculate the offset between the control points and the target positions after adjustment, and determine the deformation parameters that need to be applied during the deformation process.

[0070] Step S72: According to the above deformation parameters, use the grid interpolation method to calculate the positions of each pixel point in the image after deformation one by one. To ensure the visual continuity and accuracy of the image after deformation, color interpolation calculation needs to be performed on the pixel points at the new positions. Reconstruct the entire image after the deformation calculation is completed to generate the adjusted image. In the adjusted image, the relative positions of the eyes and shoulders reach the effect of horizontal alignment, ensuring the symmetry of the image and the standardized processing of the human body posture, and meeting the production requirements of standard ID photos. The local deformation algorithm can not only accurately adjust the positions of the shoulder key points but also ensure the overall quality and detail consistency of the image, realizing the efficient adjustment of the shoulder region and the optimization of the visual effect. Interpolation calculation is a technology widely used in mathematics, image processing, data analysis, and other fields. Its core purpose is to estimate the values of unknown points between known data points, thereby constructing a more continuous and smooth result.

[0071] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes and combinations of elements can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for automatically processing a portrait photo into a standard ID photo, characterized in that: The following steps are included: Step S1: input a portrait photo, use a pre-trained facial key point detection model to detect facial key points, and locate the coordinates of both eyes; Step S2: Calculate the direction and angle of rotation required according to the coordinates of both eyes to ensure that both eyes are aligned in the horizontal direction; Step S3: Use the cutout model to perform precision cutout to generate a transparent background portrait image, and rotate the transparent background portrait image to obtain a transparent background portrait image after eye alignment; Step S4: cropping the transparent background portrait image according to the rotation result, and synthesizing the transparent background with the standard background to process it into a standard ID photo; Step S5: using a pre-trained human shoulder key point detection model to detect key points of the left and right shoulders of the standard ID photo; Step S6: Calculate the difference between the average height of the shoulders and the preset reference height according to the coordinates of the key points of the left and right shoulders, and determine the shoulder that needs to be adjusted; Step S7: Using a local deformation algorithm, adjust the shoulder key points that need to be adjusted so that their heights are aligned with the reference values, and obtain a result image in which the eyes and shoulders of the original image are flush.

2. The method for automatically processing a portrait photo into a standard ID photo according to claim 1, characterized in that: The face key point detection model is a model in computer vision that can identify the structural features of the face in the image and accurately mark multiple key point features, including eyes, nose, ears, corners of the mouth, jaw, shoulder and neck below the ears, and key shoulder parts; The face keypoint detection model uses hierarchical feature extraction and feature pyramid structure to ensure that objects of various sizes are detected, and uses non-maximum suppression to refine these predictions to filter out duplicate or low-confidence boxes, thereby achieving more accurate object detection.

3. The method for automatically processing a portrait photo into a standard ID photo according to claim 1, characterized in that: The face key point detection model adopts an improved backbone network and neck architecture. The backbone network adopts a faster and more efficient variant of the cross-stage partial bottleneck structure. The improved backbone network adds a layer of cross-stage local spatial attention module after the spatial pyramid pooling fast module.

4. The method for automatically processing a portrait photo into a standard ID photo according to claim 1, characterized in that: The facial key point detection comprises the following steps: Collect and preprocess portrait data; The face landmark detection model passes the input image into a convolutional neural network to extract features to perform object detection; Self-training of facial key point detection model; The facial key point detection model automatically identifies faces in images and annotates their key points.

5. The method for automatically processing a portrait photo into a standard ID photo according to claim 1, characterized in that: The step of calculating the direction and angle of rotation required according to the coordinates of the two eyes includes the following steps: Analyze the coordinates of both eyes in the input portrait photo, calculate the horizontal axis distance and the rotation angle; According to the relative relationship between the distance difference between the two eyes on the horizontal axis and the image width, the difference is used to determine the degree of horizontal displacement of the two eyes, and further calculate the rotation angle.

6. The method for automatically processing a portrait photo into a standard ID photo according to claim 1, characterized in that: The cutout model is used to accurately separate the foreground from the background and generate a high-quality transparent background portrait. The cutout model includes a shared encoder module, a pyramid pooling module, a semantic context branch module, and a high-resolution detail branch module. The backbone network of the cutout model adopts a high-resolution network, which is also used as a shared encoder. The output of the last stage of the backbone network is sent to the pyramid pooling module to obtain richer semantic context information.

7. The method for automatically processing a portrait photo into a standard ID photo according to claim 1, characterized in that: The semantic The following branch module consists of five blocks, each of which includes convolution, batch normalization, activation function and a bilinear upsampling module. The output of the semantic context branch is also used to supervise the semantic segmentation task, which includes three categories: foreground, background and transition area.

8. The method for automatically processing a portrait photo into a standard ID photo according to claim 1, characterized in that: The input of the high-resolution detail branch is concatenated from the intermediate features of the shared encoder after upsampling.

9. The method for automatically processing a portrait photo into a standard ID photo according to claim 1, characterized in that: The intermediate features extracted by the shared encoder at different levels are first restored to the same resolution through upsampling operations, and then these feature maps are concatenated in the channel dimension to form the initial input of the high-resolution detail branch.

10. The method for automatically processing a portrait photo into a standard ID photo according to claim 1, characterized in that: The semantic context feature map of the semantic context branch module and the detail feature map of the high-resolution detail branch module are subjected to guided flow processing, wherein the guided flow processing is to splice the semantic context feature map from the semantic context branch and the detail feature map of the high-resolution detail branch, and then send them into a process including convolution, batch normalization, rectified linear unit activation function and convolution, batch normalization, S-shaped activation function to generate a guided map, and then perform dot multiplication and addition operations on the guided map and the detail feature map of the high-resolution detail branch to generate a feature map of the next stage of high-resolution detail.

Citation Information

Patent Citations

  • Identification photo camera capable of performing human image posture photography prompting and human image posture detection method

    CN105046246A

  • Certificate picture camera and certificate picture photographing method

    CN105120167A

  • Identification camera capable of detecting photographing quality and photographing quality detecting method

    CN105139404A

  • Certificate photo detection method and device, electronic equipment and storage medium

    CN111401242A

  • Image synthesis method, terminal and storage medium

    CN113810588A

Cited By

  • Identification photo automatic generation method and system

    CN120510638A