Face direction correction method and device, computer device and storage medium

By using face detection and correction matrix processing, the problem of decreased recognition rate caused by inconsistent face image orientation is solved, achieving efficient face orientation correction and recognition.

CN115661883BActive Publication Date: 2026-05-19TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2021-07-08
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

When users shoot short videos using terminal devices, inconsistent orientation of facial images leads to a decrease in facial recognition rate, and existing technologies struggle to effectively correct facial poses.

Method used

The face detection module obtains face localization points and candidate boxes, determines the confidence level, filters target detection boxes, obtains face annotation points, calculates the orientation correction matrix, and performs face orientation correction based on the correction matrix.

Benefits of technology

It improves the face recognition rate and enhances the efficiency of face orientation correction, enabling simultaneous correction of the orientation of multiple face images, and is suitable for video calls, video editing, and video playback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115661883B_ABST
    Figure CN115661883B_ABST
Patent Text Reader

Abstract

The application relates to a face direction correction method and device, computer equipment and a storage medium. The method comprises the following steps: obtaining a face image to be detected, performing face detection on the face image to be detected, obtaining at least one group of face positioning points and at least one face candidate frame corresponding to each group of face positioning points; determining a confidence degree according to positioning accuracy and a classification probability value; screening a target detection frame corresponding to each group of face positioning points from the face candidate frame based on the confidence degree of each face candidate frame; obtaining face marking points; the face marking points are face positioning points in a face image in a standard posture; determining a direction correction matrix corresponding to each group of face positioning points according to the face marking points and the face positioning points; and performing face direction correction based on the target detection frame corresponding to each group of face positioning points and the direction correction matrix corresponding to the face positioning points. The face direction can be corrected by using the method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, computer device, and storage medium for facial orientation correction. Background Technology

[0002] With the development of science and technology, facial recognition technology is being used more and more widely in the field of image processing. For example, when users shoot short videos using their devices, facial recognition technology can locate faces in the frame and apply stickers or beautification effects to the identified faces. However, when users shoot short videos, they often use landscape, portrait, or other angles, resulting in short videos containing multiple facial images from different angles.

[0003] Face detection, as one of the most crucial modules in face recognition technology, directly impacts the recognition rate by determining the orientation of faces in the acquired images. Therefore, correcting facial poses in face images has become an urgent problem to solve. Summary of the Invention

[0004] Therefore, it is necessary to provide a method, apparatus, computer device, and storage medium for facial orientation correction that can perform facial orientation correction, in response to the above-mentioned technical problems.

[0005] A method for facial orientation correction, the method comprising:

[0006] Acquire a face image to be detected, perform face detection on the face image to be detected, and obtain at least one set of face localization points and at least one face candidate box corresponding to each set of face localization points;

[0007] Determine the localization accuracy and classification probability value of each face candidate box, and determine the confidence level corresponding to each face detection box based on the localization accuracy and the classification probability value;

[0008] Based on the confidence level of each face candidate box, target detection boxes corresponding to each group of face localization points are selected from the face candidate boxes.

[0009] Obtain face annotation points; the face annotation points are face localization points in a face image under standard pose.

[0010] Based on the face annotation points and each group of face positioning points, determine the orientation correction matrix corresponding to each group of face positioning points;

[0011] Based on the target detection boxes corresponding to each group of face localization points, and according to the orientation correction matrix corresponding to the respective face localization points, face orientation correction is performed.

[0012] A facial orientation correction device, the device comprising:

[0013] The face detection module is used to acquire a face image to be detected, perform face detection on the face image to be detected, and obtain at least one set of face localization points and at least one face candidate box corresponding to each set of face localization points.

[0014] The matrix determination module is used to determine the localization accuracy and classification probability value of each face candidate box, and to determine the confidence level of each face detection box based on the localization accuracy and classification probability value; based on the confidence level of each face candidate box, to select target detection boxes corresponding to each group of face localization points from the face candidate boxes; to obtain face annotation points; the face annotation points are face localization points in the face image under standard pose; and to determine the orientation correction matrix corresponding to each group of face localization points based on the face annotation points and each group of face localization points.

[0015] The orientation correction module is used to perform face orientation correction based on the target detection boxes corresponding to each group of face localization points and the orientation correction matrix corresponding to the respective face localization points.

[0016] In one embodiment, the face detection module further includes a feature extraction module, used to extract face features from the face image to be detected at multiple scales to obtain face feature images at multiple scales; adjust the number of channels in each face feature image at multiple scales to obtain multiple face feature images with the same number of channels; and determine at least one set of face localization points in the face image to be detected, and at least one face candidate box corresponding to each set of face localization points, based on the multiple face feature images with the same number of channels.

[0017] In one embodiment, the face correction method is executed through a face correction model, which includes a face detection structure and a direction detection structure. The face detection module is further configured to perform face region detection on the face image to be detected using the face detection structure and based on multiple face feature images with the same number of channels, to obtain at least one face candidate box corresponding to each face region; and to perform face localization point detection on the face region in the face image to be detected using the direction detection structure and based on multiple face feature images with the same number of channels, to obtain a set of face localization points corresponding to each face region.

[0018] In one embodiment, the face correction method is executed through a face correction model, which includes a localization detection structure and a classification detection structure. The matrix determination module further includes a target detection box determination module, used to predict the localization accuracy of the current face candidate box for each face candidate box in at least one face candidate box using the localization detection structure. The localization accuracy characterizes the degree of difference between the face candidate box and the corresponding standard face box. For each face candidate box in at least one face candidate box, the classification detection structure predicts the classification probability value of the image region selected by the current face candidate box as a face. For each face candidate box in at least one face candidate box, the confidence level of the current face candidate box is determined based on the localization accuracy and classification probability value corresponding to the current face candidate box.

[0019] In one embodiment, the target detection box determination module is further configured to, for each group of face localization points, select the face candidate box with the highest confidence from the at least one face candidate box corresponding to the current group of face localization points; and use the face candidate box with the highest confidence as the target detection box corresponding to the current group of face localization points.

[0020] In one embodiment, the matrix determination module is further configured to determine a direction correction matrix corresponding to each group of face positioning points based on the face annotation points and each group of face positioning points, including: determining a first coordinate matrix based on the position information of the face annotation points; determining a current second coordinate matrix for each group of face positioning points based on the position information of the current face positioning point; and performing an affine transformation on the first coordinate matrix and the current second coordinate matrix to obtain the direction correction matrix corresponding to the current group of face positioning points.

[0021] In one embodiment, the orientation correction module is further configured to perform face orientation correction on the image region in the current target detection box corresponding to the current group of face positioning points for each group of face positioning points in at least one group of face positioning points, based on the current orientation correction matrix corresponding to the current group of face positioning points.

[0022] In one embodiment, the orientation correction module is further configured to generate a blank face image with the same size as the current target detection box; the blank face image is a blank image containing multiple first pixels; the second pixels in the image region selected by the current target detection box are traversed sequentially; for the currently traversed second pixel, based on the coordinates of the currently traversed second pixel and the current orientation correction matrix, a target first pixel corresponding to the currently traversed second pixel is determined from the multiple first pixels in the blank image, and the pixel information of the currently traversed second pixel is used as the pixel information of the corresponding target first pixel; after traversing all the second pixels and obtaining the pixel information of each first pixel, the correction result of face orientation correction for the image region in the current target detection box is obtained based on the pixel information of each first pixel.

[0023] In one embodiment, the face orientation correction device further includes a training module for acquiring training samples and acquiring standard key points, standard face bounding boxes, and classification labels corresponding to the training samples; performing face detection on the training samples using the face correction model to be trained, obtaining at least one set of predicted key points, at least one predicted candidate box corresponding to each set of predicted key points, a predicted localization accuracy and a standard localization accuracy corresponding to each predicted candidate box, and a predicted classification corresponding to each predicted candidate box; constructing a loss function based on a first difference between the predicted candidate box and the corresponding standard face bounding box, a second difference between the predicted key points and the corresponding standard key points, a third difference between the predicted localization accuracy and the corresponding standard localization accuracy, and a fourth difference between the predicted classification and the corresponding classification label; and training the face correction model based on the loss function until the training termination condition is met, thus obtaining a trained face correction model.

[0024] In one embodiment, the training module is further configured to: perform face region detection on the training samples using the face detection structure in the face correction model to be trained, to obtain at least one predicted candidate box corresponding to each face region; perform face localization point detection on the face regions in the training samples using the orientation detection structure in the face correction model to be trained, to obtain a set of predicted key points corresponding to each face region; perform localization accuracy detection on each predicted candidate box using the localization detection structure in the face correction model to be trained, to obtain the predicted localization accuracy corresponding to each predicted candidate box; determine the overlap between each predicted candidate box and the corresponding standard face box using the localization detection structure in the face correction model to be trained, and determine the standard localization accuracy corresponding to each predicted candidate box based on the overlap; and determine the probability value that the image region selected by each predicted candidate box is a face using the classification detection structure in the face correction model to be trained, to obtain the corresponding predicted classification.

[0025] In one embodiment, the face orientation correction device is further configured to acquire multiple initial face images, and rotate each of the multiple initial face images by a random angle to obtain a rotated face image; adjust the scale and pixel value of each rotated face image to obtain an adjusted face image; and stitch the adjusted face images together to obtain training samples.

[0026] In one embodiment, the face orientation correction device is further configured to: when the face image to be detected is a face image captured during a video call, adjust the image displayed in the video call based on the target face image obtained after face orientation correction; when the face image to be detected is a face image captured during video editing, replace the face image in the video to be edited with the target face image obtained after face orientation correction; and when the face image to be detected is a face image captured during video playback, add face orientation correction effects to the played video based on the target face image obtained after face orientation correction.

[0027] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program performing the following steps:

[0028] Acquire a face image to be detected, perform face detection on the face image to be detected, and obtain at least one set of face localization points and at least one face candidate box corresponding to each set of face localization points;

[0029] Determine the localization accuracy and classification probability value of each face candidate box, and determine the confidence level corresponding to each face detection box based on the localization accuracy and the classification probability value;

[0030] Based on the confidence level of each face candidate box, target detection boxes corresponding to each group of face localization points are selected from the face candidate boxes.

[0031] Obtain face annotation points; the face annotation points are face localization points in a face image under standard pose.

[0032] Based on the face annotation points and each group of face positioning points, determine the orientation correction matrix corresponding to each group of face positioning points;

[0033] Based on the target detection boxes corresponding to each group of face localization points, and according to the orientation correction matrix corresponding to the respective face localization points, face orientation correction is performed.

[0034] A computer-readable storage medium having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0035] Acquire a face image to be detected, perform face detection on the face image to be detected, and obtain at least one set of face localization points and at least one face candidate box corresponding to each set of face localization points;

[0036] Determine the localization accuracy and classification probability value of each face candidate box, and determine the confidence level corresponding to each face detection box based on the localization accuracy and the classification probability value;

[0037] Based on the confidence level of each face candidate box, target detection boxes corresponding to each group of face localization points are selected from the face candidate boxes.

[0038] Obtain face annotation points; the face annotation points are face localization points in a face image under standard pose.

[0039] Based on the face annotation points and each group of face positioning points, determine the orientation correction matrix corresponding to each group of face positioning points;

[0040] Based on the target detection boxes corresponding to each group of face localization points, and according to the orientation correction matrix corresponding to the respective face localization points, face orientation correction is performed.

[0041] A computer program product or computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the following steps: acquiring a face image to be detected; performing face detection on the face image to be detected to obtain at least one set of face localization points and at least one face candidate box corresponding to each set of face localization points; determining the localization accuracy and classification probability value of each face candidate box, and based on... Based on the localization accuracy and the classification probability value, determine the confidence level corresponding to each face detection box; based on the confidence level of each face candidate box, select target detection boxes corresponding to each group of face localization points from the face candidate boxes; obtain face annotation points; the face annotation points are face localization points in the face image under standard pose; determine the orientation correction matrix corresponding to each group of face localization points according to the face annotation points and each group of face localization points; perform face orientation correction based on the target detection boxes corresponding to each group of face localization points and according to the orientation correction matrix corresponding to the corresponding face localization points.

[0042] The aforementioned face orientation correction method, apparatus, computer equipment, storage medium, and computer program, in the face orientation correction method, by acquiring an image to be detected, face detection can be performed on the image to be detected, obtaining face localization points and face candidate boxes. By acquiring face candidate boxes, the confidence level of each face candidate box can be determined, thereby allowing the selection of target detection boxes corresponding to each group of face localization points from the face candidate boxes based on the confidence level. By acquiring face localization points and face annotation points, an orientation correction matrix corresponding to each group of face localization points can be determined based on the face localization points and face annotation points. Thus, face orientation correction can be performed based on the determined orientation correction matrix and the target detection boxes.

[0043] Furthermore, since at least one set of face localization points and the corresponding target detection boxes for each set of face localization points can be determined, face orientation correction can be performed simultaneously on multiple faces in the face image to be detected based on the determined at least one set of face localization points and the corresponding target detection boxes for each set of face localization points, thus greatly improving the efficiency of face orientation correction. Attached Figure Description

[0044] Figure 1 This is an application environment diagram of a face orientation correction method in one embodiment;

[0045] Figure 2 This is a flowchart illustrating a face orientation correction method in one embodiment;

[0046] Figure 3This is a schematic diagram of face detection results in one embodiment;

[0047] Figure 4 This is a schematic diagram of face orientation correction in one embodiment;

[0048] Figure 5 This is a schematic diagram of face orientation correction in another embodiment;

[0049] Figure 6 This is a schematic diagram of the framework of a face correction model in one embodiment;

[0050] Figure 7 This is a schematic diagram of training samples in one embodiment;

[0051] Figure 8 This is a schematic diagram illustrating the target face image in one embodiment;

[0052] Figure 9 This is a flowchart illustrating a face orientation correction method in a specific embodiment.

[0053] Figure 10 This is a flowchart illustrating the face orientation correction method in another specific embodiment;

[0054] Figure 11 This is a structural block diagram of a face orientation correction device in one embodiment;

[0055] Figure 12 This is a structural block diagram of a face orientation correction device in one embodiment;

[0056] Figure 13 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0057] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0058] Figure 1 This is an application environment diagram illustrating a face orientation correction method in one embodiment. (Refer to...) Figure 1The face orientation correction method is applied to a face orientation correction system 100. The face orientation correction system 100 includes a terminal 102 and a server 104. The terminal 102 and server 104 are connected via a network. The terminal 102 can be a desktop terminal or a mobile terminal, including but not limited to mobile phones, tablets, laptops, smart voice interaction devices, smart home appliances, and in-vehicle terminals. The server 104 can be implemented using a standalone server or a server cluster composed of multiple servers. Both the terminal 102 and server 104 can be used independently to execute the face orientation correction method provided in this embodiment. The terminal 102 and server 104 can also be used collaboratively to execute the face orientation correction method provided in this embodiment. Taking the example of terminal 102 and server 104 working together to execute the face orientation correction method provided in this embodiment, terminal 102-1 can display a face image to be detected and send the face image to be detected to server 104, so that server 104 can perform face orientation correction on the face image to be detected using a face correction model to obtain a target face image after face orientation correction, and return the target face image to terminal 102-2 for display. Terminal 102-1 and terminal 102-2 can be the same terminal or different terminals.

[0059] It should also be noted that this application relates to the field of Artificial Intelligence (AI) technology. AI refers to the theories, methods, technologies, and application systems that utilize digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine capable of reacting in a manner similar to human intelligence. AI essentially studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0060] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0061] This application specifically relates to Computer Vision (CV) technology. Computer vision is a science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in recognizing and measuring targets, and further performs image processing to make the computer-processed images more suitable for human observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), autonomous driving, intelligent transportation, and other technologies, as well as common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0062] In one embodiment, such as Figure 2 As shown, a method for facial orientation correction is provided, which is applied to... Figure 1 Taking a computer device as an example, this computer device can specifically be... Figure 1 The face orientation correction method for terminals or servers includes the following steps:

[0063] Step S202: Obtain the face image to be detected, perform face detection on the face image to be detected, and obtain at least one set of face localization points and at least one face candidate box corresponding to each set of face localization points.

[0064] In this context, face localization points refer to points used to locate key facial features, such as the center of the eyeball, the brow peak, the corner of the mouth, and the center of the nose. A set of face localization points can include multiple points; for example, the center of the left eyeball, the center of the right eyeball, and the center of the nose can be considered as a set of face localization points. Face candidate boxes refer to the detection boxes output by the face correction model used to select facial regions. For the same facial region, the face correction model can output multiple face candidate boxes. The face correction model is a machine learning model, which learns from samples to acquire the ability to correct facial orientation.

[0065] Specifically, when a face image to be detected is obtained, the computer device can input the face image into a face correction model. The face correction model detects the face region and the face localization points within that region, obtaining a set of face localization points for each face region, and at least one corresponding face candidate box for each face region. The number of face localization points included in a set can be freely set according to requirements. For example, refer to... Figure 3When it is determined that there are three faces in the image to be detected, the face correction model can output a set of face localization points for each of the three faces, and output at least one face candidate box for each of the three faces. Figure 3 A schematic diagram of face detection results in one embodiment is shown.

[0066] In one embodiment, a video recording application runs on the terminal, allowing the user to record videos. The video recording application may display a record button. During recording, the user can touch the record button to enter recording mode and adjust the phone angle to capture videos of the face from multiple angles, obtaining the recorded video. When recording ends, the terminal can send the user-recorded video to a server, enabling the server to extract the image of the face to be detected from the video.

[0067] In one embodiment, the face correction model performs key point detection on the center point of the left eyeball, the center point of the right eyeball, and the center point of the nose for each face region, thereby obtaining a set of face localization points corresponding to each face region.

[0068] In one embodiment, the face correction model includes a face detection structure and a orientation detection structure. When a face image to be detected is obtained, the face correction model can extract face features from the face image based on a preset face feature extraction strategy. The face detection structure can then identify face regions based on the extracted face features and output candidate bounding boxes for selecting these regions. The orientation detection structure can detect face localization points based on the extracted face features, obtaining a set of face localization points in each face region. The face features may include facial texture features. Facial texture features can reflect the pixel depth of facial organs, including the nose, ears, eyebrows, cheeks, or lips. Facial texture features may include the color value distribution and brightness value distribution of facial image pixels.

[0069] Step S204: Determine the localization accuracy and classification probability value of each face candidate box, and determine the confidence level of each face detection box based on the localization accuracy and classification probability value.

[0070] Specifically, when at least one face candidate box is generated, the face correction model can determine the localization accuracy and classification accuracy of each face candidate box based on the accuracy of the candidate boxes themselves, and then combine the localization accuracy and classification accuracy to obtain the confidence score corresponding to each face detection box. For example, the face correction model multiplies the localization accuracy corresponding to the current face candidate box by the classification accuracy corresponding to the current face candidate box to obtain the confidence score of the current face candidate box.

[0071] In one embodiment, the face correction method is executed through a face correction model, which includes a localization detection structure and a classification detection structure. The method determines the localization accuracy and classification probability value of each face candidate box, and determines the confidence level of each face detection box based on the localization accuracy and classification probability value. This includes: for each face candidate box in at least one face candidate box, predicting the localization accuracy of the current face candidate box using the localization detection structure; the localization accuracy characterizes the degree of difference between the face candidate box and the corresponding standard face box; for each face candidate box in at least one face candidate box, predicting the classification probability value of the image region selected by the current face candidate box as a face using the classification detection structure; and for each face candidate box in at least one face candidate box, determining the confidence level of the current face candidate box based on the localization accuracy and classification probability value corresponding to the current face candidate box.

[0072] Specifically, when at least one face candidate bounding box is obtained, for each face candidate bounding box within that box, the localization detection structure can determine the localization accuracy of the current face candidate bounding box based on its position coordinates. The localization accuracy of the face candidate bounding box characterizes the degree of difference between the face candidate bounding box and the corresponding standard face bounding box; it can be used as a value to express a model's confidence in its ability to select face regions. It is easy to understand that the more complete the face selected by a face candidate bounding box, the higher its corresponding localization accuracy.

[0073] For each face candidate box within at least one face candidate box, the classification and detection structure performs classification prediction on the image region selected by the current face candidate box, outputting the classification probability value that the image region selected by the current face candidate box is a face. Further, the face correction model combines the localization accuracy and classification probability value corresponding to the current face candidate box to obtain the confidence level of the current face candidate box. For example, the face correction model can perform a weighted summation of the localization accuracy and classification probability value to obtain the confidence level of the current face candidate box. It is worth noting that both the localization accuracy and classification probability value can be controlled within a preset range, for example, both can be within the range [0, 1].

[0074] In one embodiment, for each face candidate box in at least one face candidate box, the localization detection structure can determine the overlap between the current face candidate box and the corresponding standard candidate box, and determine the localization accuracy of the current face candidate box based on the overlap. It is easily understood that in model applications, the image to be detected input into the face correction model does not include the standard face detection box; therefore, the standard face detection box is actually a face detection box predicted by the localization detection structure to characterize the standard box.

[0075] In the above embodiments, the confidence level of the corresponding face candidate box is determined by combining the positioning accuracy and the classification probability value. Compared with the traditional method of using the classification probability value as the confidence level of the face candidate box, the embodiments of this application can reduce the probability of face candidate boxes with high positioning accuracy but low classification probability value being filtered in the process of determining the target detection box based on the confidence level, thereby improving the accuracy of the determined target detection box.

[0076] Step S206: Based on the confidence level of each face candidate box, select the target detection boxes corresponding to each group of face localization points from the face candidate boxes.

[0077] Specifically, when at least one face candidate box is generated, the face correction model selects target detection boxes corresponding to each group of face localization points from the face candidate boxes based on the confidence level of the face candidate boxes. The confidence level characterizes the accuracy of the face candidate boxes. For example, the face correction model uses face candidate boxes with a confidence level exceeding a preset threshold as target detection boxes. Another example is that the face correction model performs non-maximum suppression (NMS) on at least one face candidate box to obtain the target detection box corresponding to each face localization point, that is, the target detection box corresponding to each face region.

[0078] Step S208: Obtain face annotation points; face annotation points are face localization points in a face image under standard pose.

[0079] Specifically, a set of face annotation points can be pre-set. When face orientation correction is needed, the face correction model can obtain this pre-annotated set of face annotation points. These face annotation points are the face localization points in a face image under standard pose. Standard pose refers to the face pose after orientation correction; for example, when the corrected face is a frontal face, the standard pose is a frontal face pose. A face image under standard pose refers to a face image that includes a face in a standard pose. It's easy to understand that the face parts annotated by the annotation points correspond to the face parts annotated by the key points. For example, when the face correction model needs to perform key point detection on the center points of the left and right eyeballs and the tip of the nose, it can mark these points in the face image under standard pose, thus obtaining a set of face annotation points.

[0080] In one embodiment, when an initial face image in a standard pose is obtained, the computer device can scale the initial face image to a preset size to obtain a standard face image, and use preset face positioning points in the standard face image as face annotation points. For example, the center points of the left eyeball, right eyeball, and nose tip in the standard face image can be marked to obtain a set of face annotation points.

[0081] In one embodiment, multiple initial face images in standard poses can be acquired, and each initial face image is scaled to obtain multiple standard face images. A computer device marks preset face localization points in each standard face image and calculates the average coordinates of face localization points at the same face location. The average coordinates of face localization points at multiple face locations are used as a set of face annotation point coordinates. For example, the average coordinates of multiple left eyeball center points can be calculated to obtain the coordinates of the face annotation point corresponding to the left eyeball center point.

[0082] Step S210: Based on the face annotation points and each group of face positioning points, determine the orientation correction matrix corresponding to each group of face positioning points.

[0083] Specifically, since the image formed by a set of face localization points and the image formed by a set of face annotation points can be considered to have undergone at least one of the three transformations: translation, rotation, and scaling, the image formed by a set of face annotation points can be translated, rotated, or scaled to obtain the image formed by a set of face localization points. In other words, an affine transformation occurs between the face localization points and the face annotation points. An affine transformation, also known as an affine mapping, is a geometric transformation in which a vector space undergoes a linear transformation followed by a translation to transform it into another vector space. This linear transformation can be either a rotation or a scaling transformation.

[0084] When a set of face annotation points and at least one set of face localization points are obtained, the orientation detection structure in the face correction model can determine the coordinate values ​​of each face annotation point in the set of face annotation points. Similarly, for each set of face localization points in the at least one set, the orientation detection structure can determine the coordinate values ​​of each current face localization point in the current set. Based on the coordinate values ​​of the face annotation points and the current face localization points, the orientation correction matrix corresponding to the current set of face localization points can be determined. For example, an affine transformation can be performed on the coordinate values ​​of the face annotation points and the current face localization points to obtain the corresponding orientation correction matrix.

[0085] In one embodiment, determining the orientation correction matrix corresponding to each group of face positioning points based on face annotation points and each group of face positioning points includes: determining a first coordinate matrix based on the position information of the face annotation points; determining a current second coordinate matrix for each group of face positioning points in at least one group of face positioning points based on the position information of the current face positioning point; and performing an affine transformation on the first coordinate matrix and the current second coordinate matrix to obtain the orientation correction matrix corresponding to the current group of face positioning points.

[0086] Specifically, when a set of marker points is obtained, the orientation detection structure can determine the coordinate values ​​of each marker point in the set and determine a first coordinate matrix based on the coordinate values ​​of each marker point. For each group of face localization points in at least one set, the orientation detection structure can determine a current second coordinate matrix based on the coordinate values ​​of the current face localization points included in the current group. Further, the orientation detection structure can perform an affine transformation on the first coordinate matrix and the current second coordinate matrix to obtain the orientation correction matrix corresponding to the current group of face localization points.

[0087] In one embodiment, the orientation detection structure can determine the target detection box corresponding to the current group of face localization points, and establish a Cartesian coordinate system based on the current target detection box to obtain the coordinate value of the current face localization point. For example, a Cartesian coordinate system can be established with the lower left corner of the target detection box as the origin to calculate the coordinate value of the current face localization point. Correspondingly, the orientation detection structure can also establish a Cartesian coordinate system based on the face image in standard pose to obtain the coordinate value of each face annotation point.

[0088] In one embodiment, the orientation detection structure can be based on the affine transformation formula A*T. tol =A′ determines the direction correction matrix T tol Where A is the first coordinate matrix and A′ is the second coordinate matrix.

[0089] In one embodiment, the first coordinate matrix may be: Where (x1, y1) can be the coordinates of the first face marker in a set of face markers, and (xn, yn) can be the coordinates of the nth face marker in a set of face markers. The third column of the first coordinate matrix represents homogeneous coordinates. Similarly, the second coordinate matrix can be... (x1′,y1′) can be the coordinates of the first face location point in a set of face location points, and (xn′,yn′) can be the coordinates of the nth face location point in a set of face location points.

[0090] In the above embodiments, the corresponding orientation correction matrix can be quickly obtained by performing an affine transformation on the first coordinate matrix and the second coordinate matrix, thereby improving the efficiency of face orientation correction.

[0091] Step S212: Based on the target detection boxes corresponding to each group of face localization points, and according to the orientation correction matrix corresponding to the corresponding face localization points, face orientation correction is performed.

[0092] Specifically, since the orientation correction matrix represents the angle of rotation, translation distance, and scaling ratio of the face to be corrected compared to the face in the standard pose, the face correction model can perform face orientation correction on each face in the face image to be detected based on the target detection box corresponding to each group of face localization points and according to the orientation correction matrix corresponding to the corresponding face localization points.

[0093] In one embodiment, reference Figure 4 When the image to be detected contains only one target detection box, that is, when the image to be detected contains only one face, the face correction model can perform face orientation correction on the image region in the image to be detected based on the orientation correction matrix corresponding to the face localization point. For example, the face correction model can multiply the coordinate value of each pixel in the image to be detected by the inverse of the orientation correction matrix to obtain the corrected coordinate value of each pixel. The face correction model then migrates each pixel in the image to be detected from its current coordinate value to the corresponding corrected coordinate value to obtain the orientation-corrected target face image. Figure 4 A schematic diagram of face orientation correction in one embodiment is shown.

[0094] In one embodiment, reference Figure 5 When the face image to be detected includes face 501 and face 502, the face correction model determines the orientation correction matrix corresponding to face 501 and the orientation correction matrix corresponding to face 502. The orientation of face 501 is corrected by the orientation correction matrix corresponding to face 501 to obtain the target face image 503 after orientation correction. The orientation of face 502 is corrected by the orientation correction matrix corresponding to face 502 to obtain the target face image 504 after orientation correction. Figure 5 A schematic diagram of face orientation correction in one embodiment is shown.

[0095] In the aforementioned face orientation correction method, face detection is performed on the image to be detected, yielding face localization points and candidate face boxes. By obtaining these candidate face boxes, the confidence level of each box can be determined, allowing for the selection of target detection boxes corresponding to each set of face localization points based on the confidence level. Furthermore, by obtaining the face localization points and face annotation points, an orientation correction matrix can be determined for each set of face localization points. Thus, face orientation correction can be performed based on the determined orientation correction matrix and the target detection boxes.

[0096] Furthermore, since at least one set of face localization points and the corresponding target detection boxes for each set of face localization points can be determined, face orientation correction can be performed simultaneously on multiple faces in the face image to be detected based on the determined at least one set of face localization points and the corresponding target detection boxes for each set of face localization points, thus greatly improving the efficiency of face orientation correction.

[0097] In one embodiment, face detection is performed on the face image to be detected to obtain at least one set of face localization points and at least one face candidate box corresponding to each set of face localization points, including: extracting face features from the face image to be detected at multiple scales to obtain face feature images at multiple scales; adjusting the number of channels in each face feature image at multiple scales to obtain multiple face feature images with the same number of channels; and determining at least one set of face localization points and at least one face candidate box corresponding to each set of face localization points in the face image to be detected based on the multiple face feature images with the same number of channels.

[0098] Specifically, when a face image to be detected is obtained, the face correction model can extract facial features at multiple scales, resulting in face feature images at multiple scales. For example, the face correction model can scale the face image to obtain an image pyramid, and then extract facial features for each image in the pyramid, resulting in face feature images at multiple scales. Alternatively, the face correction model can extract facial features from the face image to obtain face feature images at the bottom layer of the feature pyramid, and then upsample these bottom-layer images to obtain face feature images at the top layer of the feature pyramid.

[0099] Furthermore, to facilitate subsequent processing of facial feature images, the face correction model can adjust the number of channels in facial feature images at multiple scales to obtain multiple facial feature images with the same number of channels. Thus, the face correction model can input multiple facial feature images with the same number of channels into the face detection structure and the orientation detection structure to obtain at least one set of face localization points in the face image to be detected, and at least one face candidate box corresponding to each set of face localization points.

[0100] In one embodiment, reference Figure 6 , Figure 6 C3 to C5 represent face feature images at multiple scales, and Ci to Pi represent the adjustment of the number of channels for the i-th face feature image. Figure 6 A schematic diagram of the framework of a face correction model in one embodiment is shown.

[0101] In the above embodiments, by extracting facial feature images at multiple scales, subsequent recognition can be performed on multiple faces of different scales included in the face image to be detected, based on these multi-scale facial feature images. By adjusting the number of channels in the multi-scale facial feature images, it is easier to process facial feature images with a uniform number of channels, thereby improving the processing efficiency of facial feature images.

[0102] In one embodiment, the face correction method is executed through a face correction model, which includes a face detection structure and a orientation detection structure. Based on multiple face feature images with the same number of channels, at least one set of face localization points and at least one face candidate box corresponding to each set of face localization points are determined in the face image to be detected. This includes: using the face detection structure and based on multiple face feature images with the same number of channels, performing face region detection on the face image to be detected to obtain at least one face candidate box corresponding to each face region; and using the orientation detection structure and based on multiple face feature images with the same number of channels, performing face localization point detection on the face regions in the face image to be detected to obtain a set of face localization points corresponding to each face region.

[0103] Specifically, the face detection structure can use a preset face recognition algorithm and multiple face feature images with the same number of channels to detect face regions in the face image to be detected, obtaining at least one candidate face box for each face region. The face recognition algorithm can be freely configured according to requirements. For example, it can be a template-based algorithm, in which case the face detection model can compare multiple face feature images with the same number of channels with the learned face feature template, and determine the face region in the face image to be detected based on the comparison result, obtaining at least one candidate face box for each face region. Alternatively, the face recognition algorithm can be a facial feature point-based algorithm, in which the face detection structure can identify local structures such as eyes, nose, and mouth in the face image to be detected based on facial feature points in multiple face feature images with the same number of channels, and determine the face region in the face image to be detected based on the geometric relationships between these local structures.

[0104] Furthermore, the orientation detection structure can perform shape search on local structures such as eyes, nose, and mouth in the face image to be detected based on multiple face feature images with the same number of channels, and determine a set of face localization points corresponding to each face region based on the shape search results.

[0105] In this embodiment, by using different model structures to perform parallel detection of the face region and face localization points, the efficiency of face detection can be improved, thereby improving the efficiency of face orientation correction.

[0106] In one embodiment, based on the confidence level of each face candidate box, target detection boxes corresponding to each group of face localization points are selected from the face candidate boxes, including: for each group of face localization points, for at least one face candidate box, the face candidate box with the highest confidence level is selected from the at least one face candidate box corresponding to the current group of face localization points; the face candidate box with the highest confidence level is used as the target detection box corresponding to the current group of face localization points.

[0107] Specifically, for each group of face localization points in a multi-group framework, the face correction model determines at least one current face candidate box corresponding to the current group of face localization points, and determines the confidence level of each current face candidate box. The face correction model sorts the at least one current face candidate box according to its confidence level from high to low, obtaining a candidate box sequence. The face correction model uses the first current face candidate box in the candidate box sequence as the target detection box corresponding to the current group of face localization points.

[0108] In this embodiment, by using the face candidate box with the highest confidence as the target detection box, the accuracy of the target detection box can be improved, thereby improving the accuracy of face correction.

[0109] In one embodiment, face orientation correction is performed based on the target detection bounding boxes corresponding to each group of face localization points and according to the orientation correction matrix corresponding to the corresponding face localization points. This includes: for each group of face localization points in at least one group of face localization points, face orientation correction is performed on the image region in the current target detection bounding box corresponding to the current group of face localization points based on the current orientation correction matrix corresponding to the current group of face localization points.

[0110] Specifically, when the orientation correction matrix corresponding to each face localization point is determined, for each group of face localization points in at least one group of face localization points, the face correction model can perform face orientation correction on the image region in the current target detection box corresponding to the current group of face localization points based on the current orientation correction matrix corresponding to the current group of face localization points.

[0111] In one embodiment, when outputting face candidate boxes, the face detection structure can also output the coordinate values ​​of the face candidate boxes in the face image to be detected. For example, the face detection structure can establish a rectangular coordinate system with the center point of the image to be detected as the origin, and output the coordinate values ​​of the upper left corner and the lower right corner of the face candidate box in the coordinate system. Thus, the face correction model can crop out the image region selected by the face candidate box from the face image to be detected based on the coordinate values ​​of the face candidate box, and perform face orientation correction on the cropped image region based on the corresponding orientation correction matrix. For example, the cropped region can be rotated to obtain the face orientation corrected image.

[0112] In the above embodiments, by determining the orientation correction matrix corresponding to each group of face positioning points, the orientation correction of multiple faces in the face image to be detected can be performed based on the determined multiple orientation correction matrices, thereby greatly improving the efficiency of face orientation correction.

[0113] In one embodiment, face orientation correction is performed on the image region within the current target detection box corresponding to the current group of face localization points based on the current orientation correction matrix. This includes: generating a blank face image with the same size as the current target detection box; the blank face image is a blank image containing multiple first pixels; sequentially traversing the second pixels in the image region selected by the current target detection box; for the currently traversed second pixel, determining the target first pixel corresponding to the currently traversed second pixel from the multiple first pixels in the blank image based on the coordinates of the currently traversed second pixel and the current orientation correction matrix, and using the pixel information of the currently traversed second pixel as the pixel information of the corresponding target first pixel; after traversing all the second pixels and obtaining the pixel information of each first pixel, obtaining the correction result of face orientation correction for the image region within the current target detection box based on the pixel information of each first pixel.

[0114] Specifically, when the current target detection box is obtained, the face correction model determines the size of the current target detection box and generates a blank face image with the same size as the current target detection box based on the determined size. The blank face image includes multiple first pixels. The face correction model sequentially traverses the second pixels in the image region selected by the current target detection box. For the currently traversed second pixel, the face correction model determines the coordinates of the currently traversed second pixel in the image region selected by the current target detection box. It then multiplies the matrix formed by the coordinates of the currently traversed second pixel by the inverse of the current direction correction matrix to obtain the coordinates of the target first pixel in the blank image corresponding to the currently traversed second pixel, i.e., obtains the target first pixel corresponding to the currently traversed second pixel. For example, if the coordinates of the currently traversed second pixel are (xk, yk), and the inverse of the current direction correction matrix is ​​T... tol -1 At that time, the face correction model can multiply the matrix [xk,yk,1] formed by the coordinates of the second pixel by T. tol -1 We obtain [xk′, yk′, 1], where (xk′, yk′) are the coordinates of the first pixel of the target. The inverse of the orientation correction matrix refers to the matrix obtained by inverting the orientation correction matrix.

[0115] Furthermore, the face correction model acquires the pixel information of the currently traversed second pixel and uses this information as the pixel information of the corresponding target first pixel. Here, pixel information refers to information related to a pixel, which may specifically include the pixel's brightness and color values. After traversing all second pixels and obtaining the pixel information of each first pixel, the face correction model can, based on the pixel information of each first pixel, obtain the face orientation correction result for the image region within the current target detection box, that is, obtain the target face image with face orientation correction corresponding to the image region within the current target detection box.

[0116] In one embodiment, when the target face image corresponding to the current target detection box is obtained after face orientation correction, the face correction model can replace the image region selected by the current target detection box in the image to be detected with the target face image. The face correction model can also display the target face image at a preset position in the image to be detected.

[0117] In the above embodiments, the image after facial orientation correction can be generated simply by mapping the second pixel to the corresponding position of the blank face image, thus improving the efficiency of image generation.

[0118] In existing technologies, a divide-and-conquer strategy is generally used for facial orientation correction. This involves dividing the 360 ​​angles into several parts and correcting the facial orientation for each part. However, correcting the facial orientation for each part significantly increases the computational load on the neural network, thus reducing the efficiency of facial orientation correction. In contrast, the embodiments of this application do not require facial orientation correction for each part. Instead, only the orientation correction matrix needs to be determined, and facial orientation correction can be performed based on this matrix. This significantly improves the efficiency of facial orientation correction.

[0119] In one embodiment, the face orientation correction method is executed by a face correction model, which is trained through a model training step. The model training step includes: acquiring training samples and acquiring standard key points, standard face bounding boxes, and classification labels corresponding to the training samples; performing face detection on the training samples using the face correction model to be trained, obtaining at least one set of predicted key points, at least one predicted candidate box corresponding to each set of predicted key points, the predicted localization accuracy and standard localization accuracy corresponding to each predicted candidate box, and the predicted classification corresponding to each predicted candidate box; constructing a loss function based on the first difference between the predicted candidate box and the corresponding standard face bounding box, the second difference between the predicted key points and the corresponding standard key points, the third difference between the predicted localization accuracy and the corresponding standard localization accuracy, and the fourth difference between the predicted classification and the corresponding classification label; and training the face correction model based on the loss function until the training termination condition is met, thus obtaining a trained face correction model.

[0120] In this context, "standard keypoints" refer to standard face localization points, which serve as the criteria for evaluating predicted keypoints. Typically, these standard keypoints are manually marked on sample images. "Predicted keypoints" refer to the predicted face localization points output by the face correction model. "Standard face bounding boxes" refer to standard face detection boxes, which serve as the criteria for evaluating predicted candidate boxes. Typically, these standard face bounding boxes are manually selected by defining the face region on sample images. "Predicted candidate boxes" refer to the predicted face candidate boxes output by the face correction model.

[0121] Classification labels refer to the labels used to specify the category to which a corresponding image region belongs. For example, classification labels can be used to specify that the image region selected by the standard face bounding box is of the "face" category. Prediction and localization accuracy refers to the accuracy with which the face correction model predicts the face region selected by the predicted face candidate bounding box. Standard localization accuracy refers to the standard localization accuracy, which is the criterion for measuring prediction and localization accuracy.

[0122] Specifically, when training a face correction model, a large number of training samples can be obtained, and each training sample can be labeled to obtain its corresponding standard key points, standard face bounding boxes, and classification labels. The computer equipment can input the training samples into the face correction model to be trained. The model then performs face detection on the training samples, obtaining at least one set of predicted key points, at least one predicted candidate box corresponding to each set of face localization points, and the prediction accuracy of each candidate box. Further, the face correction model determines the standard face bounding box corresponding to each candidate box and determines the overlap between the candidate box and the corresponding standard face bounding box, determining the standard localization accuracy of the candidate box based on the overlap. Finally, the face correction model extracts image features from the image region selected by the candidate box and performs classification prediction on the image region selected by the candidate box based on the image features, obtaining the predicted classification.

[0123] Furthermore, the face correction model constructs a loss function based on the first difference between the predicted candidate bounding box and the corresponding standard face bounding box, the second difference between the predicted keypoints and the corresponding standard keypoints, the third difference between the predicted localization accuracy and the corresponding standard localization accuracy, and the fourth difference between the predicted classification and the corresponding classification label. The model parameters of the face correction model are adjusted using this constructed loss function until the training stopping condition is met, resulting in a well-trained face correction model. The training stopping condition can be freely set according to requirements; for example, reaching the required number of prediction iterations can be considered a stopping condition.

[0124] In one embodiment, the loss function can be determined using the following formula:

[0125] L total =a*L cls +b*L loc +c*L iou +d*L pose

[0126]

[0127] Where, N pos Let N be the number of positive training samples, M be the number of negative training samples, and L be the number of negative training samples. loc For the first difference, L pose For the second difference, L iou For the third difference, L cls The fourth difference is represented by FL, which is a binary classification function. For standard face frames, For predicting candidate boxes, (cx, cy) is the center point, w is the length, and h is the height, joint j For standard keypoints, joint′j To predict key points, iou i For standard positioning accuracy, iou′ i To predict positioning accuracy, p i For category labels, p′ i For predictive classification.

[0128] In the above embodiments, by constructing a loss function, the face correction model can be trained using the constructed loss function to obtain a well-trained face correction model, so that the face orientation of the face image to be detected can be corrected based on the well-trained face correction model.

[0129] In one embodiment, face detection is performed on training samples using a face correction model to be trained, resulting in at least one set of predicted keypoints, at least one predicted candidate box corresponding to each set of predicted keypoints, a prediction localization accuracy and an accurate localization accuracy corresponding to each predicted candidate box, and a prediction classification corresponding to each predicted candidate box. This includes: performing face region detection on the training samples using the face detection structure in the face correction model to be trained, obtaining at least one predicted candidate box corresponding to each face region; and performing face localization point detection on the face regions in the training samples using the orientation detection structure in the face correction model to be trained. The system performs several steps: First, it obtains a set of predicted keypoints for each face region. Second, it uses the localization detection structure in the face correction model to perform localization accuracy detection on each predicted candidate box, obtaining the predicted localization accuracy for each candidate box. Third, it uses the localization detection structure in the face correction model to determine the overlap between each predicted candidate box and the corresponding standard face box, and determines the standard localization accuracy for each candidate box based on the overlap. Finally, it uses the classification detection structure in the face correction model to determine the probability that the image region selected by each predicted candidate box is a face, obtaining the corresponding predicted classification.

[0130] Specifically, the face correction model includes a face detection structure, a orientation detection structure, a localization detection structure, and a classification detection structure. When training samples are obtained, the face correction model extracts facial features from the training samples and inputs these features into the face detection structure to be trained. The face detection structure then performs face region detection on the training samples, obtaining at least one predicted candidate box for each face region. The face correction model can also input the extracted facial features into the orientation detection structure to be trained. The orientation detection structure then detects facial localization points in the training samples, obtaining a set of predicted keypoints for each face region.

[0131] Furthermore, the face correction model inputs facial features and the location information of predicted candidate boxes into the localization detection structure to be trained. The localization detection structure then performs localization accuracy detection on each predicted candidate box, obtaining the predicted localization accuracy for each candidate box. The localization detection structure uses predicted candidate boxes and standard face detection boxes containing the same set of predicted keypoints as corresponding detection boxes. Based on the location information of the predicted candidate boxes and the standard face boxes, it determines the overlap between each predicted candidate box and the corresponding standard face box, and uses this overlap to determine the standard localization accuracy of the corresponding predicted candidate box. For example, the localization detection structure can determine the intersection-union ratio (IUR) between the predicted candidate box and the standard face box, and use the IUR as the standard localization accuracy.

[0132] Furthermore, the face correction model can input the extracted facial features into the classification and detection structure to be trained. The classification and detection structure can then classify and predict the image region selected by the predicted candidate box based on the facial features, thereby obtaining the probability value that the image region selected by the predicted candidate box is a face, and using this probability value as the predicted classification.

[0133] In this embodiment, by inputting facial features into the face detection structure, orientation detection structure, localization detection structure, and classification detection structure respectively, the facial features can be processed in parallel through the face detection structure, orientation detection structure, localization detection structure, and classification detection structure, thereby improving training efficiency.

[0134] In one embodiment, the above-mentioned face orientation correction method further includes a training sample generation step, which includes: acquiring multiple initial face images, and rotating each of the multiple initial face images by a random angle to obtain a rotated face image; adjusting the scale and pixel value of each rotated face image to obtain an adjusted face image; and stitching the adjusted face images together to obtain training samples.

[0135] Specifically, to increase the number of training samples, when acquiring multiple initial face images, the computer device can rotate each face image at a random angle to obtain rotated face images. For each rotated face image, the computer device can adjust its scale and pixel values, setting the scale to a preset scale and adjusting the pixel values ​​of each pixel by a preset amount to obtain adjusted face images. Further, the computer device stitches together a preset number of adjusted face images to obtain training sample images. For example, referencing... Figure 7 The computer equipment stitches together every 5 adjusted facial images to obtain training sample images. Figure 7 A schematic diagram of training samples in one embodiment is shown.

[0136] In this embodiment, by adjusting and stitching the initial face images, not only is the diversity of the training samples improved, but the number of small-sized faces is also increased, thereby enhancing the richness of the training samples.

[0137] In one embodiment, the above-mentioned face orientation correction method further includes, when the face image to be detected is a face image captured during a video call, adjusting the image displayed in the video call based on the target face image obtained after face orientation correction; when the face image to be detected is a face image captured during video editing, replacing the face image in the video to be edited with the target face image obtained after face orientation correction; and when the face image to be detected is a face image captured during video playback, adding face orientation correction effects to the played video based on the target face image obtained after face orientation correction.

[0138] Specifically, the aforementioned face orientation correction method can be applied to video call scenarios, video editing scenarios, and video playback scenarios. When a user is making a video call through a terminal, the terminal can acquire the video stream generated during the video call. This video stream can be generated by an image acquisition device capturing the user's face, or it can be a video stream sent by the terminal of the user in the video call. The terminal can perform frame segmentation processing on the video stream to obtain multiple video frames to be processed. These frames are then input into a trained face correction model. The face correction model corrects the face orientation of the video frames to be processed, resulting in a target face image after face orientation correction. Furthermore, when the target face image after face orientation correction is obtained, the terminal can replace the corresponding video frame in the video stream with the corrected target face image and display the video stream after the replacement. Thus, during the video call, the user can view a video stream containing a positively oriented face, thereby improving the user experience.

[0139] When the aforementioned face correction method is applied to video editing scenarios, users can input the video to be edited through a video editing application. The application can then perform frame-by-frame processing on the video to identify detectable images containing faces. These detectable images are then input into a trained face correction model, which corrects the facial orientation of the detected images to obtain the target face image. Furthermore, the video editing application can replace corresponding video frames in the video to be edited with the target face image to obtain the edited video.

[0140] Accordingly, when the aforementioned face correction method is applied to a video editing scenario, the user can play a video through a video playback application. The video playback application can then perform frame-by-frame processing on the video to obtain the image to be detected, and use the face correction model to correct the face orientation of the image to obtain the target face image after face orientation correction. Furthermore, the video playback application can add face orientation correction effects to the played video based on the target face image after face orientation correction, for example, by referencing... Figure 8 The video playback application can display the target face image 801 at a preset position in the corresponding video screen. Figure 8 A schematic diagram illustrating a target face image is shown in one embodiment.

[0141] In this embodiment, by obtaining the target face image after face orientation correction, the screen displayed in the video call can be adjusted based on the target face image, the face image in the video to be edited can be replaced, and face orientation correction effects can be added to the video screen, thereby greatly improving the user experience.

[0142] In one embodiment, such as Figure 9 The face orientation correction method provided in this application includes the following steps:

[0143] S902, acquire the face image to be detected, extract face features at multiple scales from the face image to be detected, and obtain face feature images at multiple scales.

[0144] S904 adjusts the number of channels in each face feature image at multiple scales to obtain multiple face feature images with the same number of channels.

[0145] S906, through the face detection structure in the face correction model, and based on multiple face feature images with the same number of channels, performs face region detection on the face image to be detected, and obtains at least one face candidate box corresponding to each face region.

[0146] S908 uses the orientation detection structure in the face correction model and multiple face feature images with the same number of channels to perform face localization point detection on the face region in the face image to be detected, and obtains a set of face localization points corresponding to each face region.

[0147] S910, for each face candidate box in at least one face candidate box, the localization accuracy of the current face candidate box is predicted by the localization detection structure in the face correction model. The localization accuracy characterizes the degree of difference between the face candidate box and the corresponding standard face box.

[0148] S912, for each face candidate box in at least one face candidate box, the classification probability value of the image region selected by the current face candidate box is predicted as a face through the classification detection structure in the face correction model.

[0149] S914, for each face candidate box in at least one face candidate box, the confidence level of the current face candidate box is determined based on the localization accuracy and classification probability value corresponding to the current face candidate box.

[0150] S916, for each group of face localization points corresponding to at least one face candidate box, select the face candidate box with the highest confidence from the at least one face candidate box corresponding to the current group of face localization points, and use the face candidate box with the highest confidence as the target detection box corresponding to the current group of face localization points.

[0151] S918, Obtain face annotation points. Face annotation points are face positioning points in a face image under standard pose. Determine the first coordinate matrix based on the position information of the face annotation points.

[0152] S920, for each set of face positioning points in at least one set of face positioning points, the current second coordinate matrix is ​​determined based on the position information of the current face positioning point.

[0153] S922, perform an affine transformation on the first coordinate matrix and the current second coordinate matrix to obtain the orientation correction matrix corresponding to the current group of face positioning points.

[0154] S924, Generate a blank face image with the same size as the current target detection box; the blank face image is a blank image containing multiple first pixels and sequentially traverses the second pixels in the image region selected by the current target detection box.

[0155] S926, for the currently traversed second pixel, based on the coordinates of the currently traversed second pixel and the current direction correction matrix, determine the target first pixel corresponding to the currently traversed second pixel from multiple first pixels in the blank image, and use the pixel information of the currently traversed second pixel as the pixel information of the corresponding target first pixel.

[0156] S928: After traversing all the second pixels and obtaining the pixel information of each first pixel, the correction result of face orientation correction of the image region in the current target detection box is obtained based on the pixel information of each first pixel.

[0157] In the aforementioned face orientation correction method, face detection is performed on the image to be detected, yielding face localization points and candidate face boxes. By obtaining these candidate face boxes, the confidence level of each box can be determined, allowing for the selection of target detection boxes corresponding to each set of face localization points based on the confidence level. Furthermore, by obtaining the face localization points and face annotation points, an orientation correction matrix can be determined for each set of face localization points. Thus, face orientation correction can be performed based on the determined orientation correction matrix and the target detection boxes.

[0158] Furthermore, since at least one set of face localization points and the corresponding target detection boxes for each set of face localization points can be determined, face orientation correction can be performed simultaneously on multiple faces in the face image to be detected based on the determined at least one set of face localization points and the corresponding target detection boxes for each set of face localization points, thereby greatly improving the efficiency of face orientation correction.

[0159] In one embodiment, such as Figure 10 The face orientation correction method provided in this application includes the following steps:

[0160] S1002, acquire multiple initial face images, and rotate each of the multiple initial face images by a random angle to obtain a rotated face image.

[0161] S1004: Adjust the scale and pixel values ​​of each rotated face image to obtain an adjusted face image. Then, stitch the adjusted face images together to obtain training samples.

[0162] S1006, obtain the standard key points, standard face bounding boxes and classification labels corresponding to the training samples.

[0163] S1008, using the face detection structure in the face correction model to be trained, perform face region detection on the training samples to obtain at least one prediction candidate box corresponding to each face region.

[0164] S1010 uses the orientation detection structure in the face correction model to perform face localization point detection on the face region in the training sample, and obtains a set of predicted key points corresponding to each face region.

[0165] S1012, using the localization detection structure in the face correction model to be trained, the localization accuracy of each predicted candidate box is detected, and the prediction localization accuracy corresponding to each predicted candidate box is obtained.

[0166] S1014, using the localization detection structure in the face correction model to be trained, determine the overlap between each predicted candidate box and the corresponding standard face box, and determine the standard localization accuracy corresponding to each predicted candidate box based on the overlap.

[0167] S1016: Using the classification and detection structure in the face correction model to be trained, determine the probability value of the image region selected by each prediction candidate box as a face, and obtain the corresponding prediction classification.

[0168] S1018. Based on the first difference between the predicted candidate box and the corresponding standard face box, the second difference between the predicted key points and the corresponding standard key points, the third difference between the predicted localization accuracy and the corresponding standard localization accuracy, and the fourth difference between the predicted classification and the corresponding classification label, a loss function is constructed.

[0169] S1020: The face correction model is trained based on the loss function until the training termination condition is met, and the trained face correction model is obtained.

[0170] This application also provides an application scenario in which the above-described face orientation correction method is applied. Specifically, the face orientation correction method is applied in this scenario as follows:

[0171] When a face image to be detected is acquired, the computer device can correct the face orientation of the face image using the aforementioned face orientation correction method to obtain a face orientation-corrected target face image. This corrected target face image is then input to a server for processing downstream target tasks, enabling the server to perform downstream task processing based on the target face image. For example, the server can determine the key facial contour points in the corresponding face image to be detected based on the target face image, and crop the face from the corresponding face image to be detected.

[0172] The above application scenarios are merely illustrative. It is understood that the application of the business-related data reporting methods provided in the embodiments of this application is not limited to the above scenarios.

[0173] It should be understood that, although Figure 1 , Figures 9-10 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 , Figures 9-10At least some of the steps in the process may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but may be executed at different times. The execution order of these steps or stages is not necessarily sequential, but may be executed in turn or alternately with other steps or at least some of the steps or stages in other steps.

[0174] In one embodiment, such as Figure 11 As shown, a face orientation correction device 1100 is provided. This device can be a software module, a hardware module, or a combination of both as part of a computer device. Specifically, the device includes: a face detection module 1102, a matrix determination module 1104, and an orientation correction module 1106, wherein:

[0175] The face detection module 1102 is used to acquire a face image to be detected, perform face detection on the face image to be detected, and obtain at least one set of face localization points and at least one face candidate box corresponding to each set of face localization points.

[0176] The matrix determination module 1104 is used to determine the localization accuracy and classification probability value of each face candidate box, and to determine the confidence level of each face detection box based on the localization accuracy and classification probability value; based on the confidence level of each face candidate box, to select the target detection boxes corresponding to each group of face localization points from the face candidate boxes; to obtain face annotation points; the face annotation points are the face localization points in the face image under standard pose; and to determine the orientation correction matrix corresponding to each group of face localization points based on the face annotation points and each group of face localization points.

[0177] The orientation correction module 1106 is used to perform face orientation correction based on the target detection boxes corresponding to each group of face localization points and according to the orientation correction matrix corresponding to the corresponding face localization points.

[0178] In one embodiment, reference Figure 12 The face detection module 1102 also includes a feature extraction module 1121, which is used to extract face features from the face image to be detected at multiple scales to obtain face feature images at multiple scales; adjust the number of channels in each face feature image at multiple scales to obtain multiple face feature images with the same number of channels; and determine at least one set of face localization points in the face image to be detected, and at least one face candidate box corresponding to each set of face localization points, based on the multiple face feature images with the same number of channels.

[0179] In one embodiment, the face correction method is executed through a face correction model, which includes a face detection structure and a orientation detection structure. The face detection module 1102 is further configured to perform face region detection on the face image to be detected using the face detection structure and based on multiple face feature images with the same number of channels, to obtain at least one face candidate box corresponding to each face region. The orientation detection module is also configured to perform face localization point detection on the face region in the face image to be detected using the orientation detection structure and based on multiple face feature images with the same number of channels, to obtain a set of face localization points corresponding to each face region.

[0180] In one embodiment, the face correction method is executed through a face correction model, which includes a localization detection structure and a classification detection structure. The matrix determination module 1104 further includes a target detection box determination module 1141, which is used to predict the localization accuracy of the current face candidate box for each face candidate box in at least one face candidate box through the localization detection structure. The localization accuracy characterizes the degree of difference between the face candidate box and the corresponding standard face box. For each face candidate box in at least one face candidate box, the classification probability value of the image region selected by the current face candidate box is predicted to be a face through the classification detection structure. For each face candidate box in at least one face candidate box, the confidence of the current face candidate box is determined based on the localization accuracy and classification probability value corresponding to the current face candidate box.

[0181] In one embodiment, the target detection box determination module 1141 is further configured to, for each group of face localization points, select the face candidate box with the highest confidence from the at least one face candidate box corresponding to the current group of face localization points; and use the face candidate box with the highest confidence as the target detection box corresponding to the current group of face localization points.

[0182] In one embodiment, the matrix determination module 1104 is further configured to determine the orientation correction matrix corresponding to each group of face positioning points based on the face annotation points and each group of face positioning points, including: determining a first coordinate matrix based on the position information of the face annotation points; determining a current second coordinate matrix for each group of face positioning points based on the position information of the current face positioning point; and performing an affine transformation on the first coordinate matrix and the current second coordinate matrix to obtain the orientation correction matrix corresponding to the current group of face positioning points.

[0183] In one embodiment, the orientation correction module 1106 is further configured to perform face orientation correction on the image region in the current target detection box corresponding to the current group of face positioning points for each group of face positioning points in at least one group of face positioning points, based on the current orientation correction matrix corresponding to the current group of face positioning points.

[0184] In one embodiment, the orientation correction module 1106 is further configured to generate a blank face image with the same size as the current target detection box; the blank face image is a blank image containing multiple first pixels; the second pixels in the image region selected by the current target detection box are traversed sequentially; for the currently traversed second pixel, based on the coordinates of the currently traversed second pixel and the current orientation correction matrix, the target first pixel corresponding to the currently traversed second pixel is determined from the multiple first pixels in the blank image, and the pixel information of the currently traversed second pixel is used as the pixel information of the corresponding target first pixel; after traversing all the second pixels and obtaining the pixel information of each first pixel, the correction result of face orientation correction of the image region in the current target detection box is obtained based on the pixel information of each first pixel.

[0185] In one embodiment, the face orientation correction device 1100 further includes a training module 1108, used to acquire training samples and acquire standard key points, standard face bounding boxes, and classification labels corresponding to the training samples; perform face detection on the training samples using the face correction model to be trained, to obtain at least one set of predicted key points, at least one predicted candidate box corresponding to each set of predicted key points, the predicted localization accuracy and standard localization accuracy corresponding to each predicted candidate box, and the predicted classification corresponding to each predicted candidate box; construct a loss function based on the first difference between the predicted candidate box and the corresponding standard face bounding box, the second difference between the predicted key point and the corresponding standard key point, the third difference between the predicted localization accuracy and the corresponding standard localization accuracy, and the fourth difference between the predicted classification and the corresponding classification label; train the face correction model based on the loss function until the training termination condition is met, to obtain the trained face correction model.

[0186] In one embodiment, the training module 1108 is further configured to: perform face region detection on the training samples using the face detection structure in the face correction model to be trained, to obtain at least one predicted candidate box corresponding to each face region; perform face localization point detection on the face regions in the training samples using the orientation detection structure in the face correction model to be trained, to obtain a set of predicted key points corresponding to each face region; perform localization accuracy detection on each predicted candidate box using the localization detection structure in the face correction model to be trained, to obtain the predicted localization accuracy corresponding to each predicted candidate box; determine the overlap between each predicted candidate box and the corresponding standard face box using the localization detection structure in the face correction model to be trained, and determine the standard localization accuracy corresponding to each predicted candidate box based on the overlap; and determine the probability value that the image region selected by each predicted candidate box is a face using the classification detection structure in the face correction model to be trained, to obtain the corresponding predicted classification.

[0187] In one embodiment, the face orientation correction device 1100 is further configured to acquire multiple initial face images, and rotate each initial face image in the multiple initial face images by a random angle to obtain a rotated face image; adjust the scale and pixel value of each rotated face image to obtain an adjusted face image; and stitch the adjusted face images together to obtain training samples.

[0188] In one embodiment, the face orientation correction device 1100 is further configured to: when the face image to be detected is a face image captured during a video call, adjust the image displayed in the video call based on the target face image obtained after face orientation correction; when the face image to be detected is a face image captured during video editing, replace the face image in the video to be edited with the target face image obtained after face orientation correction; and when the face image to be detected is a face image captured during video playback, add face orientation correction effects to the played video based on the target face image obtained after face orientation correction.

[0189] Specific limitations regarding the face orientation correction device can be found in the limitations of the face orientation correction method described above, and will not be repeated here. Each module in the aforementioned face orientation correction device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0190] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 13 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computational and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores facial orientation correction data. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a facial orientation correction method.

[0191] Those skilled in the art will understand that Figure 13The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0192] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0193] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0194] In one embodiment, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and executes the computer instructions, causing the computer device to perform the steps in the above method embodiments.

[0195] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0196] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0197] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for correcting facial orientation, characterized in that, The method includes: The process involves: acquiring a face image to be detected from a video call stream, a video to be edited, or a video to be played; extracting face features from the face image at multiple scales to obtain face feature images at multiple scales; adjusting the number of channels in each face feature image at multiple scales to obtain multiple face feature images with the same number of channels; using the face detection structure in the face correction model and based on the multiple face feature images in the face image to be detected, performing face region detection on the face image to be detected to obtain at least two candidate boxes for each face region; using the direction detection structure in the face correction model and based on the multiple face feature images, performing face localization point detection on the face regions in the face image to be detected to obtain a set of face localization points corresponding to each face region, wherein the number of sets is the same as the number of faces in the face image to be detected, and the angle between different face directions in the face image to be detected and a preset direction is different; Using the localization detection structure and classification detection structure in the face correction model, the localization accuracy and classification probability value of each face candidate box are determined respectively, and the confidence level of each face candidate box is determined according to the localization accuracy and the classification probability value. Based on the confidence level of each face candidate box, target detection boxes corresponding to each group of face localization points are selected from the face candidate boxes. Obtain face annotation points; the face annotation points are pre-annotated face positioning points in a face image under standard pose. Based on the face annotation points and each group of face positioning points, determine the orientation correction matrix corresponding to each group of face positioning points; Based on the target detection boxes corresponding to each group of face positioning points, and according to the orientation correction matrix corresponding to the corresponding face positioning points, face orientation correction is performed to obtain the target face image after face orientation correction; the target face image is used to adjust the screen displayed in video calls, replace the face image in the video to be edited, and add face orientation correction effects to the video screen.

2. The method according to claim 1, characterized in that, The facial features in the facial feature image include facial texture features, which include the color value distribution and brightness value distribution of the facial image pixels.

3. The method according to claim 1, characterized in that, The step of determining the confidence level of each face candidate box based on the positioning accuracy and the classification probability value includes: The face correction model is used to perform a weighted summation of the positioning accuracy and the classification probability value to obtain the confidence level of each face candidate box.

4. The method according to claim 1, characterized in that, The step involves determining the localization accuracy and classification probability value of each face candidate box using the localization detection structure and classification detection structure in the face correction model, and determining the confidence level of each face candidate box based on the localization accuracy and classification probability value, including: For each face candidate box in at least one face candidate box, the localization accuracy of the current face candidate box is predicted by the localization detection structure; the localization accuracy characterizes the degree of difference between the face candidate box and the corresponding standard face box. For each face candidate box in at least one face candidate box, the classification detection structure is used to predict the classification probability value of the image region selected by the current face candidate box as a face. For each face candidate box in at least one face candidate box, the confidence level of the current face candidate box is determined based on the localization accuracy and classification probability value corresponding to the current face candidate box.

5. The method according to claim 1, characterized in that, The step of filtering target detection boxes corresponding to each group of face localization points from the face candidate boxes based on the confidence level of each candidate box includes: For each group of face localization points, at least one face candidate box is selected from the at least one face candidate box corresponding to the current group of face localization points, and the face candidate box with the highest confidence is selected. The candidate face bounding box with the highest confidence is used as the target detection box corresponding to the face localization point of the current group.

6. The method according to claim 1, characterized in that, The step of determining the orientation correction matrix corresponding to each group of face positioning points based on the face annotation points and each group of face positioning points includes: The first coordinate matrix is ​​determined based on the location information of the face annotation points; For each set of face localization points in at least one set, the current second coordinate matrix is ​​determined based on the current position information of the face localization point; Perform an affine transformation on the first coordinate matrix and the current second coordinate matrix to obtain the orientation correction matrix corresponding to the current group of face positioning points.

7. The method according to claim 1, characterized in that, The step of performing face orientation correction based on the target detection bounding boxes corresponding to each group of face localization points and according to the orientation correction matrix corresponding to the respective face localization points includes: For each group of face localization points in at least one group, face orientation correction is performed on the image region in the current target detection box corresponding to the current group of face localization points based on the current orientation correction matrix corresponding to the current group of face localization points.

8. The method according to claim 7, characterized in that, The step of performing face orientation correction on the image region within the current target detection box corresponding to the current group of face localization points, based on the current orientation correction matrix corresponding to the current group of face localization points, includes: Generate a blank face image with the same size as the current target detection box; the blank face image is a blank image containing multiple first pixels; The second pixel points in the image region selected by the current target detection box are sequentially traversed; For the currently traversed second pixel, based on the coordinates of the currently traversed second pixel and the current direction correction matrix, the target first pixel corresponding to the currently traversed second pixel is determined from the multiple first pixels of the blank image, and the pixel information of the currently traversed second pixel is used as the pixel information of the corresponding target first pixel. After traversing all the second pixels and obtaining the pixel information of each first pixel, the correction result of face orientation correction for the image region in the current target detection box is obtained based on the pixel information of each first pixel.

9. The method according to claim 1, characterized in that, The face orientation correction method is executed by a face correction model, which is obtained through a model training step, which includes: Obtain training samples, and obtain standard key points, standard face bounding boxes, and classification labels corresponding to the training samples; The training samples are subjected to face detection by the face correction model to be trained, and at least one set of predicted key points, at least one predicted candidate box corresponding to each set of predicted key points, the predicted localization accuracy and standard localization accuracy corresponding to each predicted candidate box, and the predicted classification corresponding to each predicted candidate box. A loss function is constructed based on the first difference between the predicted candidate box and the corresponding standard face box, the second difference between the predicted key point and the corresponding standard key point, the third difference between the predicted localization accuracy and the corresponding standard localization accuracy, and the fourth difference between the predicted classification and the corresponding classification label. The face correction model is trained based on the loss function until the training termination condition is met, thus obtaining a well-trained face correction model.

10. The method according to claim 9, characterized in that, The step of performing face detection on the training samples using the face correction model to be trained, obtaining at least one set of predicted key points, at least one predicted candidate box corresponding to each set of predicted key points, the predicted localization accuracy and standard localization accuracy corresponding to each predicted candidate box, and the predicted classification corresponding to each predicted candidate box, includes: By using the face detection structure in the face correction model to be trained, face region detection is performed on the training samples to obtain at least one prediction candidate box corresponding to each face region. By using the orientation detection structure in the face correction model to be trained, face localization point detection is performed on the face regions in the training samples to obtain a set of predicted key points corresponding to each face region. Using the localization detection structure in the face correction model to be trained, the localization accuracy of each predicted candidate box is detected, and the prediction localization accuracy corresponding to each predicted candidate box is obtained. The overlap between each predicted candidate box and the corresponding standard face box is determined by the localization detection structure in the face correction model to be trained, and the standard localization accuracy corresponding to each predicted candidate box is determined according to the overlap. By using the classification and detection structure in the face correction model to be trained, the probability value of the image region selected by each prediction candidate box is determined as a face, and the corresponding prediction classification is obtained.

11. The method according to claim 9, characterized in that, The method further includes: Multiple initial face images are acquired, and each of the multiple initial face images is rotated at a random angle to obtain a rotated face image. Each rotated face image is adjusted in scale and pixel value to obtain an adjusted face image; The adjusted face images are stitched together to obtain training samples.

12. The method according to any one of claims 1 to 11, characterized in that, The method further includes: When the face image to be detected is a face image captured during a video call, the screen displayed in the video call is adjusted based on the target face image obtained after face orientation correction. When the face image to be detected is a face image captured during video editing, the face image in the video to be edited will be replaced by the target face image obtained after face orientation correction. When the face image to be detected is a face image captured during video playback, a face orientation correction effect is added to the video frame based on the target face image obtained after face orientation correction.

13. A facial orientation correction device, characterized in that, The device includes: A face detection module is used to acquire face images to be detected from call video streams, videos to be edited, or videos to be played; extract face features from the face images to be detected at multiple scales to obtain face feature images at multiple scales; adjust the number of channels in each face feature image at multiple scales to obtain multiple face feature images with the same number of channels; use the face detection structure in the face correction model and based on the multiple face feature images in the face images to be detected to perform face region detection on the face images to be detected to obtain at least two face candidate boxes corresponding to each face region; use the direction detection structure in the face correction model and based on the multiple face feature images to perform face localization point detection on the face regions in the face images to be detected to obtain a set of face localization points corresponding to each face region, wherein the number of sets is the same as the number of faces in the face images to be detected, and the angle between different face directions in the face images to be detected and a preset direction is different; The matrix determination module is used to determine the localization accuracy and classification probability value of each face candidate box through the localization detection structure and classification detection structure in the face correction model, and to determine the confidence level of each face candidate box based on the localization accuracy and classification probability value; based on the confidence level of each face candidate box, to filter out target detection boxes corresponding to each group of face localization points from the face candidate boxes; to obtain face annotation points; the face annotation points are pre-annotated face localization points in a face image under standard pose; and to determine the orientation correction matrix corresponding to each group of face localization points based on the face annotation points and each group of face localization points. The orientation correction module is used to perform face orientation correction based on the target detection boxes corresponding to each group of face positioning points and the orientation correction matrix corresponding to the corresponding face positioning points to obtain the target face image after face orientation correction; the target face image is used to adjust the screen displayed in video calls, replace the face image in the video to be edited, and add face orientation correction effects to the video screen.

14. The apparatus according to claim 13, characterized in that, The facial features in a facial feature image include facial texture features, which include the color value distribution and brightness value distribution of the facial image pixels.

15. The apparatus according to claim 13, characterized in that, The matrix determination module is further configured to perform a weighted summation of the positioning accuracy and the classification probability value using the face correction model to obtain the confidence level corresponding to each face candidate box.

16. The apparatus according to claim 13, characterized in that, The face orientation correction device is executed through a face correction model, which includes a positioning detection structure and a classification detection structure. The matrix determination module is further configured to, for each face candidate box in at least one face candidate box, predict the localization accuracy of the current face candidate box using the localization detection structure; the localization accuracy characterizes the degree of difference between the face candidate box and the corresponding standard face box; for each face candidate box in at least one face candidate box, predict the classification probability value of the image region selected by the current face candidate box as a face using the classification detection structure; and for each face candidate box in at least one face candidate box, determine the confidence level of the current face candidate box based on the localization accuracy and classification probability value corresponding to the current face candidate box.

17. The apparatus according to claim 13, characterized in that, The matrix determination module is further configured to, for each group of face localization points, select the face candidate box with the highest confidence from the at least one face candidate box corresponding to the current group of face localization points; and use the face candidate box with the highest confidence as the target detection box corresponding to the current group of face localization points.

18. The apparatus according to claim 13, characterized in that, The matrix determination module is further configured to determine a first coordinate matrix based on the position information of the face annotation points; for each group of face positioning points in at least one group of face positioning points, a current second coordinate matrix is ​​determined based on the position information of the current face positioning point; and an affine transformation is performed on the first coordinate matrix and the current second coordinate matrix to obtain the orientation correction matrix corresponding to the current group of face positioning points.

19. The apparatus according to claim 13, characterized in that, The orientation correction module is further configured to, for each group of face positioning points in at least one group of face positioning points, perform face orientation correction on the image region in the current target detection box corresponding to the current group of face positioning points based on the current orientation correction matrix corresponding to the current group of face positioning points.

20. The apparatus according to claim 19, characterized in that, The orientation correction module is further configured to generate a blank face image with the same size as the current target detection box; the blank face image is a blank image containing multiple first pixels; the second pixels in the image region selected by the current target detection box are sequentially traversed; for the currently traversed second pixel, based on the coordinates of the currently traversed second pixel and the current orientation correction matrix, the target first pixel corresponding to the currently traversed second pixel is determined from the multiple first pixels in the blank image, and the pixel information of the currently traversed second pixel is used as the pixel information of the corresponding target first pixel; After traversing all the second pixels and obtaining the pixel information of each first pixel, the correction result of face orientation correction for the image region in the current target detection box is obtained based on the pixel information of each first pixel.

21. The apparatus according to claim 13, characterized in that, The face orientation correction device is executed by a face correction model, which is trained through a model training step. The device also includes: The training module is used to acquire training samples and obtain standard key points, standard face bounding boxes, and classification labels corresponding to the training samples; perform face detection on the training samples using the face correction model to be trained, and obtain at least one set of predicted key points, at least one predicted candidate box corresponding to each set of predicted key points, a predicted localization accuracy and a standard localization accuracy corresponding to each predicted candidate box, and a predicted classification corresponding to each predicted candidate box; construct a loss function based on the first difference between the predicted candidate box and the corresponding standard face bounding box, the second difference between the predicted key point and the corresponding standard key point, the third difference between the predicted localization accuracy and the corresponding standard localization accuracy, and the fourth difference between the predicted classification and the corresponding classification label; train the face correction model based on the loss function until the training termination condition is met, and obtain the trained face correction model.

22. The apparatus according to claim 21, characterized in that, The training module is also used to perform face region detection on the training samples using the face detection structure in the face correction model to be trained, and obtain at least one prediction candidate box corresponding to each face region; and to perform face localization point detection on the face regions in the training samples using the orientation detection structure in the face correction model to be trained, and obtain a set of prediction key points corresponding to each face region. Using the localization detection structure in the face correction model to be trained, the localization accuracy of each predicted candidate box is detected to obtain the predicted localization accuracy corresponding to each predicted candidate box; using the localization detection structure in the face correction model to be trained, the overlap between each predicted candidate box and the corresponding standard face box is determined, and the standard localization accuracy corresponding to each predicted candidate box is determined based on the overlap; using the classification detection structure in the face correction model to be trained, the probability value of the image region selected by each predicted candidate box being a face is determined to obtain the corresponding predicted classification.

23. The apparatus according to claim 21, characterized in that, The face orientation correction device is also used to acquire multiple initial face images, and rotate each of the multiple initial face images by a random angle to obtain a rotated face image; adjust the scale and pixel value of each rotated face image to obtain an adjusted face image; and stitch the adjusted face images together to obtain training samples.

24. The apparatus according to any one of claims 13 to 23, characterized in that, The face orientation correction device is further configured to: when the face image to be detected is a face image captured during a video call, adjust the image displayed in the video call based on the target face image obtained after face orientation correction; when the face image to be detected is a face image captured during video editing, replace the face image in the video to be edited with the target face image obtained after face orientation correction; and when the face image to be detected is a face image captured during video playback, add face orientation correction effects to the played video based on the target face image obtained after face orientation correction.

25. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 12.

26. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 12.

27. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 12.