Face correction method, device, equipment and storage medium
Patent Information
- Application Number
- CN202610778950.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-01
- Publication Date
- 2026-09-04
AI Technical Summary
当此类非标准朝向的图像被输入到现有检测模型时,模型预先学习的特征空间结构被完全破坏,导致无法提取有效的人脸特征,进而造成检测任务的彻底失败
[0016]The beneficial effects of this invention are as follows: By normalizing the pose of facial images acquired in smart home scenarios and correcting facial orientation, the interference of head tilt on downstream algorithms is effectively eliminated. This significantly improves the accuracy and robustness of tasks such as face recognition, key point localization, and attribute analysis, providing stable and high-quality standardized input data for various visual processing systems, thereby improving the accuracy of face recognition in smart home appliances such as smart door locks. Simultaneously, by binding rotation information to the corrected image, the integrity of the image data and the reversibility of processing are ensured. Furthermore, this rotation information can serve as high-value quantified metadata, which can be directly used by downstream applications, avoiding redundant calculations and improving the overall operating efficiency of the system.
Smart Images

Figure CN122694718A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, specifically to a face correction method, device, equipment, and storage medium. Background Technology
[0002] With the rapid development of IoT technology, the smart home ecosystem has become increasingly widespread. Against this backdrop, facial recognition technology serves as a key support, providing the foundation for core functions such as identity verification for smart locks, security monitoring for smart cameras, and user interaction with smart home appliances. This technology enhances the convenience and security of modern home life through automation and intelligence, demonstrating broad application prospects and enormous market value.
[0003] Existing face detection technologies primarily rely on deep learning models, especially convolutional neural networks. These models are trained on massive datasets of standard frontal face images to learn and master the inherent spatial feature hierarchy of the face, such as the relative spatial layout of the eyes, nose, and mouth. Therefore, the detection performance of these models is highly dependent on the standard, vertical orientation of the face in the input image. Under ideal or controlled conditions, when the face orientation conforms to the model's preset parameters, this technology can achieve efficient and accurate detection, and has become the mainstream solution in the industry.
[0004] However, the above solutions have certain shortcomings in the actual deployment and application of smart homes. First, the installation positions and angles of smart terminal devices (such as smart door locks and wall-mounted cameras) are diverse, often resulting in fixed tilts or rotations of the cameras themselves. Second, users' interaction postures are not static. These real-world factors cause the captured facial images to frequently exhibit discrete rotation states such as 90 degrees, 180 degrees, or 270 degrees. When such non-standard orientation images are input into existing detection models, the pre-learned feature space structure of the model is completely destroyed, making it impossible to extract effective facial features, thus causing the detection task to fail completely. This defect directly leads to serious problems such as smart door locks failing to recognize legitimate users and security cameras missing critical events, greatly damaging product reliability and user experience, and constituting a core technical bottleneck restricting the development of facial recognition applications in smart homes. Therefore, there is an urgent need in this field for a solution that can overcome the influence of image rotation and achieve omnidirectional face detection. Summary of the Invention
[0005] This application provides a face correction method, apparatus, device, and storage medium for correcting the orientation of face images collected in smart home scenarios, so that the face images can meet the image requirements of face recognition, thereby improving the face recognition accuracy of smart home appliances such as smart door locks.
[0006] The technical solution adopted by this invention to solve the problem is as follows: Firstly, this application provides a face correction method, including: Get the initial image; Face detection is performed on the initial image to obtain a confidence score, which is used to characterize whether the initial image contains a face image; When the confidence level indicates that the initial image contains a face image, the initial face orientation corresponding to the initial image is determined; Based on the initial face orientation, the initial image is rotated to obtain the image to be identified. The image to be identified is associated with its corresponding rotation information. The image to be identified is the corrected image corresponding to the initial image.
[0007] In some embodiments of this application, determining the initial face orientation corresponding to the initial image includes: Extract the set of key feature points of the face image from the initial image; Calculate the coordinate difference between each key feature point in the set of key feature points; The initial face orientation corresponding to the initial image is determined based on this coordinate difference.
[0008] In some embodiments of this application, determining the initial face orientation corresponding to the initial image includes: Extract the set of key feature points and detection boxes of the face image in the initial image. The detection boxes are used to represent the coordinate information of the face image. Calculate the coordinate difference between each key feature point in the set of key feature points; The initial face orientation corresponding to the initial image is determined based on the detection box and the coordinate difference.
[0009] In some embodiments of this application, rotating the initial image based on the initial face orientation to obtain the image to be identified includes: Based on the initial face orientation, the initial image is rotated to obtain the image to be identified, including: When the initial face orientation is not a frontal orientation, the rotation information is determined based on the initial face orientation, and the rotation information includes the rotation angle and the rotation direction; Based on this rotation information, the initial image is rotated to obtain the corrected image; The face orientation of the corrected image is determined until the difference between the face orientation of the corrected image and the frontal face orientation meets a preset threshold. Then, the corrected image is output as the image to be identified. When the initial face orientation is frontal, the output initial image is the image to be recognized.
[0010] In some embodiments of this application, rotating the initial image according to the rotation information to obtain a corrected image includes: Extract the detection box of the face image in the initial image. The detection box is used to represent the coordinate information of the face image. Using the midpoint of the detection box as the rotation center, the initial image is rotated based on the rotation information to obtain the corrected image.
[0011] In some embodiments of this application, the method further includes: The face recognition model is invoked to perform face recognition on the image to be recognized, so as to obtain the recognition result.
[0012] In some embodiments of this application, after calling a face recognition model to perform face recognition on the image to be recognized to obtain the recognition result, the method further includes: The image to be identified is recovered based on the rotation information to obtain the initial image; The image is displayed based on this initial image.
[0013] Secondly, this application provides a face correction device, comprising: The acquisition module is used to acquire the initial image; The processing module is used to perform face detection on the initial image to obtain a confidence score, which is used to characterize whether the initial image contains a face image; when the confidence score indicates that the initial image contains a face image, the initial face orientation corresponding to the initial image is determined. Based on the initial face orientation, the initial image is rotated to obtain the image to be identified. The image to be identified is associated with its corresponding rotation information. The image to be identified is the corrected image corresponding to the initial image.
[0014] Thirdly, this application also provides a computer device, which includes: One or more processors; Memory; and One or more applications, wherein the applications are stored in memory and configured to be executed by a processor to implement the face correction method of any of the first aspects.
[0015] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, the computer program being loaded by a processor to perform the steps in the face correction method of any of the first aspects.
[0016] The beneficial effects of this invention are as follows: By normalizing the pose of facial images acquired in smart home scenarios and correcting facial orientation, the interference of head tilt on downstream algorithms is effectively eliminated. This significantly improves the accuracy and robustness of tasks such as face recognition, key point localization, and attribute analysis, providing stable and high-quality standardized input data for various visual processing systems, thereby improving the accuracy of face recognition in smart home appliances such as smart door locks. Simultaneously, by binding rotation information to the corrected image, the integrity of the image data and the reversibility of processing are ensured. Furthermore, this rotation information can serve as high-value quantified metadata, which can be directly used by downstream applications, avoiding redundant calculations and improving the overall operating efficiency of the system. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic diagram of a system architecture provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of an embodiment of the face correction method provided by the present invention; Figure 3 This is a schematic flowchart of a face correction method provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of a specific embodiment of the face correction device provided in this invention. Figure 5 This is a schematic diagram of an embodiment of the computer device provided in this invention. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] In the description of this application, the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first," "second," "third," etc., may explicitly or implicitly include one or more features.
[0021] In this application, the term "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be made without using these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of this application with unnecessary detail. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.
[0022] It should be noted that since the method in this application embodiment is executed in a computer device, the processing objects of each computer device exist in the form of data or information, such as time, which is essentially time information. It is understood that if size, quantity, position, etc. are mentioned in subsequent embodiments, they are all corresponding data that exist so that the computer device can process them. Specific details will not be elaborated here.
[0023] With the rapid development of IoT technology, the smart home ecosystem has become increasingly widespread. Against this backdrop, face detection technology serves as a key support, providing the foundation for core functions such as identity verification for smart locks, security monitoring for smart cameras, and user-sensory interaction for smart home appliances. This technology enhances the convenience and security of modern home life through automation and intelligence, demonstrating broad application prospects and significant market value. Existing face detection technologies primarily rely on deep learning models, especially convolutional neural networks. These models are trained on massive datasets of standard frontal face images, learning and mastering the inherent spatial feature hierarchy of the face, such as the relative spatial layout of the eyes, nose, and mouth. Therefore, the detection performance of these models is highly dependent on the standard, vertical orientation of the face in the input image. Under ideal or controlled conditions, when the face orientation conforms to the model's preset values, this technology can achieve efficient and accurate detection, and has become the mainstream technical solution in the industry.
[0024] However, in the actual deployment and application of smart homes, the above solutions have revealed significant shortcomings. First, the installation positions and angles of smart terminal devices (such as smart door locks and wall-mounted cameras) are diverse, often resulting in fixed tilts or rotations of the cameras themselves. Second, users' interaction postures are not static. These real-world factors cause the captured facial images to frequently exhibit discrete rotation states such as 90 degrees, 180 degrees, or 270 degrees. When such non-standard orientation images are input into existing detection models, the pre-learned feature space structure of the model is completely destroyed, making it impossible to extract effective facial features, thus causing the detection task to fail completely. This defect directly leads to serious problems such as smart door locks failing to identify legitimate users and security cameras missing critical events, greatly damaging product reliability and user experience, and constituting a core technical bottleneck restricting the development of facial recognition applications in smart homes. Therefore, there is an urgent need in this field for a solution that can overcome the influence of image rotation and achieve omnidirectional face detection.
[0025] To address this technical problem, this application provides the following technical solution: acquiring an initial image; performing face detection on the initial image to obtain a confidence score, which characterizes whether the initial image contains a face image; when the confidence score indicates that the initial image contains a face image, determining the initial face orientation corresponding to the initial image; rotating the initial image based on the initial face orientation to obtain an image to be recognized, which is associated with its corresponding rotation information, and the image to be recognized is the corrected image corresponding to the initial image. This method of normalizing the pose of face images acquired in smart home scenarios and correcting their face orientation effectively eliminates the interference of head tilt on downstream algorithms, significantly improving the accuracy and robustness of tasks such as face recognition, key point localization, and attribute analysis. It provides stable, high-quality standardized input data for various visual processing systems, thereby improving the face recognition accuracy of smart home appliances such as smart door locks. Simultaneously, by binding the rotation information to the corrected image, the integrity of the image data and the reversibility of the processing are ensured. On the other hand, this rotation information can be used directly by downstream applications as a high-value quantitative metadata, avoiding redundant calculations and improving the overall operating efficiency of the system.
[0026] This application provides a face correction method, apparatus, device, and storage medium for correcting the orientation of face images captured in a smart home scenario, ensuring the face images meet the requirements for face recognition, thereby improving the face recognition accuracy of smart home appliances such as smart door locks. The electronic device provided in this application can be implemented as various types of user terminals or as a server.
[0027] Electronic devices can use the face correction method provided in this application to correct the orientation of face images collected in smart home scenarios, so that the face images can meet the image requirements of face recognition, thereby improving the face recognition accuracy of smart home appliances such as smart door locks.
[0028] The above methods can be applied to many smart home devices or image processing fields, such as smart door locks, video access control, facial effects processing, etc.
[0029] In one exemplary solution, this face correction method can be applied to the face recognition scenario of a smart door lock. For example, in a smart lock's facial recognition scenario, when the user is within the camera's field of view, the camera captures an initial image of the user (which may or may not contain a face depending on the user's posture; if a face is present, it may be facing forward or backward). The smart lock's built-in processor performs face detection on this initial image to obtain a corresponding confidence score, which indicates whether the initial image contains a face. When the confidence score indicates that the initial image contains a face, the initial face orientation is determined. Then, based on this initial face orientation, appropriate correction processing is performed to obtain a face recognition image. Finally, face recognition is performed on this face recognition image to obtain a recognition result. If the recognition result indicates that the user is a registered user, the lock can be opened. If the recognition result indicates that the user is not a registered user, appropriate prompts or alarms can be issued (e.g., sending a notification to registered users that "a stranger is trying to open the door," etc.).
[0030] In one exemplary solution, the face correction method can be applied to face effect processing scenarios. For example, in face effect processing, an initial image of the user is captured by the camera of a mobile phone (which may have an application for face effect processing installed). (Depending on the user's posture, the initial image may or may not contain a face; if it does contain a face, the face may be facing forward or backward.) The phone's built-in processor can perform face detection on the initial image to obtain a corresponding confidence score. This confidence score is used to characterize whether the initial image contains a face image. When the confidence score indicates that the initial image contains a face image, the initial face orientation of the face image in the initial image is determined. Then, based on the initial face orientation, corresponding correction processing is performed to obtain an image to be recognized that conforms to face recognition; finally, face recognition is performed based on the image to be recognized to obtain the recognition result (at this time, the recognition result can indicate the position of multiple key feature points of the face, such as the eyes); after determining the user's eye position (or other positions) based on the recognition result, a special effect image (such as glasses) is added at this position; then, based on the rotation information bound during the correction process, the image to be recognized is restored, and the special effect image is also processed accordingly based on the rotation information to obtain an image with special effects processing based on the initial image, which is then displayed on the phone screen.
[0031] It should be understood that the above is only an example of the application scenario of face correction. There are many other possible application scenarios, which are not limited here.
[0032] The face correction method provided in this application embodiment is applied to, for example, Figure 1 The system architecture diagram shown is for your reference. Figure 1 To support a face correction method, the terminal device 100 connects to the server 300 via network 200, and the server 300 connects to the database 400. Network 200 can be a wide area network (WAN), a local area network (LAN), or a combination of both. The client for implementing the face correction scheme is deployed on the terminal device 100, or it can run on the terminal device 100 as a standalone application. The specific form of the client is not limited here.
[0033] The server 300 involved in this application can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms.
[0034] Terminal equipment 100, also known as user equipment (UE), mobile station (MS), mobile terminal (MT), customer premises equipment (CPE), etc., can be a device that includes both receiving and transmitting hardware, that is, a device with receiving and transmitting hardware capable of performing bidirectional communication on a bidirectional communication link. Such equipment can include cellular or other communication devices with single-line displays, multi-line displays, or no multi-line displays. Examples include handheld devices with wireless connectivity, vehicle-mounted devices, machine-type communication (MTC) terminals, etc. Currently, terminal devices 100 can include: mobile phones, tablets, laptops, PDAs, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving vehicles, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, or wireless terminals in smart homes, etc. For example, wireless terminals in self-driving vehicles can be drones, helicopters, or airplanes. For example, wireless terminals in vehicle-to-everything (V2X) systems can be in-vehicle equipment, vehicle-mounted equipment, in-vehicle modules, vehicles, or ships, etc. Wireless terminals in industrial control can be cameras, robots, or robotic arms, etc. Wireless terminals in smart homes can be televisions, air conditioners, robot vacuums, speakers, or set-top boxes, etc.
[0035] It should be noted that the terminal device 100 may be a device or apparatus with a chip, or a device or apparatus with integrated circuitry, or a chip, module, or control unit in the device or apparatus shown above; this application does not impose any specific limitations. The solution provided in this application can be implemented by the terminal device 100 and the server 300 working together.
[0036] In short, a database can be viewed as an electronic filing cabinet—a place to store electronic files, where users can perform operations such as adding, querying, updating, and deleting data. A "database" is a collection of data stored together in a certain way, shared by multiple users, with minimal redundancy, and independent of application programs. A Database Management System (DBMS) is a computer software system designed to manage databases, generally possessing basic functions such as storage, retrieval, security, and backup. DBMSs can be classified according to the database model they support, such as relational or Extensible Markup Language (XML); or according to the type of computer they support, such as server clusters or mobile phones; or according to the query language used, such as Structured Query Language (SQL) or XQuery; or according to performance priorities, such as maximum scale or highest operating speed; or other classification methods. Regardless of the classification method used, some DBMSs can cross categories, for example, supporting multiple query languages simultaneously. In this application, the database 400 can be used to store data such as initial images, images to be recognized, face detection models, or face recognition models.
[0037] Those skilled in the art will understand that Figure 1 The system architecture diagram shown is one possible system architecture for this application and does not constitute a limitation on the system architecture of this application. Other system architectures may include more advanced architectures. Figure 1 The number of more or fewer terminal devices or servers shown, for example Figure 1 The diagram shows one server. It is understood that the system architecture may also include one or more other terminal devices or servers, which are not limited here.
[0038] It should be noted that, Figure 1 The system architecture shown is an example. The servers and scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of servers and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0039] like Figure 2 The diagram shown is a flowchart of an embodiment of the face correction method in this application. The following description, using a terminal device as the execution subject, details the face correction method, which may include the following steps 201-204: 201. Obtain the initial image.
[0040] The terminal device captures images of its shooting area using a built-in or external camera and uses them as the initial image.
[0041] The initial image may contain a face or not, or if a face is present, its orientation may not meet the requirements for facial recognition, depending on the installation location of the smart home device or the user's posture when entering the shooting area.
[0042] In one exemplary solution, when the terminal device is a small smart home device, it can be configured as a low-power device to reduce energy consumption. That is, the facial recognition system of the terminal device is normally set to a sleep state, and only activated to capture and recognize facial images when a trigger event is detected.
[0043] The triggering event could be a user entering the smart home's camera range; or a user touching or pressing a specific button or area of the smart home. The smart home can detect whether a user has entered its camera range using infrared sensors, microwave sensors, and / or radar sensors.
[0044] Specifically, this infrared sensor can detect changes in the movement of infrared radiation (body temperature) at a specific wavelength emitted by the human body. When a heat source enters and moves within the sensor's fan-shaped detection area, it triggers the smart home system to activate the facial recognition system.
[0045] This microwave sensor and / or radar sensor can actively emit low-power microwave / radar waves and detect the Doppler effect of the reflected waves to determine if there is movement. When movement is detected, the smart home system is triggered to activate the facial recognition system.
[0046] The principle behind a user touching or pressing a specific button or area of a smart home device is as follows: The user generates an electrical signal through physical action to wake up the facial detection system deployed on the smart home. For example, touching the numeric keypad area of a door lock, pressing the doorbell button, or lifting / pressing the handle generates an electrical signal to wake up the system.
[0047] 202. Perform face detection on the initial image to obtain a confidence score, which is used to characterize whether the initial image contains a face image.
[0048] The terminal device calls a face detection model to perform face detection on the initial image to obtain a confidence score, and determines whether the initial image contains a face image based on the confidence score.
[0049] In one exemplary approach, the face detection model can employ a lightweight, edge-device-optimized deep learning model that can quickly scan images and output face location and confidence scores. For example, a lightweight convolutional neural network could be used, such as variants of MTCNN, BlazeFace, or RetinaFace. These models are specifically designed and pruned to maintain high accuracy while achieving real-time operation on embedded processors (or even dedicated NPUs).
[0050] During the operation of the face detection model, an appropriate algorithm can be used to scan the initial image. At each location, the model determines whether the region matches facial features. For each region determined to match facial features, the face detection model can output a confidence score, which can be a value between 0 and 1, representing the degree to which the face detection model confirms that the region is indeed a face. A confidence threshold can be set when the face detection model makes its decision. When the confidence score output by the face detection model is greater than the confidence threshold, it can be determined that a face has been successfully detected; when the confidence score output by the face detection model is less than or equal to the confidence threshold, it can be determined that there is no face image in the image.
[0051] Optionally, the face detection model can also output a detection box when outputting the results. In this case, the detection box is used to define the coordinates of the face's location, that is, to indicate the specific position and size of the face in the image. In an exemplary scheme, the detection box can be a rectangular region defined in a two-dimensional image coordinate system. It is usually represented by a set of four values, and the two most common representations are as follows: In one example scheme, the top-left and bottom-right coordinates are used: (x_min, y_min, x_max, y_max).
[0052] In another exemplary scheme, the top-left corner coordinates and the width and height of the rectangle are used: (x, y, width, height).
[0053] It should be understood that these two representations can be converted to each other. After outputting the representation of the detection box, the terminal device can draw based on the representation of the detection box in the display interface, thereby outlining the face region in the initial image.
[0054] 203. When the confidence level indicates that the initial image contains a face image, determine the initial face orientation corresponding to the initial image.
[0055] When the terminal device determines that the initial image contains a face, it can prepare to perform subsequent operations such as face recognition and face effects processing. To ensure the accuracy of these subsequent operations, the terminal device can detect the face orientation of the face image and perform corresponding image correction if the face orientation does not meet the requirements of the subsequent operations. The face orientation detection process can be implemented as follows: In one exemplary solution, to be applicable to platforms with limited computing resources such as mobile phones and embedded devices, and to achieve fast and accurate quantification of the deflection and pitch angles of the face, this embodiment can determine the face orientation by using key point coordinates. That is, the terminal device can extract a set of key feature points from the face image in the initial image; then calculate the coordinate difference between each key feature point in the set; and finally determine the initial face orientation corresponding to the initial image based on the coordinate difference.
[0056] The set of key feature points may include the following key feature points: corners of the eyes, eyebrows, tip of the nose, nostrils, corners of the mouth, chin contour, etc.
[0057] Calculating coordinate differences can be understood as calculating yaw, pitch, and roll angles. The yaw angle indicates the head's left-right movement around the vertical axis (head shaking); the pitch angle indicates the head's up-down rotation around the horizontal axis (head nodding); and the roll angle indicates the head's left-right tilt around the front-rear axis (head tilting). During the calculation, the tip of the nose can be used as the center reference point of the face. The coordinate difference between the corners of the eyes is used to determine left-right and tilt; the coordinate difference between the corners of the mouth is used to help determine left-right rotation; and the coordinate difference between the tip of the chin is used to determine up-down nodding.
[0058] In practical applications, when the head is kept level, the line connecting the left and right eyes or the left and right corners of the mouth should be nearly horizontal; when the head is tilted, this line will tilt. Based on this principle, the roll angle can be calculated by measuring the slope of the line connecting the corners of the eyes.
[0059] When a person's face is directly facing the camera, the tip of the nose should be roughly at the horizontal midpoint of the facial contour, meaning the distance from the tip of the nose to the left side of the face is approximately equal to the distance to the right side. When the head turns to the left, the left side of the face appears narrower (perspective effect), and the distance from the tip of the nose to the left side of the face becomes significantly smaller than the distance to the right side. Based on this principle, the yaw angle can be calculated by measuring the horizontal distance between the tip of the nose and the left or right corners of the mouth or eyes.
[0060] When a person's face is level, the vertical distance from the eyes to the tip of the nose has a relatively fixed ratio to the vertical distance from the tip of the nose to the mouth (or chin). When a person tilts their head up, their eyes move closer to their nose, while their chin moves further away, causing the upper part of the face to be compressed vertically; the opposite is true when the head is tilted down. Based on this principle, the pitch angle can be calculated by measuring the vertical distances between the tip of the nose and upper and lower reference points on the face. The upper reference point can be the center of the two corners of the eyes, and the lower reference point can be the tip of the chin.
[0061] Finally, the orientation of the face in the initial image is determined based on the yaw angle, pitch angle, and roll angle.
[0062] In one exemplary solution, to achieve efficient and low-cost face orientation determination on resource-constrained platforms such as mobile phones and embedded devices, and to improve the accuracy of face orientation determination, this embodiment may also provide the following technical solution: extracting a set of key feature points and a detection box from the face image in the initial image, wherein the detection box is used to represent the coordinate information of the face image; calculating the coordinate difference of each key feature point in the set of key feature points; and determining the initial face orientation corresponding to the initial image based on the detection box and the coordinate difference.
[0063] In this scheme, the terminal device can first normalize the coordinates of each key feature point in the key feature point set based on the detection box to obtain the normalized coordinates of each key feature point relative to the detection box; then, it can calculate the yaw angle, pitch angle, and roll angle based on the key feature points with normalized coordinates; finally, it can determine the face orientation based on the yaw angle, pitch angle, and roll angle. In this embodiment, the process of calculating the yaw angle, pitch angle, and roll angle based on the key feature points with normalized coordinates can be referred to the foregoing description, and will not be repeated here.
[0064] 204. Based on the initial face orientation, the initial image is rotated to obtain the image to be identified. The image to be identified is associated with its corresponding rotation information. The image to be identified is the corrected image corresponding to the initial image.
[0065] After the terminal device obtains the initial face orientation, it performs the following steps based on this orientation: If the initial face orientation is a frontal orientation, the initial image is output as the image to be recognized; if the initial face orientation is not a frontal orientation, rotation information, including rotation angle and rotation direction, is determined based on the initial face orientation; the initial image is rotated based on the rotation information to obtain a corrected image; the face orientation of the corrected image is determined until the difference between the face orientation of the corrected image and the frontal orientation meets a preset threshold, at which point the corrected image is output as the image to be recognized. During the correction process, the rotation information or correction information corresponding to each rotation operation is associated with the output corrected image.
[0066] Here, the "frontal orientation" can be understood as the pose that allows the computer to accurately and reliably recognize a face. In this pose, the key features of the face (eyes, nose, and mouth) are fully and completely exposed, and are less affected by perspective distortion. This allows the algorithm to extract the richest and most stable feature information, thus maintaining high recognition reliability under different lighting conditions and facial expressions. In this embodiment, the "frontal orientation" can also be understood as the calculated yaw angle, pitch angle, and roll angle being within an acceptable, small range around zero degrees.
[0067] In this embodiment, the following scheme can be adopted when determining the rotation information based on the initial face orientation: The rotation angle in the two-dimensional plane is determined based on the roll angle, and the rotation angle in the two-dimensional plane is also determined based on the direction of the roll angle. For example, if the roll angle is +10 degrees (indicating that the head tilts 10 degrees to the left), then the rotation angle can be -10 degrees.
[0068] For the correction of pitch and yaw angles, a transformation matrix can be determined based on the pitch and yaw angles, and the face image in the initial image can be corrected based on the transformation matrix.
[0069] Optionally, if a corrected image meeting facial recognition requirements is not obtained within a preset time period or preset number of corrections during the correction process, the user can be prompted to retake the photo. For example, a prompt such as "Please face the camera" can be displayed.
[0070] It should be understood that, in order to ensure the stability of the image, the rotation center of the initial image also needs to be determined when determining the rotation information.
[0071] In one exemplary scheme, the rotation center can be assumed to be the center of the image.
[0072] In another exemplary scheme, the terminal device can extract a detection box of the face image in the initial image, the detection box being used to characterize the coordinate information of the face image; and rotate the initial image based on the rotation information with the midpoint of the detection box as the rotation center to obtain the corrected image.
[0073] Optionally, in order to ensure image quality during the correction process, this embodiment may also introduce an integrated bilinear interpolation algorithm to smoothly fill the pixel missing areas during the rotation process, ensuring that the image distortion rate after rotation is ≤2%, which meets the clarity requirements of the traditional face detection model for the input image.
[0074] In this embodiment, after the image correction of the initial image is completed, the corresponding subsequent operations can be performed.
[0075] In one exemplary solution, the terminal device can invoke a face recognition model to perform face recognition on the image to be recognized in order to obtain a recognition result. For example, for a smart door lock, after obtaining the corrected image to be recognized, the smart door lock performs face recognition on the image to determine whether the current user is a registered user; if so, the door can be opened; otherwise, the door opening operation is not performed, or a prompt is sent to the registered user.
[0076] In this embodiment, the image to be recognized is associated with the rotation information, so that after high-precision recognition is performed using the image to be recognized, it is restored to its original tilted state for display. This balances the accuracy of machine recognition with the naturalness of human-computer interaction. That is, the terminal device can restore the image to be recognized based on the rotation information to obtain the initial image; and display the image based on the initial image. For example, in the application of smart door locks, when unlocking based on facial recognition, the smart door lock performs facial recognition based on the corrected image to be recognized; after the recognition is confirmed, the smart door lock displays the initial image captured by its camera on its display screen, thereby providing users with real and lossless visual feedback, faithfully recording the real posture at the moment of interaction, greatly improving the user experience and credibility in verification scenarios.
[0077] Optionally, the image restoration based on this rotation information can also be applied to facial effects processing scenarios. After the terminal device identifies the part of the face to be recognized based on the image to be recognized, it adds effects to the part to be recognized based on the effects options; when displaying, the effects need to be displayed on the initial image. At this time, the image to be recognized and the effects need to be restored based on the rotation information to complete the overlay of the effects on the initial image.
[0078] As described above, this embodiment normalizes the pose of facial images collected in smart home scenarios and corrects facial orientation, effectively eliminating the interference of head tilt on downstream algorithms. This significantly improves the accuracy and robustness of tasks such as facial recognition, key point localization, and attribute analysis, providing stable, high-quality standardized input data for various visual processing systems, thereby improving the facial recognition accuracy of smart home appliances such as smart door locks. Simultaneously, by binding rotation information to the corrected image, data integrity and processing reversibility are ensured. Furthermore, this rotation information, as a high-value quantified metadata, can be directly used by downstream applications, avoiding redundant calculations and improving the overall system operating efficiency.
[0079] The following is based on Figure 3 The flowchart shown illustrates the face correction method of this application: like Figure 3 As shown, the system first captures an image using a camera; then, it performs face detection on the image, extracting key feature points upon confirmation of a face; it then determines the face orientation based on these key feature points; if the face is facing forward, face recognition is performed; if not, the face is rotated based on its orientation to obtain a corrected image; the corrected image is then re-evaluated for face orientation; if the corrected image is facing forward, face recognition is performed; if it is still not facing forward, it is rotated again until a face-oriented image is output; if no face-oriented image is output within a preset time period or a preset number of corrections, the user is prompted to retake the image; after face recognition, the recognition result is directly output; if display is required, the image can be restored based on the rotation information obtained during the correction process to display the initial image.
[0080] To better implement the face correction method in the embodiments of this application, based on the face correction method, the embodiments of this application also provide a face correction device, such as... Figure 4 As shown, the face correction device 400 includes: Module 401 is used to acquire the initial image; The processing module 402 is used to perform face detection on the initial image to obtain a confidence score, which is used to characterize whether the initial image contains a face image; when the confidence score indicates that the initial image contains a face image, the initial face orientation corresponding to the initial image is determined. Based on the initial face orientation, the initial image is rotated to obtain the image to be identified. The image to be identified is associated with its corresponding rotation information. The image to be identified is the corrected image corresponding to the initial image.
[0081] In this embodiment, the pose normalization of facial images acquired in smart home scenarios is used for facial orientation correction, effectively eliminating the interference of head tilt on downstream algorithms. This significantly improves the accuracy and robustness of tasks such as facial recognition, key point localization, and attribute analysis, providing stable, high-quality standardized input data for various visual processing systems, thereby improving the facial recognition accuracy of smart home appliances such as smart door locks. Simultaneously, by binding rotation information to the corrected image, data integrity and processing reversibility are ensured. Furthermore, this rotation information, as a high-value quantified metadata, can be directly used by downstream applications, avoiding redundant calculations and improving the overall system operating efficiency.
[0082] In some embodiments of this application, the processing module 402 is specifically used for: Extract the set of key feature points of the face image from the initial image; Calculate the coordinate difference between each key feature point in the set of key feature points; The initial face orientation corresponding to the initial image is determined based on this coordinate difference.
[0083] In this embodiment, the deflection and pitch angles of a face can be quickly and accurately quantified by performing simple geometric calculations on the coordinates of key points. This achieves low-cost and high-efficiency real-time judgment, and is especially suitable for platforms with limited computing resources, such as mobile phones and embedded devices. It has extremely high engineering practical value.
[0084] In some embodiments of this application, the processing module 402 is specifically used for: Extract the set of key feature points and detection boxes of the face image in the initial image. The detection boxes are used to represent the coordinate information of the face image. Calculate the coordinate difference between each key feature point in the set of key feature points; The initial face orientation corresponding to the initial image is determined based on the detection box and the coordinate difference.
[0085] In this embodiment, the detection box is used as a stable reference frame, making the pose estimation based on coordinate difference more robust to changes in face size and position. At the same time, the method has a very small computational load and strong real-time performance, realizing efficient and low-cost face orientation judgment on resource-constrained platforms such as mobile phones and embedded devices, and has extremely high engineering application value.
[0086] In some embodiments of this application, the processing module 402 is specifically used for: When the initial face orientation is not a frontal orientation, the rotation information is determined based on the initial face orientation, and the rotation information includes the rotation angle and the rotation direction; Based on this rotation information, the initial image is rotated to obtain the corrected image; The face orientation of the corrected image is determined until the difference between the face orientation of the corrected image and the frontal face orientation meets a preset threshold. Then, the corrected image is output as the image to be identified. When the initial face orientation is frontal, the output initial image is the image to be recognized.
[0087] In this embodiment, the rotation information is determined based on the user's initial facial orientation and continuously tracked and corrected. This transforms static, absolute correction into dynamic, relative tracking, which not only establishes a personalized benchmark but also achieves more stable and efficient correction and avoids unnecessary jitter.
[0088] In some embodiments of this application, the processing module 402 is specifically used for: Extract the detection box of the face image in the initial image. The detection box is used to represent the coordinate information of the face image. Using the midpoint of the detection box as the rotation center, the initial image is rotated based on the rotation information to obtain the corrected image.
[0089] In this embodiment, image correction is performed with the center of the face detection box as the rotation center. This ensures that the face remains near its original position in the image after rotation, achieving "rotation in place". This maximizes the integrity of the corrected face data and provides stable and high-quality input for subsequent recognition, analysis and other tasks, significantly improving the robustness and accuracy of the entire system.
[0090] In some embodiments of this application, the processing module 402 is further configured to: The face recognition model is invoked to perform face recognition on the image to be recognized, so as to obtain the recognition result.
[0091] In this embodiment, the pose normalization of facial images acquired in smart home scenarios is used for facial orientation correction, effectively eliminating the interference of head tilt on downstream algorithms. This significantly improves the accuracy and robustness of tasks such as facial recognition, key point localization, and attribute analysis, providing stable, high-quality standardized input data for various visual processing systems, thereby improving the facial recognition accuracy of smart home appliances such as smart door locks. Simultaneously, by binding rotation information to the corrected image, data integrity and processing reversibility are ensured. Furthermore, this rotation information, as a high-value quantified metadata, can be directly used by downstream applications, avoiding redundant calculations and improving the overall system operating efficiency.
[0092] In some embodiments of this application, the processing module 402 is further configured to: The image to be identified is recovered based on the rotation information to obtain the initial image; The image is displayed based on this initial image.
[0093] In this embodiment, after high-precision recognition is achieved using the corrected image, it is then restored to its original tilted state for display. This balances the accuracy of machine recognition with the naturalness of human-computer interaction. For the machine, using the corrected image ensures recognition accuracy; for the user, displaying the original image provides realistic and lossless visual feedback, faithfully recording the actual posture at the moment of interaction, greatly enhancing the user experience and credibility in verification scenarios.
[0094] This application also provides a computer device that integrates any of the face correction devices provided in this application. The computer device includes: One or more processors; Memory; and One or more applications, wherein the applications are stored in memory and configured to be executed by a processor from the steps of the face correction method in any of the above embodiments of the face correction method.
[0095] This application also provides a computer device that integrates any of the face correction devices provided in this application. For example... Figure 5 As shown, it illustrates a schematic diagram of the computer device involved in the embodiments of this application, specifically: The computer device may include components such as a processor 501 with one or more processing cores, a memory 502 with one or more computer-readable storage media, a power supply 503, and an input unit 504. Those skilled in the art will understand that... Figure 5 The computer device structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein: The processor 501 is the control center of the computer device. It connects various parts of the computer device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 502, and by calling data stored in the memory 502, it performs various functions of the computer device and processes data, thereby providing overall monitoring of the computer device. Optionally, the processor 501 may include one or more processing cores; preferably, the processor 501 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 501.
[0096] The memory 502 can be used to store software programs and modules. The processor 501 executes various functional applications and data processing by running the software programs and modules stored in the memory 502. The memory 502 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 502 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 502 may also include a memory controller to provide the processor 501 with access to the memory 502.
[0097] The computer equipment also includes a power supply 503 that supplies power to the various components. Preferably, the power supply 503 can be logically connected to the processor 501 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 503 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0098] The computer device may also include an input unit 504, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0099] Although not shown, the computer device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 501 in the computer device loads the executable files corresponding to the processes of one or more application programs into the memory 502 according to the following instructions, and the processor 501 runs the application programs stored in the memory 502 to realize various functions, as follows: Get the initial image; Face detection is performed on the initial image to obtain a confidence score, which is used to characterize whether the initial image contains a face image; When the confidence level indicates that the initial image contains a face image, the initial face orientation corresponding to the initial image is determined; Based on the initial face orientation, the initial image is rotated to obtain the image to be identified. The image to be identified is associated with its corresponding rotation information. The image to be identified is the corrected image corresponding to the initial image.
[0100] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0101] Therefore, embodiments of this application provide a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), a disk, or an optical disk, etc. A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in any of the face correction methods provided in embodiments of this application. For example, the computer program loaded by the processor can execute the following steps: Get the initial image; Face detection is performed on the initial image to obtain a confidence score, which is used to characterize whether the initial image contains a face image; When the confidence level indicates that the initial image contains a face image, the initial face orientation corresponding to the initial image is determined; Based on the initial face orientation, the initial image is rotated to obtain the image to be identified. The image to be identified is associated with its corresponding rotation information. The image to be identified is the corrected image corresponding to the initial image.
[0102] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the detailed descriptions of other embodiments above, which will not be repeated here.
[0103] In practice, each of the above units or structures can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units or structures, please refer to the previous method embodiments, which will not be repeated here.
[0104] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0105] The foregoing has provided a detailed description of a face correction method, apparatus, computer device, and computer-readable storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A face correction method, characterized in that, include: Get the initial image; Face detection is performed on the initial image to obtain a confidence score, which is used to characterize whether the initial image contains a face image; When the confidence level indicates that the initial image contains a face image, the initial face orientation corresponding to the initial image is determined; The initial image is rotated based on the initial face orientation to obtain an image to be identified. The image to be identified is associated with its corresponding rotation information. The image to be identified is the corrected image corresponding to the initial image.
2. The method according to claim 1, characterized in that, Determining the initial face orientation corresponding to the initial image includes: Extract the set of key feature points of the face image from the initial image; Calculate the coordinate difference between each key feature point in the set of key feature points; The initial face orientation corresponding to the initial image is determined based on the coordinate difference.
3. The method according to claim 1, characterized in that, Determining the initial face orientation corresponding to the initial image includes: Extract the set of key feature points and detection boxes of the face image in the initial image, wherein the detection boxes are used to represent the coordinate information of the face image; Calculate the coordinate difference between each key feature point in the set of key feature points; The initial face orientation corresponding to the initial image is determined based on the detection box and the coordinate difference.
4. The method according to claim 1, characterized in that, The initial image is rotated based on the initial face orientation to obtain the image to be recognized, including: When the initial face orientation is not a frontal orientation, the rotation information is determined based on the initial face orientation, and the rotation information includes the rotation angle and the rotation direction; The initial image is rotated based on the rotation information to obtain a corrected image; The face orientation of the corrected image is determined until the difference between the face orientation of the corrected image and the frontal face orientation meets a preset threshold. Then, the corrected image is output as the image to be recognized. When the initial face orientation is frontal, the initial image is output as the image to be recognized.
5. The method according to claim 4, characterized in that, The initial image is rotated based on the rotation information to obtain the corrected image, including: Extract the detection box of the face image in the initial image, the detection box is used to represent the coordinate information of the face image; Using the midpoint of the detection box as the rotation center, the initial image is rotated based on the rotation information to obtain the corrected image.
6. The method according to claim 1, characterized in that, The method further includes: The face recognition model is invoked to perform face recognition on the image to be recognized in order to obtain the recognition result.
7. The method according to claim 6, characterized in that, After calling a face recognition model to perform face recognition on the image to be recognized to obtain the recognition result, the method further includes: The image to be identified is recovered based on the rotation information to obtain the initial image; The image is displayed based on the initial image.
8. A face correction device, characterized in that, include: The acquisition module is used to acquire the initial image; The processing module is used to perform face detection on the initial image to obtain a confidence score, which is used to characterize whether the initial image contains a face image; when the confidence score indicates that the initial image contains a face image, the module determines the initial face orientation corresponding to the initial image. The initial image is rotated based on the initial face orientation to obtain an image to be identified. The image to be identified is associated with its corresponding rotation information. The image to be identified is the corrected image corresponding to the initial image.
9. A computer device, characterized in that, The computer device includes: One or more processors; Memory; and One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the method of any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, It stores a computer program, which is loaded by a processor to perform the steps of the method according to any one of claims 1 to 7.