Identification method, device, equipment, medium and product

By applying different transformation matrices to facial images to transform facial and target area images to different image spaces, the problem of low recognition accuracy of specific parts in facial images is solved, and efficient visible area recognition is achieved on terminal devices.

CN121768050APending Publication Date: 2026-03-31BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies suffer from low recognition accuracy due to excessive loss of detail when identifying specific parts of facial images, such as lips and mouth. This is especially true on terminal devices with limited computing power, where it is difficult to achieve high-precision recognition of visible areas.

Method used

By transforming the face image to different image spaces, the pixel ratio of the face and target parts is optimized using the first transformation matrix and the second transformation matrix, respectively, to ensure that the face and target parts are described better with as little loss of detail as possible, thereby improving the recognition effect.

Benefits of technology

It effectively avoids the loss of detail when recognizing different parts in the same image space, improves the recognition accuracy of visible areas of faces and target parts, and is especially suitable for terminal devices with limited computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121768050A_ABST
    Figure CN121768050A_ABST
Patent Text Reader

Abstract

The invention discloses an identification method and device, equipment, a medium and a product, and the method comprises the steps: firstly obtaining a face image, the face comprises a target part; the face image is transformed into a first image space according to a first transformation matrix to obtain a first transformed image, so that the pixel proportion of the face in the first transformed image is as large as possible, and the face image is transformed into a second image space according to a second transformation matrix to obtain a second transformed image, so that the pixel proportion of the face in the second transformed image is as large as possible. The pixel proportion of the target part in the second transformation image is as large as possible; and then, according to the first transformation image, the first transformation matrix, the second transformation image and the second transformation matrix, determining a visible region identification result of the face and a visible region identification result of the target part. Visibly, according to the application, the visible region recognition processing is realized by means of different image spaces for different parts, so that the recognition effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a recognition method, apparatus, device, medium, and product. Background Technology

[0002] For some scenarios, such as facial effects or facial beautification, there is a need to identify certain parts, such as the visible areas of the face, lips, and mouth, from a facial image so that these areas can be used for other processing, such as adding lipstick effects.

[0003] However, how to achieve the above identification is a technical problem that urgently needs to be solved. Summary of the Invention

[0004] This application provides an identification method, apparatus, device, medium, and product that helps improve identification effectiveness.

[0005] To achieve the above objectives, the technical solution provided in this application is as follows:

[0006] This application provides a recognition method, the method comprising: acquiring a face image, the face including a target region; transforming the face image to a first image space according to a first transformation matrix to obtain a first transformed image; and transforming the face image to a second image space according to a second transformation matrix to obtain a second transformed image, wherein the first transformation matrix is ​​determined based on facial key points of the face image and key points of the first image space, the key points of the first image space being used to describe the cropping constraints of the face, the facial key points including key points of the target region; the second transformation matrix being determined based on the key points of the target region and key points of the second image space, the key points of the second image space being used to describe the cropping constraints of the target region, the pixel percentage of the target region in the second transformed image being higher than the pixel percentage of the target region in the first transformed image; and determining the visible region recognition result of the face and the visible region recognition result of the target region based on the first transformed image, the first transformation matrix, the second transformed image, and the second transformation matrix.

[0007] In one possible implementation, the visible region identification result of the face is obtained by transforming the region prediction result of the first transformed image back to the image space of the face image according to the inverse transformation matrix corresponding to the first transformation matrix. The region prediction result of the first transformed image is used to describe the position of the visible region of the face in the first transformed image. The visible region identification result of the target part is obtained by transforming the region prediction result of the second transformed image back to the image space of the face image according to the inverse transformation matrix corresponding to the second transformation matrix. The region prediction result of the second transformed image is used to describe the position of the visible region of the target part in the second transformed image.

[0008] In one possible implementation, the method further includes: stitching the first transformed image and the second transformed image together to obtain a stitched image; and determining the region prediction result of the first transformed image and the region prediction result of the second transformed image based on the stitched image, wherein the region prediction result of the first transformed image is used to describe the position of the visible region of the face in the first transformed image, and the region prediction result of the second transformed image is used to describe the position of the visible region of the target part in the second transformed image.

[0009] The step of determining the visible region recognition result of the face and the visible region recognition result of the target part based on the first transformed image, the first transformed matrix, the second transformed image and the second transformed matrix includes: determining the visible region recognition result of the face and the visible region recognition result of the target part based on the region prediction result of the first transformed image, the first transformed matrix, the region prediction result of the second transformed image and the second transformed matrix.

[0010] In one possible implementation, the width of the first transformed image is the same as the width of the second transformed image; the height of the stitched image is determined based on the sum of the heights of the first transformed image and the second transformed image.

[0011] In one possible implementation, the width of the first transformed image is smaller than the height of the first transformed image.

[0012] In one possible implementation, the ratio between the height of the first transformed image and the height of the second transformed image is determined based on the ratio between the height of the face and the maximum height of the target part. The target part includes multiple shapes, each with a different height, and the height of each shape is not greater than the maximum height of the target part.

[0013] In one possible implementation, the target area includes the lips and the oral cavity, and the visible area recognition result of the target area includes the visible area recognition result of the lips and the visible area recognition result of the oral cavity.

[0014] In one possible implementation, the identification method is applied to a terminal device.

[0015] This application provides a recognition device, comprising: an acquisition unit for acquiring a face image, the face including a target region; a transformation unit for transforming the face image to a first image space according to a first transformation matrix to obtain a first transformed image, and transforming the face image to a second image space according to a second transformation matrix to obtain a second transformed image, wherein the first transformation matrix is ​​determined based on facial key points of the face image and key points of the first image space, the key points of the first image space being used to describe the cropping constraints of the face, the facial key points including key points of the target region, the second transformation matrix being determined based on the key points of the target region and key points of the second image space, the key points of the second image space being used to describe the cropping constraints of the target region, and the pixel percentage of the target region in the second transformed image being higher than the pixel percentage of the target region in the first transformed image; and a determination unit for determining the visible region recognition result of the face and the visible region recognition result of the target region based on the first transformed image, the first transformation matrix, the second transformed image, and the second transformation matrix.

[0016] This application provides an electronic device, the device comprising: a processor and a memory; the memory for storing instructions or computer programs; the processor for executing the instructions or computer programs in the memory, so that the electronic device performs the identification method provided in this application.

[0017] This application provides a computer-readable medium storing instructions or computer programs that, when executed on a device, cause the device to perform the identification method provided in this application.

[0018] This application provides a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the identification method provided in this application.

[0019] Compared with related technologies, this application has at least the following advantages:

[0020] The technical solution provided in this application first acquires a face image, ensuring that the face depicted in the image includes target areas such as lips and mouth. Then, according to a first transformation matrix, the face image is transformed to a first image space to obtain a first transformed image, maximizing the pixel proportion of the face in the first transformed image. This allows the first transformed image to better describe the face presented in the face image with minimal loss of facial details. Finally, according to a second transformation matrix, the face image is transformed to a second image space to obtain a second transformed image, maximizing the pixel proportion of the target area in the second transformed image. The proportion is made as large as possible so that the second transformed image can better describe the target part presented in the face image with as little loss of detail as possible; then, based on the first transformed image, the first transformed matrix, the second transformed image and the second transformed matrix, the visible region recognition result of the face and the visible region recognition result of the target part are determined, so that the visible region recognition result of the face is used to describe the position of the visible region of the face in the face image, and the visible region recognition result of the target part is used to describe the position of the visible region of the target part in the face image.

[0021] The first transformation matrix is ​​determined based on the facial key points of the face image and the key points of the first image space. The key points of the first image space are used to describe the cropping constraints of the face. This allows the first transformed image determined based on the first transformation matrix to better represent the cropping result of the face in the face image. As a result, the first transformed image contains little or no information that interferes with the recognition of the face, such as background information. This allows the first transformed image to better describe the face presented in the face image with as little loss of facial details as possible. Thus, the visible area of ​​the face determined based on the first transformed image is more accurate.

[0022] Furthermore, since the second transformation matrix is ​​determined based on the key points of the target region and the key points of the second image space, which are used to describe the cropping constraints of the target region, the second transformed image determined based on the second transformation matrix can better represent the cropping result of the target region in the face image. This results in the second transformed image containing a small amount or no information that interferes with the recognition of the target region, such as background information, other facial information, etc. Also, since the pixel ratio of the target region in the second transformed image is higher than that in the first transformed image, the target region details described by the second transformed image are more than those described by the first transformed image, thus making the visible area of ​​the target region determined based on the second transformed image more accurate.

[0023] As can be seen, this application achieves visible area recognition processing by using different image spaces for different parts. This effectively avoids the defects caused by excessive loss of details in some parts when using the same image space to achieve visible area recognition processing for different parts, thus improving the recognition effect. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 A flowchart illustrating an identification method provided in an embodiment of this application;

[0026] Figure 2 A schematic diagram of an identification process provided for an embodiment of this application;

[0027] Figure 3 This is a schematic diagram of the structure of an identification device provided in an embodiment of this application;

[0028] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0029] Research has revealed that some recognition schemes involve the following steps: After acquiring a face image, face detection processing is performed to obtain a bounding box, such as a rectangle; this bounding box is then expanded outward to obtain a square box, ensuring that the square box is larger than the bounding box and that its geometric center overlaps with the bounding box's geometric center; the image region within this square box is then cropped from the face image to obtain a cropped image that includes the face and its surrounding area; this cropped image is then scaled to a target size, such as H×W, to obtain a scaled image; and finally, this scaled image is input into a pre-built machine learning model, enabling the model to predict the visible areas of certain parts, such as the face, lips, and mouth, based on the scaled image.

[0030] The study also found that when the recognition scheme shown above is applied to devices with limited computing power, such as terminal devices, due to real-time requirements and computing power limitations, the recognition scheme needs to meet the following constraints: the scaled image involved in the recognition scheme has a small image resolution (also known as size), such as 160×160 or 128×128, and the model involved in the recognition scheme uses an ultra-lightweight neural network model, such as a model computational cost (Floating Point Operations, FLOPs) of less than 20 megabytes.

[0031] Further research revealed the following drawbacks in the recognition schemes described above: Because the visible areas of some parts have a very low pixel percentage in the scaled image (e.g., the visible area of ​​the lips accounts for approximately 3.125% of the pixel percentage), when the image resolution of the scaled image is low, the model cannot accurately perceive changes in the semantic meaning of edge pixels in these parts due to the significant loss of detail in the visible areas. This results in low accuracy in predicting the visible areas of these parts, leading to poor recognition performance.

[0032] Based on the above research, in order to better improve the recognition effect, this application provides a recognition method, which includes: firstly acquiring a face image, such that the face presented in the face image includes target parts, such as lips and mouth; then transforming the face image to a first image space according to a first transformation matrix to obtain a first transformed image, so that the pixel proportion of the face in the first transformed image is as large as possible, thereby enabling the first transformed image to better describe the face presented in the face image with minimal loss of facial details; and finally transforming the face image to a second image space according to a second transformation matrix to obtain a second transformed image, so that the target parts... The proportion of pixels in the second transformed image is made as large as possible, so that the second transformed image can better describe the target part presented in the face image with as little loss of detail as possible. Then, based on the first transformed image, the first transformed matrix, the second transformed image, and the second transformed matrix, the visible region recognition result of the face and the visible region recognition result of the target part are determined, so that the visible region recognition result of the face is used to describe the position of the visible region of the face in the face image, and the visible region recognition result of the target part is used to describe the position of the visible region of the target part in the face image.

[0033] The first transformation matrix is ​​determined based on the facial key points of the face image and the key points of the first image space. The key points of the first image space are used to describe the cropping constraints of the face. This allows the first transformed image determined based on the first transformation matrix to better represent the cropping result of the face in the face image. As a result, the first transformed image contains little or no information that interferes with the recognition of the face, such as background information. This allows the first transformed image to better describe the face presented in the face image with as little loss of facial details as possible. Thus, the visible area of ​​the face determined based on the first transformed image is more accurate.

[0034] Furthermore, since the second transformation matrix is ​​determined based on the key points of the target region and the key points of the second image space, which are used to describe the cropping constraints of the target region, the second transformed image determined based on the second transformation matrix can better represent the cropping result of the target region in the face image. This results in the second transformed image containing a small amount or no information that interferes with the recognition of the target region, such as background information, other facial information, etc. Also, since the pixel ratio of the target region in the second transformed image is higher than that in the first transformed image, the target region details described by the second transformed image are more than those described by the first transformed image, thus making the visible area of ​​the target region determined based on the second transformed image more accurate.

[0035] As can be seen, this application achieves visible area recognition processing by using different image spaces for different parts. This effectively avoids the defects caused by excessive loss of details in some parts when using the same image space to achieve visible area recognition processing for different parts, thus improving the recognition effect.

[0036] Furthermore, this application does not limit the executing entity of the identification method provided in the embodiments of this application. For example, the identification method provided in the embodiments of this application can be applied to devices with limited computing power, such as terminal devices. The terminal device can be a smartphone, computer, personal digital assistant (PDA), tablet computer, etc.

[0037] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0038] To better understand the technical solution provided in this application, the identification method provided in this application will be explained below with reference to some accompanying drawings. For example... Figure 1 As shown, the identification method provided in this application includes S1-S3 as described below. Wherein, the... Figure 1 A flowchart of an identification method provided in an embodiment of this application.

[0039] S1: Acquire a face image, which includes the target area.

[0040] Among them, facial images refer to images that include faces, such as... Figure 2 The image shown is a facial image. It can be seen that, in one possible implementation, the facial image can be an image with dimensions H1×W1, where H1 represents the height of the facial image, W1 represents the width of the facial image, and H1×W1 represents the image resolution of the facial image. It should be noted that this application does not limit the relationship between H1 and W1; for example, they may be the same, or they may be different.

[0041] In addition, the face image can at least satisfy the following constraints: the face in the face image includes the target region, and the size of the face is larger than the size of the target region, so that visible region recognition processing can be performed on the face and the target region subsequently. Here, the target region refers to the part of the face that needs to be identified as a visible region.

[0042] Furthermore, this application does not limit the implementation method of the target part. For example, to improve flexibility, the target part can be defined according to the actual application scenario. Example 1: In some scenarios, such as those focusing on the lips and mouth, the target part may include the lips and mouth, so that subsequent visible area recognition processing can be performed on at least the lips and mouth. Example 2: In some scenarios, such as those focusing on the mouth, the target part may include the mouth, so that subsequent visible area recognition processing can be performed on at least the mouth. Example 3: In some scenarios, such as those focusing on the eyes, the target part may include the eyes, so that subsequent visible area recognition processing can be performed on at least the eyes.

[0043] As can be seen, in one possible implementation, the target part mentioned above can be used to describe any facial feature, such as eyebrows, eyes, ears, nose, mouth and other facial features; and the target part can include at least one part, such as the lips and the mouth.

[0044] Furthermore, this application does not limit the implementation of the facial image described above. For example, in some scenarios, such as unobstructed scenarios, the facial image can be used to describe a face in an unobstructed state, such that the visible area of ​​the face described by the facial image includes the entire face, and the visible area of ​​the target part described by the facial image includes the entire target part. As another example, in some scenarios, such as occluded scenarios, the facial image can be used to describe a face in an occluded state, such that the visible area of ​​the face described by the facial image includes the unoccluded portion of the face, and the visible area of ​​the target part described by the facial image includes the unoccluded portion of the target part.

[0045] It should be noted that the occlusion described above refers to a phenomenon caused by an opaque object existing between the face and the camera during the capture of a facial image. This obstructs the view of the face in the captured image, preventing the obscured portion, such as part of the face or mouth, from being shown, resulting in an incomplete view of the face and / or target area. Furthermore, this application does not limit the implementation of this object; for example, it can be a hand, hair, or sunglasses.

[0046] Furthermore, this application does not limit the method of obtaining the facial images mentioned above. For example, it can be implemented using any image acquisition method, such as taking pictures with a camera.

[0047] Based on the relevant content of S1 above, in some scenarios, such as facial effects scenarios, for terminal devices with limited computing power, the terminal device receives a facial image provided by the user through some input device, so that the terminal device can subsequently process the facial image to obtain the visible area of ​​the face and the visible area of ​​the target part.

[0048] S2: Transform the face image to a first image space according to the first transformation matrix to obtain a first transformed image, and transform the face image to a second image space according to the second transformation matrix to obtain a second transformed image. The first transformation matrix is ​​determined based on the facial key points of the face image and the key points of the first image space. The key points of the first image space are used to describe the cropping constraints of the face. The facial key points include the key points of the target part. The second transformation matrix is ​​determined based on the key points of the target part and the key points of the second image space. The key points of the second image space are used to describe the cropping constraints of the target part. The pixel ratio of the target part in the second transformed image is higher than the pixel ratio of the target part in the first transformed image.

[0049] Here, the first image space refers to the image space used when performing face cropping processing on an image, so that the first image space is used to describe the constraints that need to be satisfied during the face cropping processing, such as the screenshot obtained through the face cropping processing (e.g., Figure 2 The constraints that the screenshot of the face shown needs to satisfy are such that the first image space can describe some characteristics of the face in the screenshot, such as the face being located in the center of the screenshot, the aspect ratio of the screenshot being a preset fixed value, the roll angle of the face in the screenshot being 0, and the pixel proportion of the face region in the screenshot being greater than a preset threshold (such as 80%). It should be noted that the fixed value can be determined according to the actual application scenario; for example, the fixed value can be 1 or 4 / 3.

[0050] As can be seen, in one possible implementation, the first image space mentioned above can be used to describe face cropping constraints. These cropping constraints refer to the constraints that need to be satisfied when performing face cropping processing on an image, so that the cropping constraints can represent the constraints that need to be satisfied when mapping the face in the image to the first image space. Furthermore, this application does not limit the cropping constraints. For example, for a screenshot obtained through the face cropping processing, the cropping constraints may include at least one of the following constraints: the face is located at the center of the screenshot; the aspect ratio of the screenshot is a preset fixed value; the Roll angle of the face in the screenshot is 0; the pixel proportion of the face region in the screenshot is greater than a preset threshold.

[0051] Based on the above two paragraphs, when the first image space mentioned above is used as the image space of the screenshot obtained through face cropping processing, the first image space can at least satisfy the following constraints: pixels located at or near the center of the first image space are used to describe the face (that is, the center pixel of the first image space and its neighboring pixels are used to describe the face); the aspect ratio of the first image space is a preset fixed value (e.g., 1 or 4 / 3); the roll angle of the face presented by the first image space is 0; the ratio between the number of pixels in the first image space used to describe the face and the number of pixels in the first image space is greater than a preset threshold (e.g., 80%).

[0052] In addition, in some scenarios, such as applications with low computational requirements, the first image space mentioned above may also satisfy the following constraint: the image resolution in the first image space does not exceed the maximum image resolution set in advance for the application scenario, so that the image with the first image space can meet the image processing requirements of the application scenario.

[0053] Furthermore, this application does not limit the implementation of the first image space described above. For example, it can be implemented using a pre-defined image space of an average face, so that the first image space can represent the characteristics of the average face, such as the characteristics described by the face cropping constraints described above.

[0054] Furthermore, this application does not limit the representation of the first image space described above; for example, it can be represented using a two-dimensional image. Thus, in one possible implementation, the first image space can be implemented using a pre-defined two-dimensional image of an average face. Since the average face in this two-dimensional image satisfies the characteristics described by the face cropping constraints above, the screenshot obtained by performing face cropping processing using this two-dimensional image also satisfies these characteristics, thereby enabling the screenshot to better represent the facial state.

[0055] The first transformation matrix refers to the transformation matrix used when projecting a face image onto a first image space, such as an affine transformation matrix; and the first transformation matrix is ​​used to describe the correspondence between some or all of the pixels in the face image (e.g., pixels used to describe the face) and the corresponding pixels in the first image space.

[0056] Furthermore, this application does not limit the method of obtaining the first transformation matrix mentioned above. For example, in order to improve flexibility, the first transformation matrix can be determined based on the facial key points of the face image and the key points of the first image space, so that the first transformation matrix can at least describe the correspondence between the key points used to describe the face in the face image and the key points used to describe the face in the first image space. This allows the first transformation matrix to represent the alignment between the key points in the image space of the face image and the key points in the first image space, and further allows the first transformation matrix to represent the alignment between the pixels in the image space of the face image and the pixels in the first image space.

[0057] Regarding the facial key points in the facial image above, such as Figure 2 The facial key points shown are used to describe the facial state presented in the facial image, such as the location of the facial region, the shape of the face, and the roll angle of the face. Moreover, this application does not limit the method of obtaining the facial key points. For example, it can be implemented using any existing or future facial key point detection method.

[0058] Regarding the key points in the first image space mentioned above, these key points are used to describe the facial state presented in the first image space, such as the facial state of an average face, so that the key points in the first image space can describe the face cropping constraints mentioned above. Furthermore, this application does not limit the implementation method of the key points in the first image space. For example, the key points in the first image space may include key points of the face (such as an average face) in the first image space. Additionally, this application does not limit the method of obtaining the key points in the first image space. For example, it can be implemented using a manual annotation method. Alternatively, it can be implemented using any existing or future facial key point detection method.

[0059] Based on the above three paragraphs, it can be seen that, in one possible implementation, the process of determining the first transformation matrix can be as follows: using a preset algorithm, such as the least squares method or any optimal solution search method, to calculate the transformation matrix that minimizes the error between the facial key points of the face image and the key points of the first image space, and use this as the first transformation matrix, so that the similarity between the state of the face in the result obtained by transforming the face image according to the first transformation matrix and the state of the face described by the first image space is maximized, thereby enabling the face described by the result to satisfy the cropping constraints of the face described by the first image space as much as possible, and thus making the face cropping processing based on the first transformation matrix more effective.

[0060] The first transformed image refers to the result obtained by transforming a face image according to a first transformation matrix, so that the first transformed image can represent the state of the face described by the face image in a first image space. This ensures that the face described by the first transformed image satisfies the cropping constraints of the face described by the first image space as much as possible, thereby enabling the first transformed image to better represent the result of face cropping processing on the face image. Figure 2 The screenshot shown is of the face.

[0061] In addition, in some scenarios, such as applications with low computational requirements, the first transformed image mentioned above can at least satisfy the following constraints: the image resolution of the first transformed image is less than the image resolution of the face image, so that the height of the first transformed image is less than the height of the face image, and the width of the first transformed image is less than the width of the face image, thereby enabling the first transformed image to include some information from the face image.

[0062] Research has found that because the height of most faces is greater than their width, when the screenshot obtained by face cropping is square, there are large background areas on the left and right sides of the face in the screenshot. This results in a lot of background information in the screenshot, which in turn causes a lot of interference.

[0063] Based on the above research, in order to better reduce the interference caused by background information, this application also provides a possible implementation of the first image space mentioned above. In this implementation, the first image space can at least satisfy the following constraints: the width of the first image space is less than the height of the first image space, so that the width of the first transformed image obtained by transforming the face image to the first image space is also less than the height of the first transformed image, thereby making the left and right sides of the face in the first transformed image have no or a small background area, thus effectively reducing the interference caused by background information.

[0064] The study also found that for screenshots obtained by cropping faces, if the screenshot is square, when the height of the screenshot is similar to the height of the face in the screenshot, the ratio between the width of the face and the width of the screenshot is between 60% and 65%.

[0065] Based on the above research, in order to better reduce the interference caused by background information, this application also provides a possible implementation of the first image space mentioned above. In this implementation, the first image space can at least satisfy the following constraints: the height of the first image space is slightly higher than the height of the face in the first image space, and the aspect ratio of the first image space is 4 / 3, so that the height of the first transformed image obtained by transforming the face image to the first image space is similar to the height of the face in the first transformed image, and the aspect ratio of the first transformed image is 4 / 3, thereby making the first transformed image contain less background area, which can effectively reduce the interference caused by background information.

[0066] As can be seen, in one possible implementation, when the height of the first transformed image is H2, the width of the first transformed image is 3 / 4 times H2. This can effectively reduce the pixel proportion of background information in the first transformed image, thereby effectively increasing the pixel proportion of the face in the first transformed image, and thus effectively reducing the interference caused by background information.

[0067] Based on the above content regarding the first transformed image, in some scenarios, after acquiring a face image, facial keypoint detection is first performed on the face image to obtain the facial keypoints, so that these keypoints can describe the facial state presented in the face image. Then, these facial keypoints are aligned with keypoints in a pre-defined first image space (such as an average face) to obtain a first transformation matrix. This minimizes the error between the keypoints obtained by mapping the facial keypoints to the first image space according to the first transformation matrix and the keypoints in the first image space, thereby enabling the first transformation matrix to represent the face image. The alignment between pixels in the image space and pixels in the first image space is determined. Then, the first transformation matrix is ​​regarded as an affine transformation matrix, and the face image is transformed to the first image space through affine transformation to obtain a first transformed image. This allows the first transformed image to display the face according to the cropping constraints (such as face size, face display position, etc.) described by the first image space. This enables the first transformed image to better represent the face region cropped from the face image, and the background information in the first transformed image is much less than the background information in the face image, thus effectively avoiding interference caused by background information.

[0068] The second image space refers to the image space used when performing target region cropping on an image. This second image space describes the constraints that need to be satisfied during the target region cropping process, such as the screenshot obtained through the target region cropping process (e.g., ...). Figure 2 The screenshot of the mouth shown needs to meet certain constraints, such as being able to describe some characteristics of the target area in the screenshot, such as the target area being located at the center of the screenshot, the aspect ratio of the screenshot being a preset fixed value, the Roll angle of the target area in the screenshot being 0, and the pixel proportion of the target area in the screenshot being greater than a preset threshold. It should be noted that this fixed value can be determined according to the actual application scenario; for example, the fixed value can be 2 / 3.

[0069] As can be seen, in one possible implementation, the second image space described above can be used to describe the cropping constraints of the target region. These cropping constraints refer to the constraints that must be satisfied when cropping a target region from an image, so that the cropping constraints can represent the constraints that must be satisfied when mapping the target region in the image to the second image space. Furthermore, this application does not limit the cropping constraints. For example, for a screenshot obtained through the target region cropping process, the cropping constraints may include at least one of the following constraints: the target region is located at the center of the screenshot; the aspect ratio of the screenshot is a preset fixed value; the Roll angle of the target region in the screenshot is 0; and the pixel proportion of the target region in the screenshot is greater than a preset threshold.

[0070] Based on the above two paragraphs, when the second image space mentioned above is used as the image space of the screenshot obtained by the target part cropping process, the second image space can at least satisfy the following constraints: pixels located at or near the center of the second image space are used to describe the target part (that is, the center pixel of the second image space and its neighboring pixels are used to describe the target part); the aspect ratio of the second image space is a preset fixed value (e.g., 2 / 3); the roll angle of the target part presented by the second image space is 0; the ratio between the number of pixels in the second image space used to describe the target part and the number of pixels in the second image space is greater than a preset threshold (e.g., 80%).

[0071] In addition, in some scenarios, such as applications with low computational requirements, the second image space mentioned above can also satisfy the following constraint: the image resolution in the second image space does not exceed the maximum image resolution set in advance for the application scenario, so that the image with the second image space can meet the image processing requirements of the application scenario.

[0072] Furthermore, this application does not limit the implementation of the second image space described above. For example, it can be implemented using a pre-defined image space of an average target part (such as an average mouth) so that the second image space can represent the characteristics of the average target part, such as the characteristics described by the cropping constraints of the target part described above.

[0073] Furthermore, this application does not limit the representation of the second image space described above; for example, it can be represented using a two-dimensional image. Therefore, in one possible implementation, the second image space can be a two-dimensional image of a pre-defined average target region. Since the average target region in this two-dimensional image satisfies the characteristics described by the target region cropping constraints described above, the screenshot obtained by cropping the target region using this two-dimensional image also satisfies these characteristics, thereby enabling the screenshot to better represent the state of the target region.

[0074] The second transformation matrix refers to the transformation matrix used when projecting a face image onto a second image space, such as an affine transformation matrix; and the second transformation matrix is ​​used to describe the correspondence between some or all of the pixels in the face image (e.g., pixels used to describe the target area) and the corresponding pixels in the second image space.

[0075] Furthermore, this application does not limit the method of obtaining the second transformation matrix mentioned above. For example, in order to improve flexibility, when the facial key points of the facial image mentioned above include key points of the target part, the second transformation matrix can be determined based on the key points of the target part and the key points in the second image space, so that the second transformation matrix can at least describe the correspondence between the key points in the facial image used to describe the target part and the key points in the second image space used to describe the target part. This allows the second transformation matrix to represent the alignment between some key points in the image space of the facial image and the corresponding key points in the second image space, and further allows the second transformation matrix to represent the alignment between pixels in the image space of the facial image and pixels in the second image space.

[0076] Regarding the key points of the target area mentioned above, the key points of the target area refer to the key points existing in the facial key points of the face image that are used to describe the target area, so that the key points of the target area can describe the state of the target area presented in the face image.

[0077] Regarding the key points in the second image space mentioned above, these key points describe the state of the target region within the second image space, enabling them to describe the truncating constraints of the target region. Furthermore, this application does not limit the implementation method of the key points in the second image space. For example, the key points in the second image space may include key points of the target region (such as the average mouth). Additionally, this application does not limit the method of obtaining the key points in the second image space. For example, it can be implemented using manual annotation methods. Alternatively, it can be implemented using any existing or future key point detection method.

[0078] Based on the above three paragraphs, it can be seen that in one possible implementation, when the facial key points of the above face image include the key points of the target part, the process of determining the second transformation matrix can be as follows: using a preset algorithm, such as the least squares method or any optimal solution search method, to calculate the transformation matrix that minimizes the error between the key points of the target part and the key points of the second image space, and use it as the second transformation matrix, so that the similarity between the state of the target part in the result obtained by transforming the face image according to the second transformation matrix and the state of the target part described by the second image space is maximized, so that the target part described by the result can satisfy the cropping constraint of the target part described by the second image space as much as possible, thereby making the target part cropping processing based on the second transformation matrix more effective.

[0079] The second transformed image refers to the result obtained by transforming the face image according to the second transformation matrix, so that the second transformed image can represent the state of the target part described by the face image in the second image space. This ensures that the target part described by the second transformed image satisfies the cropping constraints of the target part described by the second image space as much as possible, thereby enabling the second transformed image to better represent the result of target part cropping processing on the face image. Figure 2 The screenshot shown is of the mouth.

[0080] In addition, in some scenarios, such as applications with low computational requirements, the second transformed image mentioned above can at least satisfy the following constraints: the image resolution of the second transformed image is less than the image resolution of the face image, so that the height of the second transformed image is less than the height of the face image, and the width of the second transformed image is less than the width of the face image, thereby enabling the second transformed image to include some information from the face image.

[0081] Research has found that if, in some forms, the width of most target parts is greater than the height of the target parts, then in order to better reduce the interference caused by background information, this application also provides a possible implementation of the second image space mentioned above. In this method, the second image space can at least satisfy the following constraints: the width of the second image space is determined based on the width of the target parts, and the height of the second image space is determined based on the maximum height of the target parts (such as the height of the mouth when it is open), so that the aspect ratio of the second image space is determined based on the ratio between the maximum height of the target parts and the width of the target parts. This ensures that the second transformed image obtained by transforming the face image to the second image space contains as few pixels as possible for describing background information, thus effectively reducing the interference caused by background information.

[0082] For the maximum height of the target part, this maximum height refers to the maximum height that the target part can reach. Furthermore, when the target part includes multiple forms, such as different degrees of mouth opening, this maximum height satisfies the following constraints: the heights of different forms are different, and the height of each form does not exceed the maximum height of the target part. Therefore, the maximum height of the target part can be obtained through maximum value analysis of the heights of each form of the target part.

[0083] The study also found that, in order to better reduce the overhead of computing resources, the first transformed image and the second transformed image mentioned above can be stitched together along the Y-axis, such as... Figure 2 The vertical stitching shown is designed to minimize the number of useless pixels, such as blank pixels or background pixels, in the resulting image, thereby minimizing the computational resources consumed when processing the stitched image.

[0084] Based on the above research, in order to better reduce resource consumption, this application also provides a possible implementation of the second image space mentioned above. In this implementation, the second image space can at least satisfy the following constraint: the width of the second image space is the same as the width of the first image space mentioned above, so that the width of the second transformed image obtained by transforming to the second image space is the same as the width of the first transformed image obtained by transforming to the first image space, thereby enabling the second transformed image and the first transformed image to be stitched together along the Y-axis direction, which is beneficial to reducing resource consumption.

[0085] The study also found that if, in some forms, the width of most target parts is greater than the height of the target parts, then the relative ratio between the maximum height of the target parts and the height of the face including the target parts is almost constant.

[0086] Based on the above research, in order to better reduce interference caused by background information, this application also provides a possible implementation of the second image space mentioned above. In this implementation, the second image space can at least satisfy the following constraint: the ratio between the height of the second image space and the height of the first image space mentioned above is a preset ratio, such as 1 / 2, so that the ratio between the height of the second transformed image obtained by transforming to the second image space and the height of the first transformed image obtained by transforming to the first image space is the preset ratio. This preset ratio is obtained by analyzing a large number of facial images, so that the preset ratio can better describe the height correlation between the face and the target area.

[0087] It should be noted that, for the preset ratio shown in the above paragraph, when the width of the second image space above is the same as the width of the first image space above, and the ratio between the height of the second image space and the height of the first image space is the preset ratio, the preset ratio is used to represent the relative ratio between the height of the target part and the height of the face including the target part when the target part is enlarged to the width of the face, such as a relative ratio of 1:2.

[0088] Based on the above two paragraphs, it can be understood that, in one possible implementation, the ratio between the height of the first transformed image and the height of the second transformed image can be determined based on the ratio between the height of the face and the maximum height of the target part, so that the second transformed image can independently represent the target part presented in the first transformed image. This allows the second transformed image to represent a local magnified result of the target part in the first transformed image, effectively overcoming the defect caused by the small pixel proportion of the target part in the first transformed image, thereby improving the recognition effect.

[0089] Based on the above content regarding the second transformed image, for some scenarios, after acquiring a face image, facial keypoint detection is first performed on the face image to obtain the facial keypoints, ensuring that these keypoints include keypoints of various parts of the face, such as the keypoints of the target area. Then, the keypoints of the target area are aligned with the keypoints of a pre-defined second image space (such as the average mouth area) to obtain a second transformation matrix. This minimizes the error between the keypoints obtained by mapping the keypoints of the target area to the second image space using the second transformation matrix and the keypoints in the second image space, thus enabling the second transformation matrix to represent the image space of the face image. The alignment between the intermediate pixels and the pixels in the second image space is determined. Then, the second transformation matrix is ​​treated as an affine transformation matrix, and the face image is transformed to the second image space through an affine transformation to obtain a second transformed image. This allows the second transformed image to display the target part according to the cropping constraints (such as mouth size, mouth display position, etc.) described by the second image space. As a result, the second transformed image can represent the target part region cropped from the face image. Consequently, the second transformed image contains far less information other than the target part information than the face image, thus effectively avoiding interference caused by other information.

[0090] Based on the content related to S2 above, for some scenarios, if the target area includes the lips and mouth, or if the target area includes the mouth, then after acquiring the face image, facial keypoint detection processing is first performed on the face image to obtain the facial keypoints of the face image, so that the facial keypoints include the keypoints of various parts of the face, such as the keypoints of the mouth; then, by aligning the facial keypoints with the pre-set average facial keypoints, face cropping processing is achieved on the face image, such as... Figure 2 The facial region cropping process shown yields a first transformed image, which represents the facial region cropped from the facial image. This allows the first transformed image to describe the facial features presented in the facial image with minimal interference. Furthermore, mouth cropping is achieved by aligning the key points of the mouth with pre-defined average mouth key points. Figure 2 The mouth region shown is cropped to obtain a second transformed image, which is able to represent the mouth region cropped from the face image, thereby enabling the second transformed image to describe the mouth features presented in the face image with as little interference as possible.

[0091] S3: Based on the first transformed image, the first transformed matrix, the second transformed image, and the second transformed matrix, determine the visible region recognition result of the face and the visible region recognition result of the target part.

[0092] The visible region recognition result of the face is used to describe the location of the visible region of the face in the face image, so that the visible region recognition result can indicate which pixels in the face image are used to describe the visible region of the face.

[0093] The visible region identification result of the target part is used to describe the location of the visible region of the target part in the face image, so that the visible region identification result can indicate which pixels in the face image are used to describe the visible region of the target part.

[0094] Therefore, in one possible implementation, when the target area includes the lips and the mouth, the visible area recognition result of the target area includes the visible area recognition result of the lips and the visible area recognition result of the mouth. Specifically, the visible area recognition result of the lips describes the position of the visible area of ​​the lips in the face image, so that the visible area recognition result of the lips can indicate which pixels in the face image are used to describe the visible area of ​​the lips. The visible area recognition result of the mouth describes the position of the visible area of ​​the mouth in the face image, so that the visible area recognition result of the mouth can indicate which pixels in the face image are used to describe the visible area of ​​the mouth.

[0095] Furthermore, this application does not limit the implementation of S3 above. For example, it can be implemented using a pre-built machine learning model with region recognition function.

[0096] In addition, to further improve the recognition effect, this application also provides a possible implementation of S3 above. In this implementation, S3 may specifically include: transforming the region prediction result of the first transformed image back to the image space of the face image according to the inverse transformation matrix corresponding to the first transformation matrix to obtain the visible region recognition result of the face; and transforming the region prediction result of the second transformed image back to the image space of the face image according to the inverse transformation matrix corresponding to the second transformation matrix to obtain the visible region recognition result of the target part.

[0097] For the inverse transformation matrix corresponding to the first transformation matrix mentioned above, the transformation process described by the inverse transformation matrix is ​​the inverse of the transformation process described by the first transformation matrix, so that the inverse transformation matrix can represent the transformation matrix required when transforming an image from the first image space back to the original image space (such as the image space of the face image mentioned above). The inverse transformation matrix is ​​determined by inverse reasoning based on the first transformation matrix, and this application does not limit the method of obtaining the inverse transformation matrix; for example, it can be implemented using any existing or future method for determining the inverse affine transformation matrix.

[0098] Regarding the region prediction result of the first transformed image mentioned above, this region prediction result is used to describe the location of the visible region of the face in the first transformed image, so that the region prediction result can indicate which pixels in the first transformed image are used to describe the visible region of the face. Furthermore, this application does not limit the representation method of the region prediction result; for example, it can be represented using a binary mask. Moreover, this application does not limit the acquisition method of the region prediction result; for example, it can be implemented using any machine learning model with facial visible region prediction functionality. Furthermore, this application does not limit the implementation method of the machine learning model; for example, the machine learning model can at least include a feature extractor and a feature decoder arranged sequentially.

[0099] For the inverse transformation matrix corresponding to the second transformation matrix mentioned above, the transformation process described by the inverse transformation matrix is ​​the inverse of the transformation process described by the second transformation matrix, so that the inverse transformation matrix can represent the transformation matrix required when transforming an image from the second image space back to the original image space (such as the image space of the face image mentioned above). The inverse transformation matrix is ​​determined by inverse reasoning based on the second transformation matrix, and this application does not limit the method of obtaining the inverse transformation matrix; for example, it can be implemented using any existing or future method for determining the inverse affine transformation matrix.

[0100] Regarding the region prediction result of the second transformed image mentioned above, this region prediction result is used to describe the position of the visible region of the target part in the second transformed image, so that the region prediction result can indicate which pixels in the second transformed image are used to describe the visible region of the target part. Furthermore, this application does not limit the representation method of the region prediction result; for example, it can be represented using a binary mask. Moreover, this application does not limit the acquisition method of the region prediction result; for example, it can be implemented using any machine learning model with mouth visible region (or lip visible region and oral cavity visible region) prediction function. Also, this application does not limit the implementation method of the machine learning model; for example, the machine learning model can at least include a feature extractor and a feature decoder arranged sequentially.

[0101] Based on the above five paragraphs, it can be seen that in some scenarios, the visible region recognition result of the face can be determined based on the region prediction result of the first transformation matrix and the first transformed image, and the visible region recognition result of the target part can be determined based on the region prediction result of the second transformation matrix and the second transformed image, so that the two determination processes do not interfere with each other, which helps to reduce the computational resource overhead.

[0102] Furthermore, this application does not limit the process of determining the region prediction results mentioned above. For example, the region prediction result of the first transformed image can be obtained by a machine learning model performing face region detection processing on the first transformed image, and the region prediction result of the second transformed image can be obtained by another machine learning model performing target region detection processing on the second transformed image, so that the two detection processes are independent, which is beneficial to improving the detection effect.

[0103] In addition, in order to better reduce resource consumption, S3 above may specifically include steps 11-13 below.

[0104] Step 11: Stitch the first transformed image and the second transformed image together to obtain the stitched image.

[0105] It should be noted that this application does not limit the implementation of step 11 above. For example, it can be implemented using any image stitching method.

[0106] For example, in some scenarios, such as when the width of the target part is greater than its height in certain shapes, in order to better reduce the amount of computation, when the width of the first transformed image is the same as the width of the second transformed image, step 11 above can specifically be: stitching the first transformed image and the second transformed image along the height direction to obtain a stitched image, so that the height of the stitched image is equal to the sum of the heights of the first transformed image and the second transformed image, thereby minimizing the amount of useless information, such as background information, in the stitched image, thus effectively reducing the interference caused by the useless information.

[0107] As can be seen, when the height of the first transformed image above is H2, the width of the first transformed image is H2×3 / 4, the height of the second transformed image is H3 (for example, H3 = H2 / 2), and the width of the second transformed image is H2×3 / 4, the height of the stitched image above is H2+H3, and the width of the stitched image is H2×3 / 4.

[0108] For example, in some scenarios, such as when the size of the stitched image does not match the size of the machine learning model with region prediction function, step 11 above can be specifically as follows: first, stitch the first transformed image and the second transformed image along the height direction to obtain the stitched image; then, scale the stitched image according to the size requirements of the model.

[0109] Based on the above three paragraphs, it can be seen that when the width of the first transformed image is the same as the width of the second transformed image, the height of the stitched image is determined by the sum of the heights of the first transformed image and the second transformed image.

[0110] Step 12: Based on the stitched image above, determine the region prediction result of the first transformed image and the region prediction result of the second transformed image. The region prediction result of the first transformed image is used to describe the position of the visible area of ​​the face in the first transformed image, and the region prediction result of the second transformed image is used to describe the position of the visible area of ​​the target part in the second transformed image.

[0111] It should be noted that this application does not limit the implementation of step 12 above. For example, it can be implemented using any machine learning model with regional prediction function.

[0112] For example, step 12 above can specifically be: performing image segmentation processing on the stitched image above to obtain the region prediction results of the first transformed image and the second transformed image. This image segmentation processing is used to determine the location of different regions in an image, such as the location of the visible area of ​​the face, the visible area of ​​the lips, and the visible area of ​​the mouth.

[0113] It should be noted that this application does not limit the implementation method of the image segmentation process. For example, it can be implemented using any existing or future machine learning model with image segmentation capabilities. Furthermore, this application does not limit the implementation method of the machine learning model. For example, the machine learning model may include a feature extractor, a feature decoder, a segmentation module, and an upsampling classification module arranged in sequence.

[0114] As can be seen, in one possible implementation, when step 12 above is implemented using a pre-built machine learning model with image segmentation or region prediction capabilities, if the model includes a feature extractor, a feature decoder, a segmentation module, and an upsampling classification module arranged sequentially, then step 12 can specifically be as follows: After inputting the stitched image above into the machine learning model, the feature extractor in the model first performs feature extraction processing on the stitched image to obtain the feature extraction result; then, the feature decoder in the model performs feature decoding processing on the feature extraction result to obtain the feature decoding result; finally, the segmentation module in the model segments the feature decoding result to obtain the first feature map (e.g., ...). Figure 2 Part 1 shown) and the second feature map (e.g.) Figure 2 Part 2 shown is used to make the first feature map describe facial features and the second feature map describe target area features; finally, the upsampling classification module in the model upsamples and classifies the first feature map to obtain the region prediction result of the first transformed image (e.g., Figure 2 The face mask shown is used, and the upsampling classification module in the model upsamples and classifies the second feature map to obtain the region prediction result of the second transformed image (e.g., the face mask image shown). Figure 2 (The mouth mask shown).

[0115] As can be seen, in some scenarios, if the height of the stitched image is H2+H3 and the width of the stitched image is H2×3 / 4, then some of the data mentioned above has the following characteristics: the height of the feature decoding result is (H2+H3) / s, and the width of the feature decoding result is (H2×3) / (4×s), where s represents the downsampling rate of the shallowest feature layer of the feature decoder; by horizontally segmenting the feature decoding result at y=H2 / s along the Y-axis, the first feature map and The second feature map is configured such that the height of the first feature map is H2 / s, the width of the first feature map is (H2×3) / (4×s), the height of the second feature map is H3 / s, and the width of the second feature map is (H2×3) / (4×s); the height of the region prediction result of the first transformed image is H2, and the width of the region prediction result of the first transformed image is H2×3 / 4; the height of the region prediction result of the second transformed image is H3, and the width of the region prediction result of the second transformed image is H2×3 / 4.

[0116] Step 13: Based on the region prediction results of the first transformed image, the first transformation matrix, the region prediction results of the second transformed image, and the second transformation matrix, determine the visible region recognition results of the face and the visible region recognition results of the target part.

[0117] It should be noted that the relevant content of step 13 above can be found in the previous text.

[0118] Based on the content of steps 11 to 13 above, it is known that for some scenarios, such as applications with low computational requirements, after obtaining the first transformed image and the second transformed image, these two images can be stitched together to obtain the stitched result, such as... Figure 2 The stitching result shown is designed to describe the facial information carried by the first transformed image and the target area information carried by the second transformed image with as few pixels as possible; then the stitching result is input into a machine learning model with region prediction capabilities to obtain a first mask image output by the model (e.g., ...). Figure 2 The face mask shown) and the second mask (as shown) Figure 2 The first mask (shown as an example) is used to represent the position of the visible area of ​​the face in the first transformed image, and the second mask is used to represent the position of the visible area of ​​the target part in the second transformed image. Then, according to the inverse transformation matrix corresponding to the first transformation matrix, the first mask is transformed back to the image space of the face image to obtain the visible area recognition result of the face. And according to the inverse transformation matrix corresponding to the second transformation matrix, the second mask is transformed back to the image space of the face image to obtain the visible area recognition result of the target part.

[0119] Based on the relevant content in S1 to S3 above, the recognition method provided in this application first acquires a face image, ensuring that the face presented in the face image includes target parts, such as lips and mouth; then, according to a first transformation matrix, the face image is transformed to a first image space to obtain a first transformed image, maximizing the pixel proportion of the face in the first transformed image, thereby enabling the first transformed image to better describe the face presented in the face image with minimal loss of facial details; and finally, according to a second transformation matrix, the face image is transformed to a second image space to obtain a second transformed image, maximizing the pixel proportion of the face in the second ... The second transformed image has the largest possible pixel proportion, so that the second transformed image can better describe the target part presented in the face image with minimal loss of detail of the target part; then, based on the first transformed image, the first transformed matrix, the second transformed image, and the second transformed matrix, the visible region recognition result of the face and the visible region recognition result of the target part are determined, so that the visible region recognition result of the face is used to describe the position of the visible region of the face in the face image, and the visible region recognition result of the target part is used to describe the position of the visible region of the target part in the face image.

[0120] As can be seen, this application achieves visible area recognition processing by using different image spaces for different parts. This effectively avoids the defects caused by excessive loss of details in some parts when using the same image space to achieve visible area recognition processing for different parts, thus improving the recognition effect.

[0121] Based on the aforementioned recognition methods, the technical solution provided in this application has the following advantages: ① Because this application utilizes an additional and independent image space to achieve visible area recognition processing for the target part, the technical solution provided in this application can achieve a higher level of recognition effect for the visible area of ​​the target part under the same or similar computational load conditions. ② Because the size of the image space involved in the face cropping processing of this application is determined based on the aspect ratio of the face, the amount of useless information (such as background information) that needs to be processed when using this image space to achieve visible area recognition processing for the face is reduced, thereby enabling the technical solution provided in this application to achieve a higher level of recognition effect for the visible area of ​​the face under the same or similar computational load conditions.

[0122] Based on the identification method provided in the embodiments of this application, the embodiments of this application also provide an identification device, which is described below in conjunction with... Figure 3 Explanation and clarification will be provided. Among them, Figure 3This is a schematic diagram of an identification device provided in an embodiment of this application. It should be noted that for technical details of the identification device provided in this embodiment, please refer to the relevant content of the identification method above.

[0123] like Figure 3 As shown, the identification device 300 provided in this application embodiment includes:

[0124] The acquisition unit 301 is used to acquire a face image, wherein the face includes a target region;

[0125] Transformation unit 302 is configured to transform the face image to a first image space according to a first transformation matrix to obtain a first transformed image, and to transform the face image to a second image space according to a second transformation matrix to obtain a second transformed image. The first transformation matrix is ​​determined based on the facial key points of the face image and the key points of the first image space. The key points of the first image space are used to describe the cropping constraints of the face, and the facial key points include the key points of the target part. The second transformation matrix is ​​determined based on the key points of the target part and the key points of the second image space. The key points of the second image space are used to describe the cropping constraints of the target part. The pixel ratio of the target part in the second transformed image is higher than the pixel ratio of the target part in the first transformed image.

[0126] The determining unit 303 is used to determine the visible region recognition result of the face and the visible region recognition result of the target part based on the first transformed image, the first transformed matrix, the second transformed image and the second transformed matrix.

[0127] In one possible implementation, the visible region identification result of the face is obtained by transforming the region prediction result of the first transformed image back to the image space of the face image according to the inverse transformation matrix corresponding to the first transformation matrix. The region prediction result of the first transformed image is used to describe the position of the visible region of the face in the first transformed image. The visible region identification result of the target part is obtained by transforming the region prediction result of the second transformed image back to the image space of the face image according to the inverse transformation matrix corresponding to the second transformation matrix. The region prediction result of the second transformed image is used to describe the position of the visible region of the target part in the second transformed image.

[0128] In one possible implementation, the determining unit 303 is specifically configured to: stitch the first transformed image and the second transformed image together to obtain a stitched image; determine the region prediction result of the first transformed image and the region prediction result of the second transformed image based on the stitched image, wherein the region prediction result of the first transformed image is used to describe the position of the visible region of the face in the first transformed image, and the region prediction result of the second transformed image is used to describe the position of the visible region of the target part in the second transformed image; and determine the visible region recognition result of the face and the visible region recognition result of the target part based on the region prediction result of the first transformed image, the first transformation matrix, the region prediction result of the second transformed image, and the second transformation matrix.

[0129] In one possible implementation, the width of the first transformed image is the same as the width of the second transformed image; the height of the stitched image is determined based on the sum of the heights of the first transformed image and the second transformed image.

[0130] In one possible implementation, the width of the first transformed image is smaller than the height of the first transformed image.

[0131] In one possible implementation, the ratio between the height of the first transformed image and the height of the second transformed image is determined based on the ratio between the height of the face and the maximum height of the target part. The target part includes multiple shapes, each with a different height, and the height of each shape is not greater than the maximum height of the target part.

[0132] In one possible implementation, the target area includes the lips and the oral cavity, and the visible area recognition result of the target area includes the visible area recognition result of the lips and the visible area recognition result of the oral cavity.

[0133] In one possible implementation, the identification device 300 is deployed on a terminal device.

[0134] Based on the aforementioned content of the recognition device 300, the working principle of the recognition device 300 provided in this application includes: firstly, acquiring a face image so that the face presented in the face image includes target parts, such as lips and mouth; then, transforming the face image to a first image space according to a first transformation matrix to obtain a first transformed image, so that the pixel proportion of the face in the first transformed image is as large as possible, thereby enabling the first transformed image to better describe the face presented in the face image with minimal loss of facial details; and finally, transforming the face image to a second image space according to a second transformation matrix to obtain a second transformed image, so that the target... The pixel proportion of the target area in the second transformed image is maximized, so that the second transformed image can better describe the target area presented in the face image with minimal loss of detail. Then, based on the first transformed image, the first transformation matrix, the second transformed image, and the second transformation matrix, the visible region recognition result of the face and the visible region recognition result of the target area are determined. The visible region recognition result of the face is used to describe the position of the visible region of the face in the face image, and the visible region recognition result of the target area is used to describe the position of the visible region of the target area in the face image. Therefore, the recognition device 300 effectively avoids the defects caused by excessive loss of detail in some areas when using the same image space to perform visible region recognition processing for different areas, thus improving the recognition effect.

[0135] In addition, this application also provides an electronic device, the device including a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device performs any implementation of the identification method provided in this application.

[0136] See Figure 4 This diagram illustrates a structural schematic of an electronic device 400 suitable for implementing embodiments of the present disclosure. The terminal devices in these embodiments may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0137] like Figure 4As shown, electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 402 or a program loaded from storage device 408 into random access memory (RAM) 403. RAM 403 also stores various programs and data required for the operation of electronic device 400. The processing device 401, ROM 402, and RAM 403 are interconnected via bus 404. Input / output (I / O) interface 405 is also connected to bus 404.

[0138] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 408 including, for example, magnetic tapes, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 An electronic device 400 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0139] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 409, or installed from a storage device 408, or installed from a ROM 402. When the computer program is executed by the processing device 401, it performs the functions defined in the methods of embodiments of this disclosure.

[0140] The electronic device provided in this embodiment belongs to the same inventive concept as the method provided in the above embodiments. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0141] This application also provides a computer-readable medium storing instructions or a computer program that, when executed on a device, causes the device to perform any implementation of the identification method provided in this application.

[0142] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0143] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0144] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0145] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, enable the electronic device to perform the aforementioned methods.

[0146] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0147] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0148] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units / modules do not necessarily limit the specific unit itself.

[0149] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0150] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0151] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems or apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and relevant parts can be referred to the method section.

[0152] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0153] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0154] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0155] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method of identification, characterized in that, The method comprises: obtaining a face image, the face comprising a target part; transforming the face image to a first image space according to a first transformation matrix to obtain a first transformed image, and transforming the face image to a second image space according to a second transformation matrix to obtain a second transformed image, the first transformation matrix being determined according to face key points of the face image and key points of the first image space, the key points of the first image space being used to describe the interception constraint of the face, the face key points comprising key points of the target part, the second transformation matrix being determined according to the key points of the target part and key points of the second image space, the key points of the second image space being used to describe the interception constraint of the target part, the pixel proportion of the target part in the second transformed image being higher than the pixel proportion of the target part in the first transformed image; determining a visible region recognition result of the face and a visible region recognition result of the target part according to the first transformed image, the first transformation matrix, the second transformed image and the second transformation matrix.

2. The method of claim 1, wherein, The visible region recognition result of the face is obtained by transforming a region prediction result of the first transformed image back to the image space of the face image according to an inverse transformation matrix corresponding to the first transformation matrix, the region prediction result of the first transformed image being used to describe the position of the visible region of the face in the first transformed image; The visible region recognition result of the target part is obtained by transforming a region prediction result of the second transformed image back to the image space of the face image according to an inverse transformation matrix corresponding to the second transformation matrix, the region prediction result of the second transformed image being used to describe the position of the visible region of the target part in the second transformed image.

3. The method of claim 1, wherein, The method further comprises: splicing the first transformed image and the second transformed image to obtain a spliced image; determining the region prediction result of the first transformed image and the region prediction result of the second transformed image according to the spliced image, the region prediction result of the first transformed image being used to describe the position of the visible region of the face in the first transformed image, and the region prediction result of the second transformed image being used to describe the position of the visible region of the target part in the second transformed image; The determination of the visible region recognition result of the face and the visible region recognition result of the target part according to the first transformed image, the first transformation matrix, the second transformed image and the second transformation matrix comprises: determining the visible region recognition result of the face and the visible region recognition result of the target part according to the region prediction result of the first transformed image, the first transformation matrix, the region prediction result of the second transformed image and the second transformation matrix.

4. The method of claim 3, wherein, The width of the first transformed image is the same as the width of the second transformed image. The height of the spliced image is determined according to the sum of the height of the first transformed image and the height of the second transformed image.

5. The method according to any one of claims 1 to 4, characterized in that, The width of the first transformed image is less than the height of the first transformed image. And / or, The ratio between the height of the first transformed image and the height of the second transformed image is determined according to the ratio between the height of the face and the maximum height of the target part, the target part including multiple forms with different heights, and the height of each form being not greater than the maximum height of the target part.

6. The method according to any one of claims 1 to 4, characterized in that, The target part includes lips and an oral cavity, and the visible region identification result of the target part includes a visible region identification result of the lips and a visible region identification result of the oral cavity. And / or, The identification method is applied to a terminal device.

7. An identification device, characterized in that Comprise: An acquisition unit configured to acquire a face image, the face including a target part; A transformation unit configured to transform the face image to a first image space according to a first transformation matrix to obtain a first transformed image, and transform the face image to a second image space according to a second transformation matrix to obtain a second transformed image, the first transformation matrix being determined according to face key points of the face image and key points of the first image space, the key points of the first image space being used to describe the interception constraint of the face, the face key points including key points of the target part, the second transformation matrix being determined according to key points of the target part and key points of the second image space, the key points of the second image space being used to describe the interception constraint of the target part, and a pixel proportion of the target part in the second transformed image being higher than a pixel proportion of the target part in the first transformed image; A determination unit configured to determine a visible region identification result of the face and a visible region identification result of the target part according to the first transformed image, the first transformation matrix, the second transformed image, and the second transformation matrix.

8. An electronic device, comprising: The device comprises a processor and a memory; The memory is configured to store instructions or computer programs; The processor is configured to execute the instructions or computer programs in the memory to enable the electronic device to perform the method of any one of claims 1-6.

9. A computer readable medium characterized by The computer readable medium stores instructions or computer programs, which, when executed on a device, enable the device to perform the method of any one of claims 1-6.

10. A computer program product, characterised in that, It comprises a computer program carried on a non-transitory computer readable medium, the computer program containing program codes for executing the method of any one of claims 1-6.