A key point calibration method and device

By acquiring multiple target object images from different angles, using triangulation method and key point calibration model, combined with neural network training, the problems of slow key point calibration speed and low accuracy in the existing technology are solved, and accurate calibration at various angles and image quality is achieved, including calibration of two-dimensional and depth information.

CN113454684BActive Publication Date: 2025-08-01YINWANG INTELLIGENT TECHNOLOGIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180001870.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-24
Publication Date
2025-08-01
Estimated Expiration
2041-05-24

AI Technical Summary

Technical Problem

The existing key point calibration methods are slow and have low accuracy, and cannot effectively handle the situation of large face rotation angle or poor image quality, and cannot calibrate the depth information of key points.

Method used

By obtaining multiple target object images from different angles, using triangulation and key point calibration model, combined with neural network training, the position and depth of key points in the world coordinate system are automatically calibrated.

Benefits of technology

Improves the speed and accuracy of key point calibration, and can accurately calibrate key points, including two-dimensional and depth information, at various angles and image quality conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113454684B_ABST
    Figure CN113454684B_ABST
Patent Text Reader

Abstract

This application relates to the field of artificial intelligence, and specifically relates to a key point calibration method, including: obtaining multiple captured images and parameters of the capture device corresponding to the multiple captured images, where the poses of the target objects in the multiple captured images are the same and the capture angles are different, the multiple captured images include a first image and other images, the capture angle of the target object in the first image is less than a preset threshold, and the first image includes at least two images; determining the position of the key points in the world coordinate system according to the positions of the key points of the target object in the first image and the parameters of the capture device corresponding to the first image; determining the positions of the key points in the other images according to the parameters of the capture devices corresponding to the other images and the positions of the key points in the world coordinate system. It realizes the automatic calibration of key points, reduces the consumption of human resources; ensures the accuracy of key point calibration, and enables the calibration results to be put into practical use.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of autonomous driving, and particularly to a key point calibration method and device. Background Art

[0002] Identifying key points in an image is the basis for a computing device to perform vision tasks. For example, during face recognition or gesture recognition, it is necessary to first determine the positions of the key points of the face or fingers, and then, based on this, identify the current face or gesture through a series of algorithms. The model used to identify the current face or gesture needs to be trained using key point data. The larger the amount of key point data, the stronger the recognition ability of the trained model.

[0003] Existing key point data is obtained by manually calibrating images. Manual calibration has the following disadvantages: slow calibration speed, with each person being able to calibrate only about 100 - 200 images per day; different calibration personnel have inconsistent understandings of the calibration rules, and the calibration of the key points of the same image by two different calibration personnel may be different. Sometimes, even the positions of the key points of the same image calibrated by the same calibration personnel twice will be different; when the rotation angle of a face image relative to the camera is too large, resulting in part of the face being blocked, the calibration personnel can only guess approximately where the key points of the blocked part are, and the accuracy of calibration cannot be guaranteed; manual calibration can only calibrate the two-dimensional coordinates of the key points in the image and cannot calibrate the depth of the key points.

[0004] Therefore, how to obtain more key point data, ensure the accuracy of key point calibration, enable the key point calibration results to reach the level of commercial implementation, and reduce the consumption of human resources has become an urgent problem in the industry. Summary of the Invention

[0005] In view of this, this application provides a key point calibration method and device, which realizes automatic calibration of key points, reduces the consumption of human resources, and ensures the accuracy of key point calibration, enabling the calibration results to reach the level of commercial implementation.

[0006] The calibration method provided by this application can be executed by a local terminal, such as a terminal like a computer, or can be executed by a processor; it can also be executed by a server. Among them, the processor can be a central processing unit (CPU), a graphic processing unit (GPU), or a general-purpose processor, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The server can be a cloud server or a local server, and can be a physical server or a virtual server. This application does not make any limitations in this regard. The image acquisition device (such as a mobile phone, a terminal with a camera) sends the image to the local terminal. After receiving the image, the local terminal calibrates the key points of the image and stores the information of the key points obtained after calibration on the local memory; or, the image acquisition device (such as a mobile phone, a terminal with a camera) sends the image to the cloud server. After receiving the image, the server calibrates the key points of the image and stores the information of the key points obtained after calibration on the cloud memory, or transmits the information of the key points obtained after calibration (the coordinates of the key points in the image, the depth of the key points, etc.) back to the local terminal (such as a computer, a mobile phone, a camera) or transmits it back to the local memory.

[0007] In the first aspect of this application, a key point calibration method is provided, including: obtaining multiple captured images and the parameters of the capture device corresponding to the multiple captured images. The poses of the target objects in the multiple captured images are the same and the captured angles are different. The multiple captured images include a first image and other images. Among them, the captured angle of the target object in the first image is less than a preset threshold, and the first image includes at least two images; determining the position of the key points in the world coordinate system according to the position of the key points of the target object in the first image and the parameters of the capture device corresponding to the first image; and determining the position of the key points in the other images according to the parameters of the capture device corresponding to the other images and the position of the key points in the world coordinate system.

[0008] Through the above settings, multiple captured images of the target object at different angles in the same pose are obtained, increasing the quantity and variety of the captured images, thereby increasing the probability of obtaining a captured image with a captured angle less than the preset threshold. Furthermore, at least two captured images with a captured angle less than the preset threshold can be selected to determine the position of the key points in the world coordinate system, improving the accuracy of determining the position of the key points in the world coordinate system;

[0009] When the position of the key point in the world coordinate system is accurate, the position of the key point in other images can be accurately located, solving the problem that the key point cannot be accurately calibrated due to the too large acquisition angle of the target object, and improving the accuracy of calibrating the key point in the images under various acquisition angles;

[0010] The automatic calibration of the key point is realized, without manual calibration of the key point, improving the calibration efficiency of the key point and reducing human resources.

[0011] In a possible implementation, the multiple acquired images are images with standardized sizes.

[0012] Through the above settings, the size of the target object in the acquired images is unified, and further the accuracy of calibrating the position of the key point in the first image is improved.

[0013] In a possible implementation, according to the position of the key point of the target object in the first image and the parameters of the acquisition device corresponding to the first image, determining the position of the key point in the world coordinate system includes: according to the position of the key point in at least two first images and the parameters of the image acquisition device corresponding to the first image, solving the position of the key point in the world coordinate system by triangulation.

[0014] In a possible implementation, it further includes: calibrating the position of the key point in the world coordinate system to make the position of the key point in the world coordinate system located in the key area of the target object; according to the calibrated position of the key point in the world coordinate system and the parameters of the acquisition devices corresponding to the multiple acquired images, updating the position of the key point in the multiple acquired images.

[0015] Through the above settings, it is possible to obtain the position of the key point in the world coordinate system even when the position of the key point determined by the key point calibration model is inaccurate, and further update the determined position of the key point in the acquired image, ensuring the accuracy of key point recognition.

[0016] In a possible implementation, the parameters of the image acquisition device include the internal parameters of the cameras in the camera array. According to the position of the key point of the target object in the first image and the parameters of the acquisition device corresponding to the first image, determining the position of the key point in the world coordinate system includes: according to the position of the key point of the target object in the first image and the internal parameters of the cameras in the camera array, determining the position of the key point in the world coordinate system.

[0017] In a possible implementation, the target object includes a human face.

[0018] The target object of the present application is not limited to the human face, and can also be a human hand, a human body, etc.

[0019] In a possible implementation, the position of the key points in the first image is obtained through a key point calibration model, and the key point calibration model is trained as follows: obtaining other images and the positions of the determined key points in the other images; using the positions of the determined key points in the other images as the first training target, and training the key point calibration model according to the other images until the difference value between the positions of the key points obtained by the key point calibration model in the other images and the first training target converges.

[0020] Through the above settings, the accuracy of predicting the positions of the key points in the captured images by the key point calibration model can be improved, so that the prediction ability of the model can be enhanced as the input sample data increases.

[0021] In a possible implementation, the training method further includes: using the depths of the key points in multiple captured images as the second training target, and training the key point calibration model according to the multiple captured images until the difference value between the depths obtained by the key point calibration model and the second training target converges, where the depth of the second training target is obtained according to the positions of the key points in the world coordinate system and the captured angles of the target objects in the multiple captured images.

[0022] Through the above settings, the depths of the key points can be obtained, and the key point calibration model is enabled to have the function of predicting the depths of the key points.

[0023] In the second aspect of the present application, a key point calibration device is provided, including: a transceiver module and a processing module.

[0024] The transceiver module is used to obtain multiple captured images and the parameters of the capture devices corresponding to the multiple captured images. The poses of the target objects in the multiple captured images are the same and the captured angles are different. The multiple captured images include a first image and other images, where the captured angle of the target object in the first image is less than a preset threshold, and the first image includes at least two images; the processing module is used to determine the positions of the key points in the world coordinate system according to the positions of the key points of the target object in the first image and the parameters of the capture device corresponding to the first image; the processing module is further used to determine the positions of the key points in the other images according to the parameters of the capture devices corresponding to the other images and the positions of the key points in the world coordinate system.

[0025] In a possible implementation, the multiple captured images are images with standardized sizes.

[0026] In a possible implementation, the processing module is specifically used to solve the positions of the key points in the world coordinate system by using the triangulation method according to the positions of the key points in at least two first images and the parameters of the image capture devices corresponding to the first images.

[0027] In a possible implementation, the processing module is further configured to: calibrate the positions of the key points in the world coordinate system so that the positions of the key points in the world coordinate system are located in the key areas of the target object; the processing module is further configured to: update the positions of the key points in the multiple acquired images according to the positions of the calibrated key points in the world coordinate system and the parameters of the acquisition devices corresponding to the multiple acquired images.

[0028] In a possible implementation, the parameters of the image acquisition device include the internal parameters of the cameras in the camera array, and the processing module is specifically configured to determine the positions of the key points in the world coordinate system according to the positions of the key points of the target object in the first image and the internal parameters of the cameras in the camera array.

[0029] In a possible implementation, the target object includes a human face.

[0030] In a possible implementation, the positions of the key points in the first image are obtained through a key point calibration model, and the transceiver module is further configured to acquire other images and the positions of the determined key points in the other images; the processing module is further configured to use the positions of the determined key points in the other images as the first training target, and train the key point calibration model according to the other images until the difference value between the positions of the key points obtained by the key point calibration model in the other images and the first training target converges.

[0031] In a possible implementation, the processing module is further configured to: use the depths of the key points in the multiple acquired images as the second training target, and train the key point calibration model according to the multiple acquired images until the difference value between the depths obtained by the key point calibration model and the second training target converges, where the depths of the second training target are obtained according to the positions of the key points in the world coordinate system and the captured angles of the target object in the multiple acquired images.

[0032] The technical effects brought by the key point calibration device provided in the second aspect of the present application and any of its possible implementations are the same as those brought by the key point calibration method provided in the first aspect of the present application and any of its possible implementations. For the sake of brevity, they will not be elaborated here.

[0033] In the third aspect of the present application, a computing device is provided, including: a processor, the processor is coupled with a memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the computing device is enabled to execute the methods provided in the first aspect of the present application and its possible implementations.

[0034] In the fourth aspect of the present application, a computer-readable storage medium is provided. Program codes are stored in the computer-readable storage medium. When the program codes are executed by a terminal or a processor in the terminal, the methods provided in the first aspect of the present application and its possible implementations are implemented.

[0035] In a fifth aspect of the present application, a computer program product is provided. When the program code contained in the computer program product is executed by a processor in a terminal, the method provided in the first aspect of the present application and its possible implementation methods are implemented.

[0036] In the sixth aspect of the present application, a vehicle is provided, comprising: the key point calibration device provided by the second aspect of the present application and any possible implementation thereof, the computing device provided by the third aspect of the present application, the computer-readable storage medium provided by the fourth aspect of the present application, or the computer program product provided by the fifth aspect of the present application.

[0037] In the seventh aspect of the present application, a key point calibration system is provided, including: an image acquisition device and a computing device, wherein the image acquisition device is used to acquire multiple acquired images and send the multiple acquired images to the computing device, and the computing device is used to execute the key point calibration method provided by the above-mentioned first aspect and any possible implementation method thereof.

[0038] As a possible implementation of the seventh aspect, the computing device is further configured to send information of calibrated key points to the image acquisition device. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The following further illustrates the various features of the present application and the relationships between the various features with reference to the accompanying drawings. The accompanying drawings are all exemplary, and some features are not shown in actual proportion. In addition, some drawings may omit features that are customary in the field to which the present application relates and are not necessary for the present application, or additional features that are not necessary for the present application may be shown. The combination of the various features shown in the accompanying drawings is not intended to limit the present application. In addition, throughout this specification, the same figure numbers refer to the same content. The specific description of the drawings is as follows:

[0040] Figure 1 Schematic diagram of an application scenario of the key point calibration method provided in an embodiment of the present application;

[0041] Figure 2 is a flow chart of the key point calibration method provided in an embodiment of the present application;

[0042] Figure 3 This is a schematic diagram of a module of a key point calibration device provided in an embodiment of the present application;

[0043] Figure 4a This is a schematic diagram of the principle of using triangulation to locate the coordinates of key points in the world coordinate system provided by an embodiment of the present application;

[0044] Figure 4b This is a schematic diagram of the principle of using triangulation to locate the coordinates of a key point in a world coordinate system, provided in an embodiment of the present application, wherein the spatial position of the key point is not at the intersection of the straight line O1p1 and the straight line O2p2;

[0045] Figure 5a is a flowchart of the face key point calibration method provided by an embodiment of the present application;

[0046] Figure 5b is a schematic diagram of the face key point calibration rule provided by an embodiment of the present application;

[0047] Figure 6 is a structural schematic diagram of a computing device provided by an embodiment of the present application. Detailed implementation manners

[0048] To improve the accuracy of face key point calibration, a possible implementation is as follows: Obtain an initial face image, and after preprocessing the initial face image, obtain a face image to be detected; then use a first-level convolutional neural network to perform key point prediction on the face image to be detected to obtain predicted face key points; then perform second-level convolutional neural network processing and regression processing on the predicted key points to obtain target object key points, thereby improving the accuracy of face key point calibration.

[0049] However, this face key point calibration method has the following defects: 1. This method cannot make good predictions for key points in face images at various angles. For face images in which part of the face is blocked due to excessive face rotation angle, this method cannot make accurate key point predictions; 2. The key points obtained by this method are the results inferred by the model, with low credibility and cannot be directly used in practice; 3. This method has relatively high requirements for the quality of the image. When the image is blurred, this method is no longer feasible; 4. This method can only calibrate the two-dimensional information of the key points and cannot calibrate the depth information of the key points at the same time.

[0050] Another face key point calibration method is as follows: Obtain at least one frame of initial face image, and after preprocessing the initial face image, obtain at least one frame of face image to be detected; then use a convolutional neural network to perform feature extraction on the face image to be detected, and then input the extracted features into a recurrent neural network; the recurrent neural network combines the features of at least one frame of face image and the output of the previous frame of image passing through the recurrent neural network to predict multiple face key points in the current at least one frame of face image.

[0051] This face key point calibration method utilizes the temporal information between images, which means that the input images need to be several consecutive frames so that there can be a gradual change trend between the images. If there is no time correlation between each image, the key points in the image cannot be accurately recognized.

[0052] In order to enable the key points in images with low clarity, large face rotation angles, and uncorrelated time sequences to be automatically and accurately calibrated, and at the same time obtain the depth information of the key points, the embodiments of the present application provide a key point calibration method and device.

[0053] Figure 1 An exemplary application scenario of the key point calibration method provided by the embodiments of the present application is shown.

[0054] As Figure 1 shown, after an image acquisition device, such as a camera array 30, acquires images of a person 40 at different angles in the current posture, the acquired images are transmitted to a server, such as a computer 20. After receiving the images, the computer 20 calibrates the key points of the images and stores the information of the calibrated key points in a memory.

[0055] As Figure 1 shown, after the camera array 30 finishes acquiring images, the images can also be uploaded to a server 10. After receiving the images, the server 10 calibrates the key points of the images and can store the information of the calibrated key points in a cloud memory, or can transmit the information of the calibrated key points, such as the coordinates of the key points in the image (sometimes also referred to as the image coordinates of the key points in the image), the depth of the key points, etc., back to a local terminal (such as a computer, a mobile phone, a camera) or back to a local memory. Among them, the server can be a cloud server or a local server, and can be a physical server or a virtual server. The present application does not make any limitations in this regard.

[0056] Figure 2 A flowchart of the key point calibration method provided by the embodiments of the present application is shown.

[0057] The key point calibration method provided by the embodiments of the present application can be executed by a terminal, such as a terminal like a computer, or can be executed by a processor; the key point calibration method provided by the embodiments of the present application can also be executed by a server, where the processor can be a CPU, an image processor, or a general-purpose processor, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0058] Figure 2 The software code of the key point calibration method in Figure 2 can be stored in a memory, and by running the software code through a terminal or a server, the calibration of the face key points is realized. As Figure 2 shown, the key point calibration method includes the following steps:

[0059] Step S1: Obtain multiple acquired images and the parameters of the acquisition device corresponding to the multiple acquired images.

[0060] Among them, the poses of the target objects in the multiple acquired images are the same and the acquisition angles are different. The multiple acquired images include a first image and other images. Among them, the acquisition angle of the target object in the first image is less than a preset threshold, and the first image includes at least two images. The target object may include: a human face, a human hand, a human body, etc.

[0061] In some embodiments, the image acquisition device may include: a camera, a camera array, a mobile phone with a camera, and a computer with a camera. The image acquisition device may be one or more. When the image acquisition device is one, keep the target object in a fixed pose, and let the image acquisition device acquire images of the target object at different angles respectively. For example, an orbit may be set around the target object, and the image acquisition device moves along the orbit while acquiring images of the target object, and record the acquisition angle at the time of acquisition; when the image acquisition device is multiple, for example, when the image acquisition device is a camera array, let the camera array acquire images of the target object simultaneously.

[0062] In some embodiments, the types of cameras in the camera array may be the same or different. For example, the camera array may all use infrared (IR) cameras, all use red green blue (RGB) cameras or other cameras, or may also mix IR cameras and RGB cameras, so as to achieve the diversity of image data, so that the key point calibration model can support diverse image data.

[0063] In some embodiments, the parameters of the image acquisition device include the internal parameters of the cameras in the camera array. The internal parameters of the camera, also known as the camera projection matrix, are the parameters equipped with each calibrated camera. Using the camera projection matrix, the three-dimensional coordinates of the acquired target object in the world coordinate system can be converted into the two-dimensional coordinates on the acquired image.

[0064] When the three-dimensional coordinates of a key point are (X, Y, Z) and its corresponding two-dimensional coordinates (also referred to as image coordinates or coordinates in the image hereinafter) are (u, v), without considering the scaling factor, through this projection matrix The mapping performed can be described as follows:

[0065] Among them, takes a value of 0; takes a value of 1; is the ratio of the focal length of the camera to the width of the image pixels in the x-axis direction, and the x-axis is parallel to the u-axis; is the ratio of the focal length of the camera to the width of the image pixels in the y-axis direction, and the y-axis is parallel to the v-axis; is the coordinate of the intersection point of the optical axis of the camera and the image in the image.

[0066] When the server is local, it is only necessary to transmit the images collected by the camera to the local server by means of a data cable or signal transmission. The local server can perform steps S1 - S3 based on the images and the parameters of the camera. When the server is a remote server, it is also necessary to transmit the projection matrix of the camera corresponding to each image to the server, and the server executes steps S1 - S3 according to the images and the parameters of the camera.

[0067] In some embodiments, the multiple collected images are images with standardized sizes. For example, the region where the target object is located in the original image collected by the image acquisition device can be intercepted, and then the region where the target object is located is unified into images of the same size according to a preset size, which is convenient for subsequent identification of the positions of the key points of the target object in the first image. Size standardization can be implemented by a neural network. For example, it can be implemented by image segmentation models such as Regions with CNN features (RCNN) and Region Proposal Network (RPN).

[0068] In some embodiments, the captured angle of the target object can be obtained through a deflection angle recognition model. Since the multiple captured images respectively show different angles of the target object in the current posture, there must be some images that can completely display all the features of the target object. For the images that can completely and accurately display the features of the target object, the positioning of its key points in the first image must be more accurate than that of other images. It is more accurate to calculate the position of the key points in the world coordinate system using the image positions of the key points of at least two first images with a captured angle less than a preset value.

[0069] Step S2: Determine the position of the key points of the target object in the world coordinate system according to the positions of the key points of the target object in the first image and the parameters of the image acquisition device corresponding to the first image.

[0070] In some embodiments, the positions of the key points of the target object in the first image are obtained through a key point calibration model. The key point calibration model can be a neural network. In some embodiments, the neural network can be a convolutional neural network, a residual network, etc. The present application does not make any limitations in this regard.

[0071] In some embodiments, according to the positions of the key points in at least two of the first images and the parameters of the image acquisition devices corresponding to the first images, the positions of the key points in the world coordinate system are solved by the triangulation method.

[0072] In some embodiments, the positions of the key points in the image and the positions of the key points in the world coordinate system can be represented in the form of coordinates. Among them, the coordinates of any key point in at least two first images are respectively represented as: (u1, v1, 1) and (u2, v2, 1), and the camera projection matrices corresponding to at least two first images are respectively represented as:

[0073] M1: and M2:

[0074] Thus, the coordinates (X, Y, Z, 1) of the key point in the world coordinate system can be obtained according to the following formula:

[0075]

[0076]

[0077] Among them, in M1, takes the value of 0; takes the value of 1; is the ratio of the focal length of the camera of one of the first images to the width of the image pixels in the x-axis direction, and the x-axis is parallel to the u-axis; is the ratio of the focal length of the camera of one of the first images to the width of the image pixels in the y-axis direction, and the y-axis is parallel to the v-axis; is the coordinate of the intersection point of the optical axis of the camera of one of the first images and the image in the image, Z c1 represents the scaling factor of the camera corresponding to one of the first images. In M2, takes the value of 0; takes the value of 1; is the ratio of the focal length of the camera of the other first image to the width of the image pixels in the x-axis direction, and the x-axis is parallel to the u-axis; is the ratio of the focal length of the camera of the other first image to the width of the image pixels in the y-axis direction, and the y-axis is parallel to the v-axis; is the coordinate of the intersection point of the optical axis of the camera of the other first image and the image in the image, Z c2 represents the scaling factor of the camera corresponding to the other first image. After decomposing from Formula 1, it can be obtained:

[0078]

[0079] Eliminating Z c1 obtains:

[0080]

[0081] After decomposing Formula 2, it can be obtained:

[0082]

[0083] Eliminate Z c2 Obtain:

[0084]

[0085] From the above, Formula 4 and Formula 6 constitute four equations, and there are only three unknowns. Therefore, the coordinates (X, Y, Z, 1) of the key point P in the world coordinate system can be calculated.

[0086] Step S3: Determine the position of the key point in the other image according to the parameters of the acquisition device corresponding to the other image and the position of the key point in the world coordinate system.

[0087] Among them, the coordinates of any key point obtained in Step S2 in the other image are (u, v, 1), and the camera projection matrix corresponding to the other image is: Take the value of 0; Take the value of 1; is the ratio of the focal length of the camera of the other image to the width of the image pixels in the x-axis direction, and the x-axis is parallel to the u-axis; is the ratio of the focal length of the camera of the other image to the width of the image pixels in the y-axis direction, and the y-axis is parallel to the v-axis; is the coordinates of the intersection of the optical axis of the camera of the other image and the other image in the other image. Calculate the image coordinates of the key point in the other image according to the following Formula 7, where Z cb is the camera scaling factor corresponding to the other image:

[0088]

[0089] Decompose Formula 7 to obtain:

[0090]

[0091] Eliminate Z cb Obtain the image coordinates (u, v, 1) of the key point in the face image, where

[0092]

[0093] In some embodiments, the method further includes: calibrating the position of the key point in the world coordinate system so that the position of the key point in the world coordinate system is located in the key area of the target object; updating the position of the key point in the multiple acquisition images according to the calibrated position of the key point in the world coordinate system and the parameters of the acquisition devices corresponding to the multiple acquisition images.

[0094] Among them, the area within the first distance around the true position of the key point in the world coordinate system is the key area, and the key area is located on the target object; the position of the calibration key point in the world coordinate system can be calibrated by methods such as the least squares method, the gradient descent method, the Newton method, and the iterative non-linear least squares method.

[0095] When the key point calibration model cannot accurately identify the position of the key point in the first image in the early stage, the position of the obtained key point in the world coordinate system deviates from its true position. For example, as Figure 5b shown, the position of the key point 31 calculated in step S2 may not be located on the target object in the world coordinate system. For example, it is in front of the tip of the nose and deviates from its true position. Therefore, it is necessary to calibrate the position of the key point in the world coordinate system. Since the position of the calibrated key point in the world coordinate system has changed, it can be known from the calculation of the above formula that the positions of the key points in the first image and other images will also be updated.

[0096] When the key point calibration model can accurately identify the position of the key point in the first image, the calibration of its position in the world coordinate system can be omitted. When the key point recognition model can accurately identify the position of the key point in the first image, the position of the calibrated key point in the world coordinate system remains unchanged, so that the position of the updated key point in the multiple captured images is the same as the position before the update.

[0097] In some embodiments, the position of the key point in the first image can be obtained through a key point calibration model, and the key point calibration model is trained in the following manner: obtaining the other images and the determined positions of the key points in the other images; using the determined positions of the key points in the other images as the first training target, and training the key point calibration model according to the other images until the difference value between the positions of the key points obtained by the key point calibration model in the other images and the first training target converges.

[0098] Among them, the key point calibration model can be a key point convolutional neural network model, and the training samples can be multiple other images with key point positions. Among them, one form of expression of the key point position can be the coordinates of the key point in the other image, that is, the coordinates of the key point in the other image correspond to the other image, and are used to identify the key point in the other image. During training, the parameters of the key point convolutional neural network model can be initialized, and then the multiple other images are input into the key point convolutional neural network model. After being processed by the key point convolutional neural network model, the coordinates of the key points in the multiple other images are output; the coordinates of the key points in the multiple other images output are compared with the coordinates of the key points in the training samples in the multiple other images. For example, corresponding operations are performed to obtain a difference value, and the initialized key point convolutional neural network model is adjusted according to the difference value. The adjusted key point convolutional neural network model is used to process the other images of the training samples, and then a new difference value is obtained. This is repeated iteratively until the difference value converges; a preset condition that the difference value should meet can also be set. If the difference value does not meet this preset condition, the parameters of the key point convolutional neural network model are adjusted, and the adjusted key point convolutional neural network model is used to process the other images of the training samples. This is repeated iteratively until the difference value meets this preset condition.

[0099] In some embodiments, the training samples can also be multiple captured images with the positions of the updated key points in the captured images. The coordinates of the updated key points in the captured images correspond to the captured images and are used to identify the key points in the captured images. Taking the positions of the updated key points in the captured images as the third training objective, the key point calibration model is trained, and its training method is the same as the above training method and will not be elaborated here.

[0100] By initializing the key point convolutional neural network model, inputting the training samples into the initialized key point convolutional neural network model, and through cyclic iteration, the target key point convolutional neural network model is obtained, so that the positioning accuracy of the key points of the target convolutional neural network model obtained by training can be improved.

[0101] In some embodiments, the key point calibration model can also be obtained by training in the following manner: taking the depths of the key points in the multiple captured images as the second training objective, training the key point calibration model according to the multiple captured images until the difference value between the depth obtained by the key point calibration model and the second training objective converges, where the depth is obtained according to the position of the key point in the world coordinate system and the captured angles of the target objects in the multiple captured images.

[0102] In some embodiments, one form of the position of a key point in the world coordinate system is the coordinate of the key point in the world coordinate system. After obtaining the coordinate of the key point in the world coordinate system, the depth of the key point relative to the acquisition device can be obtained according to the coordinate of the acquisition device in the world coordinate system. In some embodiments, when multiple acquired images are face images, one of the key points is selected from multiple key points as a reference key point, and the depth of the reference key point relative to the acquisition device is subtracted from the depths of all key points relative to the acquisition device respectively, and then the depth of the key points of each face image is obtained according to the acquired angle. For example, as Figure 5b shown in the frontal face image (i.e., the first image), the nose tip key point 31 is selected as the reference key point, and the Z values of key points 1 - key point 68 are subtracted from the Z value of the nose tip key point 31 respectively to obtain the Z value difference of all key points relative to the nose tip key point 31. This Z value difference is used as the depth of the key point, so as to determine the depth information of the key points in the figure. For other images, the depth of the key points can be obtained according to the acquired angle of the other images and the position of the key points relative to the nose tip key point 31.

[0103] Multiple acquired images with depth can also be used as training samples and input into the key point calibration model for training in the above manner, so that the key point calibration model can have the ability to recognize the depth of key points.

[0104] Next, taking face images as an example, the key point calibration method provided by the embodiments of the present application will be described.

[0105] Specific Embodiment 1: Calibration Method for Facial Key Points

[0106] Refer to Figure 5a A specific embodiment of the key point calibration method provided by the embodiments of the present application will be described. Figure 5a The software code of the calibration method for facial key points in steps S100 - step S180 can be stored in the memory, and the processor or server of the electronic device runs the software code to implement the calibration of facial key points.

[0107] In this specific embodiment, the calibration method for facial key points includes the following steps:

[0108] Step S100: Obtain the original images of a person at different angles in the same pose and the internal parameters of the camera corresponding to the original images.

[0109] The original image can be captured by a camera array, which is composed of multiple cameras and arranged around the person being captured. Each camera in the camera array has a different angle relative to the person being captured, thereby ensuring that at least two cameras in the camera array can face the face of the person being captured, so that all facial features of the person being captured can be captured by the camera.

[0110] The cameras in a camera array can be of the same or different types. For example, an array can use all IR cameras or all RGB cameras. A camera array can also use a mixture of IR and RGB cameras. Because the intrinsic parameters of each camera are determined at the factory, the camera projection matrix can be calculated based on these camera parameters.

[0111] When the server used to calibrate key points is a local server, you only need to transfer the original image to the local server via a data cable. The local server can directly obtain the camera parameters corresponding to the original image. When the server used to calibrate key points is a server, you also need to send the camera parameters corresponding to each image to the server for use in subsequent steps.

[0112] Camera parameters can include the intrinsic parameters of each camera in the camera array. The intrinsic parameters of the camera, also known as the camera projection matrix, are the parameters of each calibrated camera. The camera projection matrix is used to convert the coordinates of the target object in the world coordinate system to the coordinates on the image.

[0113] Step S110: capturing a face region in the original image, and adjusting the captured face region to a preset size to obtain a face image to be recognized.

[0114] The facial image to be recognized obtained in step S110 is the size-standardized image captured in steps S1-S3. Because the original images of the current person captured by the camera array at different angles also contain parts of the person's body, it is necessary to locate the area containing the face in the original image, capture the facial area from the original image, and resize the captured facial area to a preset size to obtain the facial image to be recognized. In some embodiments, the original image captured by the camera array can be input into an image segmentation model, which captures the facial area in the original image. The image segmentation model then resizes the facial area to obtain the facial image to be recognized to a uniform size.

[0115] Step S120: using the captured angle recognition model to identify the posture of the face in the face image, obtain the captured angle of the face, and select at least two faces whose captured angles are smaller than a preset value as the first frontal face image and the second frontal face image.

[0116] The first frontal face image and the second frontal face image are at least two first images in steps S1 - S3 in the embodiment. In this embodiment, only two first frontal face images and the second frontal face image are selected as the first images, but the present application is not limited thereto, and three or more images can also be selected.

[0117] Step S130: Use the key point calibration model to identify the key points in the first frontal face image and the second frontal face image, and obtain the image coordinates of the key points in the first frontal face image and the second frontal face image.

[0118] In some embodiments, all the face images obtained in step S120 can also be input into the key point calibration model to obtain the image coordinates of the key points in the face images.

[0119] In some embodiments, the acquired angle recognition model and the key point calibration model can be the same neural network. For example, it can be a convolutional neural network, a residual network, etc., and the present application does not limit this.

[0120] When the acquired angle recognition model and the key point calibration model are the same neural network, input multiple face images into the same neural network, and output the coordinates of the key points in each face image and the acquired angle of the face in each face image; select at least two first frontal face images and the second frontal face image with the acquired angle of the face less than a preset value from the multiple face images.

[0121] Figure 5b Shows a face key point calibration rule provided by an embodiment of the present application. 68 key points need to be calibrated on the face image. After inputting the face image into the key point calibration model, the key point calibration model calibrates the key points of the face according to the Figure 5b shown rule, and respectively outputs the image coordinates (u i , v i , 1) of each key point, where i represents the i-th key point recognized in the face image.

[0122] Since the camera array can surround the person being captured, at least two face images with the acquired angle less than a preset value can be selected from multiple face images, which can completely display the facial features of the face without the facial features of the face being blocked or incompletely displayed due to the excessive rotation angle of the face.

[0123] Step S140: According to the parameters of the cameras corresponding to the first frontal face image and the second frontal face image and the image coordinates of the key points in the first frontal face image and the second frontal face image, use the triangulation method to determine the initial coordinates of the key points in the world coordinate system.

[0124] Figure 4aShows the schematic diagram of determining the initial coordinates of key points in the world coordinate system by triangulation. As Figure 4a shown, for any key point P, its first image point on the first camera C1 is p1, and its second image point on the second camera C2 is p2. The optical center of the first camera C1 is O1, and the optical center of the second camera C2 is O2. In an ideal situation, the position of the key point P in the world coordinate system is the intersection of the straight line O1p1 and the straight line O2p2.

[0125] In step S130, the first image coordinates (u1, v1, 1) of any key point in the first frontal face image and the second image coordinates (u2, v2, 1) in the second frontal face image have been obtained.

[0126] In step S100, the internal parameters of each camera in the camera array are known. The internal parameters of the first camera corresponding to the first frontal face image are the first camera projection matrix M1: The internal parameters of the second camera corresponding to the second frontal face image are the second camera projection matrix M2:

[0127] Thus, the coordinates (X, Y, Z, 1) of the key point P in the world coordinate system can be obtained according to the following formula:

[0128]

[0129]

[0130] Among them, in M1, takes the value of 0; takes the value of 1; is the ratio of the focal length of the camera of the first frontal face image to the width of the image pixels in the x-axis direction, and the x-axis is parallel to the u-axis; is the ratio of the focal length of the camera of the first frontal face image to the width of the image pixels in the y-axis direction, and the y-axis is parallel to the v-axis; is the coordinate of the intersection of the optical axis of the camera of the first frontal face image and the image in the image, Z c1 represents the scaling factor of the camera corresponding to the first frontal face image. In M2, takes the value of 0; takes the value of 1; is the ratio of the focal length of the camera of the second frontal face image to the width of the image pixels in the x-axis direction, and the x-axis is parallel to the u-axis; is the ratio of the focal length of the camera of the second frontal face image to the width of the image pixels in the y-axis direction, and the y-axis is parallel to the v-axis; is the coordinate of the intersection of the optical axis of the camera of the second frontal face image and the image in the image, Z c2Represents the zoom factor of the camera corresponding to the second frontal face image. After decomposing from Equation 1, it can be obtained that:

[0131]

[0132] Eliminate Z c1 Obtain:

[0133]

[0134] After decomposing Equation 2, it can be obtained that:

[0135]

[0136] Eliminate Z c2 Obtain:

[0137]

[0138] From the above, Equation 4 and Equation 6 form four equations and only have three unknowns. Therefore, the coordinates (X, Y, Z, 1) of the key point P in the world coordinate system can be calculated.

[0139] Since the first frontal face image and the second frontal face image can completely display a person's facial features, the key point calibration model can relatively accurately identify the positions of the key points in the first frontal face image and the second frontal face image. Therefore, it is relatively reliable to calculate the coordinates of the key points in the world coordinate system through the first image coordinates and the second image coordinates of the key points in the first frontal face image and the second frontal face image.

[0140] Step S150: Calibrate the initial coordinates of the key points in the world coordinate system to obtain the final coordinates of the key points in the world coordinate system.

[0141] Among them, the least squares method can be used to calibrate the coordinates of the key points in the world coordinate system, but the calibration method is not limited to the least squares method. It can also be the gradient descent method, the Newton method, and the iterative non-linear least squares method.

[0142] If the key point calibration model was not well-trained in the early stage, that is, the key point calibration model cannot very accurately identify the positions of the key points in the first frontal face image and the second frontal face image, resulting in the key points deviating from the key regions of their corresponding faces, and the intersection point of O1p1 and the straight line O2p2 is not the position of the key point P in the world coordinate system (as Figure 4b shown), therefore, step S150 needs to be executed. If the model can very accurately identify the positions of the key points in the first frontal face image and the second frontal face image, the intersection point of the straight line O1p1 and the straight line O2p2 is the position of the key point P in the world coordinate system, and step S150 can be omitted.

[0143] After step S150, the final coordinates (X, Y, Z, 1) of the 68 key points of the frontal face image shown in Figure 5b the world coordinate system can be obtained respectively.

[0144] Step S160: Determine the depth of the key points in each face image according to the final coordinates of the key points in the world coordinate system and the captured angle of the face.

[0145] Through step S150, the coordinates (X, Y, Z, 1) of any key point in the face image in the world coordinate system can be obtained. Among them, the Z value of the coordinates (X, Y, Z, 1) in the world coordinate system is the depth of the key point relative to the camera corresponding to the face image. Since the face image is obtained by size normalization on the basis of the original image of the person captured by the camera, the size of each face image is the same. In this case, when the distances between the two cameras and the face are different, from the perspective of the face image, the depth of the face relative to the camera is also the same. If the Z value of the coordinates of the key points in the world coordinate system is directly used as the depth to train the model, the model cannot accurately identify the depth of the key points.

[0146] Therefore, after calculating the coordinates (X, Y, Z, 1) of each key point in the world coordinate system, one of the key points is selected from multiple key points as a reference key point. The Z value of the reference key point is subtracted from the Z values of all key points respectively, and then the depth of the key points of each face image is obtained according to the captured angle of the target object. For example, as Figure 5b shown in the frontal face diagram, the tip-of-nose key point 31 is selected as the reference key point, and the Z values of key points 1 to key point 68 are subtracted from the Z value of the tip-of-nose key point 31 respectively, obtaining the Z value difference of all key points relative to the tip-of-nose key point 31. This Z value difference is used as the depth of the key point, so as to determine the depth information of the key points on the frontal face diagram. For other face images, the depth of the key points can be obtained according to the captured angle of the face and the position of the key points relative to the tip-of-nose key point 31.

[0147] In some embodiments, the determined depth of the key points can also be used as the training target, and the above-mentioned multiple face images are used to train the key point calibration model, so that the key point calibration model in step S130 can further have the ability to predict the depth of the key points. Of course, the depth of the key points obtained in step S160 can also be stored in the memory for other recognition operations; the depth of the key points obtained in step S160 can also be input into other neural networks for training, and the present application does not limit this.

[0148] Step S170: Determine the image coordinates of the key points in the face image according to the final coordinates of the key points in the world coordinate system and the parameters of the camera corresponding to the face image.

[0149] In some embodiments, when the final coordinates of the key points in the world coordinate system are different from the initial coordinates, step S170 updates the image coordinates of the key points obtained in step S130 in the first frontal face image and the second frontal face image.

[0150] In some embodiments, when the final coordinates of the key points in the world coordinate system are different from the initial coordinates and in step S130, all the face images obtained in step S120 are input into the key point calibration model, step S170 updates the image coordinates of the key points in all the face images.

[0151] For a face image with a large face angle, since some facial features are not shown, the key point calibration model may not accurately identify the positions of the key points corresponding to the corners of the eyes, resulting in the deviation of the key points corresponding to the corners of the eyes. By using the coordinates of the key points in the world coordinate system and the parameters of the camera corresponding to the face image, the coordinates of the key points in the face image can be determined.

[0152] In step S150 or step S140, the coordinates (X, Y, Z, 1) of the key points in the world coordinate system have been obtained. Since the coordinates of the key points in the world coordinate system are invariant, the image coordinates of the key points in other face images can be determined by using the coordinates of the key points in the world coordinate system.

[0153] According to the formula: Calculate the image coordinates of the key points in other face images.

[0154] Where, (u, v, 1) are the image coordinates of the key points in other face images, (X, Y, Z, 1) are the coordinates of the key points in the world coordinate system calculated in step S150 or step S140, takes the value of 0; takes the value of 1; is the ratio of the focal length of the camera of other face images to the width of the image pixels in the x-axis direction, and the x-axis is parallel to the u-axis; is the ratio of the focal length of the camera of other face images to the width of the image pixels in the y-axis direction, and the y-axis is parallel to the v-axis; are the coordinates of the intersection point of the optical axis of the camera of other face images and the image in the image. Calculate the coordinates of the key points in other face images according to the following formula 7, Z cb is the camera scaling factor corresponding to other face images.

[0155] Formula 7 is decomposed to obtain:

[0156]

[0157] Eliminate Z cb Obtain the image coordinates (u, v, 1) of the key points in the face image, where

[0158]

[0159] After step S170 and obtaining the image coordinates of the key points in all face images, the face images calibrated with the image coordinates can be used to train the key point calibration model, update the parameters of the key point calibration model, realize the iterative optimization of the key point calibration model, and then improve the recognition ability of the key point calibration model for the key points in the face image.

[0160] It should be noted that in step S130, when the key points in the first frontal face image and the second frontal face image recognized by the key point calibration model are relatively accurate, the initial coordinates of the key points calculated in step S140 in the world coordinate system are located in the key area, and step S150 can be omitted. Or, the final coordinates of the key points obtained after step S150 in the world coordinate system may be the same as the initial coordinates obtained before calibration. Therefore, in step S170, the image coordinates of the key points determined in the first frontal face image and the second frontal face image may be the same as the image coordinates obtained in step S130.

[0161] Step S180: Map the image coordinates of the key points in the face image to the original image to obtain the image coordinates of the key points in the original image.

[0162] Since the face image to be recognized is obtained by intercepting the face area from the image according to a preset size, the image coordinates of the key points in the face image are not the image coordinates of the key points in the original image. It is necessary to map the calibrated image coordinates of the key points in the face image to the original image, and through coordinate transformation, obtain the image coordinates of the key points in the original image.

[0163] In steps S170 and S160, the image coordinates of the key points in the face image and the depths of the key points have been obtained. To improve the recognition ability of the key point calibration model for key points, the face image calibrated with the image coordinates of the key points and / or the depths of the key points can be used as a training sample, and the image coordinates of the key points and / or the depths of the key points obtained in steps S170 and S160 can be used as training targets to train the key point calibration model. The training method can be as follows: Initialize the parameters of the key point calibration model, and then input the face image into the key point calibration model. After being processed by the key point calibration model, the face image outputs the image coordinates of the key points in the face image and / or the depths of the key points; Compare the output image coordinates of the key points in the face image and / or the depths of the key points with the image coordinates of the key points in the face image and / or the depths of the key points calibrated by the training sample. For example, perform corresponding operations to obtain a difference value, and adjust the initialized key point calibration model according to the difference value; Process the face image of the training sample with the adjusted key point calibration model, and then calculate a new difference value. Repeat this iteration until the difference value converges; A preset condition that the difference value should meet can also be set. If the difference value does not meet this preset condition, the parameters of the key point convolutional neural network model can be adjusted. Process other images of the training sample with the adjusted key point convolutional neural network model, and then calculate a new difference value. Determine whether the new difference value meets the preset condition. If it meets the preset condition, the target key point convolutional neural network model is obtained. If it does not meet the condition, continue to iterate until this preset condition is met.

[0164] Since it can be ensured that the image coordinates and depths of the key points in the face image are accurate in steps S170 and S150, using the face image, the depth, and the image coordinates of the key points in the face image to train the key point calibration model can improve the accuracy of the key point calibration model in recognizing key points, and at the same time enable the key point calibration model to have the ability to predict the depths of key points.

[0165] When the server is local, after step S180, the coordinates of the key points in the original image and the depths of the key points can be stored in a memory or folder for subsequent use; When the server is a server, the coordinates of the key points in the image and the depths of the key points can be stored in a cloud memory for subsequent use, and the coordinates of the key points in the image and the depths of the key points can also be sent back to a local terminal (such as a camera, mobile phone, computer, etc.) for subsequent use.

[0166] In the above embodiments of the present application, the expression of the coordinates of the key points in the image or the image coordinates of the key points in the image refers to the rows and columns of the image pixels corresponding to the key points, which are represented by (u, v) in the embodiments of the present application.

[0167] Figure 3 The module schematic diagram of the key point calibration device provided by the embodiments of the present application is shown. As Figure 3 shown, the key point calibration device provided by the embodiments of the present application includes: a transceiver module 1000 and a processing module 2000.

[0168] The transceiver module 1000 is used to obtain multiple captured images and the parameters of the capturing device corresponding to the multiple captured images. The poses of the target objects in the multiple captured images are the same and the capturing angles are different. The multiple captured images include a first image and other images. Among them, the capturing angle of the target object in the first image is less than a preset threshold, and the first image includes at least two images;

[0169] The processing module 2000 is used to determine the position of the key point in the world coordinate system according to the position of the key point of the target object in the first image and the parameters of the capturing device corresponding to the first image;

[0170] The processing module 2000 is further used to determine the position of the key point in the other images according to the parameters of the capturing device corresponding to the other images and the position of the key point in the world coordinate system.

[0171] In some embodiments, the multiple captured images are images with standardized sizes.

[0172] In some embodiments, the processing module 2000 is specifically used to solve the position of the key point in the world coordinate system by triangulation according to the positions of the key point in at least two first images and the parameters of the image capturing device corresponding to the first image.

[0173] In some embodiments, the processing module 2000 is further used to: calibrate the position of the key point in the world coordinate system so that the position of the key point in the world coordinate system is located in the key area of the target object;

[0174] The processing module 2000 is further used to: update the position of the key point in the multiple captured images according to the calibrated position of the key point in the world coordinate system and the parameters of the capturing device corresponding to the multiple captured images.

[0175] In some embodiments, the parameters of the image capturing device include the internal parameters of the cameras in the camera array. The processing module is specifically used to determine the position of the key point in the world coordinate system according to the position of the key point of the target object in the first image and the internal parameters of the cameras in the camera array.

[0176] In some embodiments, the target object includes a human face.

[0177] In some embodiments, the position of the key point in the first image is obtained by a key point calibration model, and the transceiver module 1000 is further configured to acquire the other image and the position of the determined key point in the other image; the processing module 2000 is further configured to use the position of the determined key point in the other image as a first training target, and train the key point calibration model according to the other image until the difference value between the position of the key point obtained by the key point calibration model in the other image and the first training target converges.

[0178] In some embodiments, the processing module 2000 is further configured to: use the depth of the key point in the multiple images as a second training target, and train the key point calibration model according to the multiple acquired images until the difference value between the depth obtained by the key point calibration model and the second training target converges, where the depth is obtained according to the position of the key point in the world coordinate system and the acquisition angles of the target object in the multiple acquired images.

[0179] It should be noted that the above-mentioned modules, that is, the transceiver module 1000 and the processing module 2000 are used to execute the relevant steps of the above method. For example, the transceiver module 1000 is used to execute the relevant content of steps S1 and S100, etc., and the processing module 2000 is used to execute the relevant content of steps S2, S3, S110 to S180, etc.

[0180] In this embodiment, the key point calibration device is presented in the form of a module. Here, the "module" may refer to an application-specific integrated circuit (ASIC), a processor and a memory that execute one or more software or firmware programs, an integrated logic circuit, and / or other devices that can provide the above functions. In addition, the above transceiver module 1000 and processing module 2000 can be implemented by Figure 6 the computing device shown.

[0181] Figure 6 FIG. 16 is a structural schematic diagram of a computing device 1500 provided in an embodiment of the present application. The computing device 1500 includes: a processor 1510 and a memory 1520 coupled to the processor 1510. The memory 1520 is used to store programs or instructions. When the programs or instructions are executed by the processor, the computing device is caused to execute the key point calibration method provided in the embodiment of the present application. Among them, the memory 1520 may be an internal storage unit of the processor 1510, an external storage unit independent of the processor 1510, or a component including an internal storage unit of the processor 1510 and an external storage unit independent of the processor 1510.

[0182] Optionally, the computing device 1500 may further include a bus and a communication interface (not shown in the figure). Among them, the memory 1520 and the communication interface may be connected to the processor 1510 through the bus. The bus may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc.

[0183] It should be understood that in the embodiments of the present application, the processor 1510 may be implemented by means such as a CPU. The processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Or the processor 1510 employs one or more integrated circuits to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.

[0184] The memory 1520 may include a read-only memory and a random access memory, and provide instructions and data to the processor 1510. A part of the processor 1510 may also include a non-volatile random access memory. For example, the processor 1510 may also store information about the device type.

[0185] When the computing device 1500 is running, the processor 1510 executes the computer-executable instructions in the memory 1520 to perform the automatic calibration of the image key points of the present application.

[0186] It should be understood that the computing device 1500 according to the embodiments of the present application may correspond to the corresponding main body executing the methods according to the embodiments of the present application, and the above and other operations and / or functions of each module in the computing device 1500 respectively implement the corresponding processes of the methods in the present embodiments. For the sake of brevity, they will not be described in detail here.

[0187] Those of ordinary skill in the art will recognize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0188] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0189] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0190] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0191] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0192] When the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

[0193] Embodiments of this application also provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it is used to execute a key point calibration method, and this method includes at least one of the solutions described in the above various embodiments.

[0194] The computer storage medium of the embodiments of this application can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, device, or component.

[0195] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, device, or component.

[0196] The program code contained on a computer-readable medium can be transmitted by any suitable medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the above.

[0197] The computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0198] The embodiments of this application also provide a computer program product. When the program code contained in the computer program product is executed by a processor in a terminal, it is used to implement the key point calibration method provided in the above embodiments.

[0199] The terms "first, second, third, etc." or similar terms such as module A, module B, module C, etc. in the specification and claims are only used to distinguish similar objects and do not represent a specific order for the objects. Understandably, under permitted circumstances, the specific order or sequence can be interchanged so that the embodiments of this application described here can be implemented in an order other than that illustrated or described here.

[0200] In the above description, the reference numerals representing steps, such as S110, S120, etc., do not necessarily mean that the steps will be executed in this order. Under permitted circumstances, the order of the front and back steps can be interchanged, or they can be executed simultaneously.

[0201] The term "comprising" used in the specification and claims should not be construed as being limited to the content listed thereafter; it does not exclude other elements or steps. Therefore, it should be interpreted as specifying the existence of the mentioned features, wholes, steps, or components, but does not exclude the existence or addition of one or more other features, wholes, steps, or components and their groups. Therefore, the expression "a device comprising device A and B" should not be limited to a device consisting only of components A and B.

[0202] As used herein, the term "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present application. Thus, the phrases "in one embodiment" or "in an embodiment" appearing throughout this specification are not necessarily all referring to the same embodiment, but may refer to the same embodiment. In addition, in one or more embodiments, the various specific features, structures, or characteristics can be combined in any suitable manner, as will be apparent to those of ordinary skill in the art from the present disclosure.

[0203] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. In case of inconsistency, the meaning as set forth in this specification or the meaning derived from the content recorded in this specification shall prevail. Additionally, the terms used herein are for the purpose of describing embodiments of the present application only and are not intended to limit the present application.

[0204] Note that the above are only the preferred embodiments of the present application and the technical principles applied. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of the present application. Therefore, although the present application has been described in detail through the above embodiments, the present application is not limited to the above embodiments. Without departing from the concept of the present application, more other equivalent embodiments can be included, all of which fall within the scope of protection of the present application.

Claims

1. A key point calibration method, characterized in that, Comprising: Obtaining a plurality of captured images and parameters of a capture device corresponding to the plurality of captured images, wherein poses of a target object in the plurality of captured images are the same and captured angles are different, the plurality of captured images includes a first image and other images, wherein a captured angle of the target object in the first image is less than a preset threshold, the first image includes at least two images, and the parameters of the capture device corresponding to the plurality of captured images include internal parameters of cameras in a camera array, and the internal parameters of the cameras include camera projection matrices; Determining positions of key points of the target object in a world coordinate system according to positions of the key points of the target object in the first image and the internal parameters of the cameras in the camera array; and Determining positions of the key points in the other images according to the parameters of the capture device corresponding to the other images and the positions of the key points in the world coordinate system, wherein the plurality of captured images are images with standardized sizes.

2. The method according to claim 1, wherein The target object includes a human face.

3. The method according to claim 1 or 2, characterized in that, The positions of the key points in the first image are obtained through a key point calibration model, and the key point calibration model is obtained through the following training method: Obtaining the other images and the determined positions of the key points in the other images; Taking the determined positions of the key points in the other images as a first training target, and training the key point calibration model according to the other images until a difference value between the positions of the key points obtained by the key point calibration model in the other images and the first training target converges.

4. The method according to claim 3, wherein The training method further includes: Taking depths of the key points in the plurality of captured images as a second training target, and training the key point calibration model according to the plurality of captured images until a difference value between the depths obtained by the key point calibration model and the second training target converges, wherein the depths of the second training target are obtained according to the positions of the key points in the world coordinate system and the captured angles of the target object in the plurality of captured images.

5. A key point calibration device, characterized in that, Comprising: A transceiver module and a processing module, The transceiver module is configured to obtain a plurality of captured images and parameters of a capture device corresponding to the plurality of captured images, wherein poses of a target object in the plurality of captured images are the same and captured angles are different, the plurality of captured images includes a first image and other images, wherein a captured angle of the target object in the first image is less than a preset threshold, the first image includes at least two images, and the parameters of the capture device corresponding to the plurality of captured images include internal parameters of cameras in a camera array, and the internal parameters of the cameras include camera projection matrices; The processing module is configured to determine positions of key points of the target object in a world coordinate system according to positions of the key points of the target object in the first image and the internal parameters of the cameras in the camera array; The processing module is further configured to determine positions of the key points in the other images according to the parameters of the capture device corresponding to the other images and the positions of the key points in the world coordinate system, wherein the plurality of captured images are images with standardized sizes.

6. The device according to claim 5, wherein The target object includes a human face.

7. The device according to claim 5 or 6, characterized in that, The position of the key point in the first image is obtained through a key point calibration model. The transceiver module is further configured to obtain the other image and the position of the determined key point in the other image. The processing module is further configured to use the position of the determined key point in the other image as a first training target, and train the key point calibration model according to the other image until the difference value between the position of the key point obtained by the key point calibration model in the other image and the first training target converges.

8. The apparatus according to claim 7, wherein The processing module is further configured to use the depth of the key point in the multiple captured images as a second training target, and train the key point calibration model according to the multiple captured images until the difference value between the depth obtained by the key point calibration model and the second training target converges, wherein the depth of the second training target is obtained according to the position of the key point in the world coordinate system and the captured angle of the target object in the multiple captured images.

9. A computing device, characterized in that, Comprising: A processor, the processor is coupled to a memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the computing device executes the method according to any one of claims 1-4.

10. A computer-readable storage medium, characterized in that, Program code is stored in the computer-readable storage medium, and is characterized in that when the program code is executed by a terminal or a processor in the terminal, the method according to any one of claims 1-4 is implemented.

11. A computer program product, characterized in that: When the program code included in the computer program product is executed by a processor in a terminal, the method according to any one of claims 1-4 is implemented.

Citation Information

Patent Citations

  • Image processing method and device, processor, electronic equipment and storage medium

    CN111160178A

  • Three-dimensional pose determination method and device, electronic equipment and storage medium

    CN112767489A