Image key point detection method, device, computer equipment and storage medium
By using a pre-trained model in image key point detection, combining the output values of regression branches and inter-frame residual branches, the problem of key point position jitter in continuous frame-type images is solved, and a more stable key point detection effect is achieved.
Patent Information
- Application Number
- CN202210448849.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-26
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-04-26
AI Technical Summary
The existing face key point detection technology is prone to jitter in continuous frame types such as videos, which affects the display effect of the application.
An image key point detection method is provided, and a multi-channel image is generated by acquiring the image to be detected and preprocessing according to its type. The multi-channel image is then input into a pre-trained keypoint detection model, which includes regression branches and inter-frame residual branches set in parallel. For discrete frame type images, only the regression prediction value output from the regression branch is used to determine the key point coordinates; for continuous frame type images, the key point coordinates are determined in combination with the output values of the regression branch and the inter-frame residual branch to improve the stability of the detection results.
This method can take into account both discrete frame and continuous frame type image detection, significantly improve the stability of adjacent frame detection results of continuous frame type images, and avoid excessive jump in key points positions.
Smart Images

Figure CN114926876B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to an image key point detection method, device, computer equipment and storage medium. Background Art
[0002] In the field of image processing technology, there is a key point detection technology, which is used to detect and track key points of interest on a target object in an image. Depending on the application field, the target object may be, for example, a human body, a face, or other objects.
[0003] Taking the target object as a human face, the corresponding facial key point detection technology is widely used in face-related applications, such as face beautification, face makeup, face recognition, etc. Figure 1 As shown in FIG, face key points are multiple key points located at multiple predefined positions on the face. These predefined key points are composed of points with clear semantic definitions and other points. In face key point detection, the position of each predefined key point on the face needs to be detected.
[0004] Existing facial key point detection technologies generally perform key point detection on single-frame images, and their main focus is on improving the accuracy of key point detection. However, for key point detection of continuous frame images such as videos, the key point detection results of the previous and next frames measured by the existing key point detection technologies are prone to large differences, resulting in jitter and other instabilities in the measured key point positions between the previous and next frames, which in turn affects the display effect of face-related applications. Summary of the invention
[0005] Based on this, it is necessary to provide an image key point detection method, device, computer equipment and storage medium that can take into account the key point detection of both discrete frame and continuous frame types of images, while improving the stability of adjacent frame detection results of continuous frame type images, in order to address the above technical problems.
[0006] A method for detecting key points of an image, the method comprising:
[0007] Acquire the image to be detected;
[0008] According to the type of the image to be detected, the image to be detected is preprocessed accordingly to generate a multi-channel image of a corresponding type; wherein the type includes a discrete frame type or a continuous frame type;
[0009] Inputting the multi-channel image into a pre-trained key point detection model; wherein the pre-trained key point detection model includes a regression branch and an inter-frame residual branch set in parallel;
[0010] When the image to be detected belongs to the discrete frame type, obtaining a regression prediction value output by the regression branch, and determining the coordinates of a plurality of key points in the image to be detected based on the regression prediction value;
[0011] When the image to be detected belongs to the continuous frame type, the regression prediction value output by the regression branch and the residual prediction value output by the inter-frame residual branch are obtained, and the coordinates of multiple key points in the image to be detected are determined based on the regression prediction value and the residual prediction value.
[0012] A device for detecting key points of an image, comprising:
[0013] An image acquisition module, used for acquiring an image to be detected;
[0014] A preprocessing module, used for performing corresponding preprocessing on the image to be detected according to the type of the image to be detected, so as to generate a multi-channel image of a corresponding type; wherein the type includes a discrete frame type or a continuous frame type;
[0015] A model prediction module, used for inputting the multi-channel image into a pre-trained key point detection model; wherein the pre-trained key point detection model includes a regression branch and an inter-frame residual branch arranged in parallel; and
[0016] A key point determination module is used to obtain the regression prediction value output by the regression branch when the image to be detected belongs to the discrete frame type, and determine the coordinates of multiple key points in the image to be detected based on the regression prediction value; and when the image to be detected belongs to the continuous frame type, obtain the regression prediction value output by the regression branch and the residual prediction value output by the inter-frame residual branch, and determine the coordinates of multiple key points in the image to be detected based on the regression prediction value and the residual prediction value.
[0017] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0018] Acquire the image to be detected;
[0019] According to the type of the image to be detected, the image to be detected is preprocessed accordingly to generate a multi-channel image of a corresponding type; wherein the type includes a discrete frame type or a continuous frame type;
[0020] Inputting the multi-channel image into a pre-trained key point detection model; wherein the pre-trained key point detection model includes a regression branch and an inter-frame residual branch set in parallel;
[0021] When the image to be detected belongs to the discrete frame type, obtaining a regression prediction value output by the regression branch, and determining the coordinates of a plurality of key points in the image to be detected based on the regression prediction value;
[0022] When the image to be detected belongs to the continuous frame type, the regression prediction value output by the regression branch and the residual prediction value output by the inter-frame residual branch are obtained, and the coordinates of multiple key points in the image to be detected are determined based on the regression prediction value and the residual prediction value.
[0023] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:
[0024] Acquire the image to be detected;
[0025] According to the type of the image to be detected, the image to be detected is preprocessed accordingly to generate a multi-channel image of a corresponding type; wherein the type includes a discrete frame type or a continuous frame type;
[0026] Inputting the multi-channel image into a pre-trained key point detection model; wherein the pre-trained key point detection model includes a regression branch and an inter-frame residual branch set in parallel;
[0027] When the image to be detected belongs to the discrete frame type, obtaining a regression prediction value output by the regression branch, and determining the coordinates of a plurality of key points in the image to be detected based on the regression prediction value;
[0028] When the image to be detected belongs to the continuous frame type, the regression prediction value output by the regression branch and the residual prediction value output by the inter-frame residual branch are obtained, and the coordinates of multiple key points in the image to be detected are determined based on the regression prediction value and the residual prediction value.
[0029] The above-mentioned image key point detection method, device, computer equipment and storage medium generate corresponding multi-channel images according to different types of images to be detected, so that different types of images to be detected, including discrete frame types or continuous frame types, can have multi-channel images of unified format, so as to use the same key point detection model to be compatible with the detection of different types of images to be detected. During detection, when the image to be detected belongs to the discrete frame type, the coordinates of multiple key points in the image to be detected are determined according to the regression prediction value output by the regression branch, so as to achieve accurate and efficient detection of key points in the discrete frame image; and when the image to be detected belongs to the continuous frame type, the coordinates of multiple key points in the image to be detected are comprehensively determined based on the regression prediction value output by the regression branch and the residual prediction value output by the inter-frame residual branch, so as to detect the key point position of the image to be detected, and at the same time, the stability of the key point position between adjacent frames is achieved, so as to avoid excessive jumps in the key point position between adjacent frames. Therefore, the scheme of the present application can take into account the key point detection of both discrete frame and continuous frame types of images, and at the same time improve the stability of the key point detection results of adjacent frames of continuous frame type images. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a schematic diagram of key points of a face in one embodiment;
[0031] Figure 2 is a schematic diagram of a terminal in one embodiment;
[0032] Figure 3 Schematic diagram of a flow chart of an image key point detection method in one embodiment;
[0033] Figure 4 is a schematic diagram of an image key point detection method in one embodiment;
[0034] Figure 5 A schematic diagram of generating a multi-channel image of a discrete frame type in one embodiment;
[0035] Figure 6 A schematic diagram of generating a multi-channel image of a continuous frame type in one embodiment;
[0036] Figure 7 is a schematic diagram of a key point detection model in one embodiment;
[0037] Figure 8 is a schematic diagram of a key point detection model in another embodiment;
[0038] Fig. 9 is a schematic diagram of a key point detection model in another embodiment;
[0039] Fig.10 A schematic diagram of a process flow of a key point detection model training process in one embodiment;
[0040] Fig.11 is a schematic diagram of a training process of a key point detection model in one embodiment;
[0041] Fig.12 is a structural block diagram of an image key point detection device in one embodiment;
[0042] Fig.13 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0043] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0044] With the popularity of deep learning, more and more key point detection algorithms are implemented using convolutional neural networks. For discrete frame images, key points can be detected through regression, heat map and other methods. For continuous frames such as videos, jitter can be eliminated through optical flow tracking, optical flow superposition of current frames for training and other methods. However, there are currently few algorithms that can adaptively take into account the key point detection of both discrete frame images and continuous frame images.
[0045] When human eyes observe objects, they usually ignore tiny jitters due to the existence of visual temporal phenomena. However, for conventional algorithms targeting single-frame images, each detection only sees the information of the current frame and ignores the differences between adjacent frames, resulting in jitters in the key point detection results between adjacent frames.
[0046] In response to the defects existing in the prior art, the present application provides an image key point detection method, which can simultaneously take into account the key point detection of two types of images, discrete frames and continuous frames, while improving the stability of adjacent frame detection results of continuous frame type images.
[0047] The image key point detection method provided in this application can be applied to Figure 2In the terminal 100 shown, the terminal 100 may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices. Among them, the image key point detection method may include a training phase and a detection phase. In the training phase, the key point detection model may be trained using a training set to obtain a trained key point detection model. The training process may be performed in the terminal 100, and the trained key point detection model may be stored in the terminal 100 for use; or, the training process may also be performed on an external device such as another terminal or server, and after obtaining the trained key point detection model, the trained key point detection model may be loaded from another external device to the terminal 100 for use. In the case where the trained key point detection model has been loaded into the terminal 100, in the detection phase, the terminal 100 may perform the image key point detection method in the embodiment of the present application on the image to be detected to determine the coordinates of multiple key points in the image to be detected.
[0048] In various examples of the present application, for ease of understanding, the target object is a human face as an example for explanation. For example, the key points of the present application include 106 predefined key points distributed in the eyebrows, eyes, nose, mouth and facial contour areas of the human face, and in the detection stage, the positions of the 106 key points in the image to be detected are detected. However, the present application can also be applied to the case where the target object is other objects, and the present application can also be set with more or fewer key points, or key points at different positions.
[0049] In one embodiment, Figure 3 and Figure 4 As shown, a method for detecting key points in an image is provided. Figure 2 The terminal 100 in FIG. 1 is taken as an example to illustrate, and the following steps S310-S350 are included:
[0050] Step S310, obtaining an image to be detected.
[0051] In one embodiment, step S310 includes: acquiring an original image; performing target object detection on the original image to determine the area where the target object is located; intercepting a sub-image containing the target object from the area; and resetting the size of the sub-image to the target size to determine the image to be detected.
[0052] For example, the original image is, for example, an image captured by the terminal 100 or an external device, and there is a target object to be detected in the original image. The target object may be, for example, a face. Accordingly, a bounding box where the face is located may be detected from the original image by a face recognition algorithm. A sub-image in the bounding box is intercepted, and the size of the sub-image is reset to a target size, which may be, for example, a resolution of 112×112, thereby obtaining an image to be detected with a resolution of 112×112.
[0053] Step S320, performing corresponding preprocessing on the image to be detected according to the type of the image to be detected, so as to generate a multi-channel image of the corresponding type; wherein the type includes a discrete frame type or a continuous frame type.
[0054] In one embodiment, the image to be detected is a multi-color channel image, and step S320 includes: when the image to be detected belongs to a discrete frame type, generating a first grayscale image with the same resolution as the image to be detected and with pixel values of zero as a grayscale residual image, and combining the multi-color channel images with the grayscale residual image to generate a multi-channel image of a discrete frame type; when the image to be detected belongs to a continuous frame type, determining a second grayscale image of the image to be detected, obtaining a previous frame image adjacent to the image to be detected, determining a third grayscale image of the previous frame image, determining a difference between the second grayscale image and the third grayscale image as a grayscale residual image, and combining the multi-color channel images with the grayscale residual image to generate a multi-channel image of a continuous frame type.
[0055] In one embodiment, the multi-color channel image may be a three-channel image of red, green, and blue (Red, Green, Blue, RGB), and accordingly, the generated multi-channel image is a four-channel image of red, green, blue, and grayscale residual. In other embodiments, the multi-color channel image may also be an image of other image modes such as a three-channel image of hue, saturation, and brightness (Hues, Saturation, Brightness, HSB), and accordingly, the generated multi-channel image may include a multi-channel image formed by combining grayscale residual images on the basis of images of other image modes.
[0056] For example, after obtaining the image to be detected with a resolution of 112×112, it can be determined whether the category of the image to be detected is a discrete frame type or a continuous frame type. A discrete frame type image refers to an image that has no significant association with other images, such as an image format image, and a continuous frame type image refers to an image that is associated with adjacent images in the time domain, such as a video format file or an image in a video stream.
[0057] For example, taking the image to be detected as an RGB three-channel image as an example, according to different categories of the image to be detected, the following two methods can be used to preprocess the image to be detected:
[0058] (1) See Figure 5 As shown in the figure, when the image to be detected with a resolution of 112×112 belongs to the discrete frame type, a zero-value image with a resolution of 112×112 with all pixel values 0 is generated, and the first grayscale image of the zero-value image is determined as the grayscale residual image. It can be understood that the value of each pixel in the first grayscale image is also zero. The images of the three RGB channels of the image to be detected in the current frame are combined into the grayscale residual image to generate a four-channel image with a resolution of 112×112.
[0059] (2) See Figure 6 As shown, when the image to be detected with a resolution of 112×112 belongs to the continuous frame type, the second grayscale image of the image to be detected is determined, the previous frame image adjacent to the image to be detected in the continuous frame sequence of the current frame is obtained, the third grayscale image of the previous frame image is determined, and the difference between the second grayscale image and the third grayscale image is determined as the grayscale residual image. The grayscale residual image is combined with the images of the three RGB channels of the image to be detected in the current frame, thereby generating a four-channel image with a resolution of 112×112.
[0060] Step S330, inputting the multi-channel image into a pre-trained key point detection model; wherein the pre-trained key point detection model includes a regression branch and an inter-frame residual branch that are set in parallel.
[0061] Before the detection phase begins, a key point detection model can be pre-trained on the terminal 100 or an external device and stored in the terminal 100, so that in this step, the terminal 100 can input the pre-processed multi-channel image into the trained key point detection model.
[0062] As will be discussed later, in the training phase, the key point detection model includes a regression branch, an inter-frame residual branch, and a type classification branch that are set in parallel, so that the regression branch, the inter-frame residual branch, and the type classification branch in the key point detection model can be trained together to comprehensively determine the layer weights in the regression branch, the inter-frame residual branch, and the type classification branch. After the training is completed, in the detection phase, referring to the following steps S340 and S350, the coordinates of the key points can only use the output values of the regression branch and the inter-frame residual branch, without using the output value of the type classification branch. Therefore, the key point detection model generated after the training is completed for the detection stage can delete the type classification branch and only retain the regression branch and the inter-frame residual branch, that is, the pre-trained key point detection model described in this step S330 can only include both the regression branch and the inter-frame residual branch to save storage and computing resources; alternatively, the key point detection model generated after the training is completed for the detection stage can also retain the regression branch, the inter-frame residual branch and the type classification branch, that is, the pre-trained key point detection model described in this step S330 can also include the regression branch, the inter-frame residual branch and the type classification branch. When determining the coordinates of the key points, the output values of the regression branch and the inter-frame residual branch are selected according to the type of the image to be detected, and the output value of the type classification branch is ignored.
[0063] For example, the generated key point detection model for the detection stage retains the regression branch, the inter-frame residual branch and the type classification branch, and the key points include 106 key points defined on the face. The image to be detected with a resolution of 112×112 is input into the trained key point detection model, and after the forward transmission of the regression branch, the inter-frame residual branch and the type classification branch in the key point detection model, the regression branch outputs a regression prediction value H of a one-dimensional vector of 1×212. loc , the regression prediction value H loc The inter-frame residual branch outputs a 1×212 one-dimensional vector residual prediction value, which represents the difference in the key point position between adjacent frames, and the type classification branch outputs a 1×1 classification prediction value H cls , the classification prediction value H cls Characterizes the type of image to be detected.
[0064] Step S340, when the image to be detected belongs to a discrete frame type, obtain the regression prediction value output by the regression branch, and determine the coordinates of multiple key points in the image to be detected based on the regression prediction value.
[0065] For example, see Figure 4 , if the image to be detected belongs to the discrete frame type, then in this step, only the regression prediction value H output by the regression branch can be used locTo determine the coordinates of the key points. The regression prediction value H loc For example, a 1×212 one-dimensional vector has 212 components, and every two adjacent components are combined into a two-dimensional coordinate in the (x, y) format to obtain 106 first two-dimensional coordinates. The 106 first two-dimensional coordinates correspond sequentially to the coordinates of 106 predefined key points in the face.
[0066] Step S350, when the image to be detected belongs to a continuous frame type, obtain the regression prediction value output by the regression branch and the residual prediction value output by the inter-frame residual branch, and determine the coordinates of multiple key points in the image to be detected based on the regression prediction value and the residual prediction value.
[0067] For example, see Figure 4 , if the image to be detected belongs to the continuous frame type, then in this step, the regression prediction value H output by the regression branch can be used loc And the residual prediction value H output by the inter-frame residual branch diff To comprehensively determine the coordinates of the key points. The regression prediction value H loc For example, a 1×212 one-dimensional vector has 212 components, and each two adjacent components are combined into a two-dimensional coordinate in the (x, y) format, thereby obtaining 106 first two-dimensional coordinates. Similarly, the residual prediction value H diff It is also a 1×212 one-dimensional vector having 212 components, and each two adjacent components form a two-dimensional coordinate in the (x, y) format, thereby obtaining 106 second two-dimensional coordinates. The 106 first two-dimensional coordinates and the 106 second two-dimensional coordinates may correspond to the 106 key points predefined in the face in sequence, so that each of the 106 key points predefined in the face may correspond to a first two-dimensional coordinate and a second two-dimensional coordinate, so that the coordinates of each key point may be a third two-dimensional coordinate determined comprehensively based on a corresponding first two-dimensional coordinate and a second two-dimensional coordinate.
[0068] In one embodiment, in step S350, when the image to be detected belongs to a continuous frame type, the coordinates of multiple key points in the image to be detected are determined based on the regression prediction value and the residual prediction value, including: taking the regression prediction value of the image to be detected as the direct regression value of the image to be detected; obtaining the regression prediction value of the previous frame image adjacent to the image to be detected and the output corresponding to the regression branch after the key point detection model is input, and taking the sum of the regression prediction value of the previous frame image and the residual prediction value of the image to be detected as the inferred regression value of the image to be detected; calculating the weighted sum of the direct regression value and the inferred regression value as the key point inference value of the image to be detected; and determining the coordinates of multiple key points in the image to be detected based on the key point inference value.
[0069] For example, when the image to be detected in the current frame belongs to the continuous frame type, the key point inference value of the image to be detected will be equal to the weighted sum of the direct regression value of the current frame and the inferred regression value of the current frame. The direct regression value of the current frame is the regression prediction value H output by the regression branch of the current frame loc , and the inferred regression value of the current frame can be obtained by the regression prediction value H output by the regression branch of the previous frame loc Add the residual prediction value H output by the inter-frame residual branch of the current frame diff The specific calculation can be seen in the following formula:
[0070]
[0071] In the above formula, k represents the current frame (i.e., the kth frame), H f,k Represents the estimated value of the key points of the current frame, H loc,k-1 represents the regression prediction value of the previous frame (i.e., the (k-1)th frame), H diff,k Represents the residual prediction value of the current frame, H loc,k Represents the regression prediction value of the current frame. α and β represent weighted weights respectively, and α+β=1. The values of α and β can be adjusted according to the situation. For example, α=0.5 and β=0.5 can be taken.
[0072] After the key point inference value of the image to be detected in the current frame is determined by the above formula, the key point inference value is also exemplarily a 1×212 one-dimensional vector, which has 212 components. Every two adjacent components are combined into a two-dimensional coordinate in the (x, y) format to obtain 106 third two-dimensional coordinates. The 106 third two-dimensional coordinates correspond sequentially to the coordinates of the 106 predefined key points in the face.
[0073] The above-mentioned image key point detection method generates corresponding multi-channel images according to different types of images to be detected, so that different types of images to be detected, including discrete frame types or continuous frame types, can have multi-channel images of unified format, so as to use the same key point detection model to be compatible with the detection of different types of images to be detected. During detection, when the image to be detected belongs to a discrete frame type, the coordinates of multiple key points in the image to be detected are determined according to the regression prediction value output by the regression branch, so as to achieve accurate and efficient detection of key points in the discrete frame image; and when the image to be detected belongs to a continuous frame type, the coordinates of multiple key points in the image to be detected are comprehensively determined based on the regression prediction value output by the regression branch and the residual prediction value output by the inter-frame residual branch, so as to detect the key point position of the image to be detected while having the stability of the key point position between adjacent frames, and avoid excessive jumps in the key point position between adjacent frames. Therefore, the scheme of the present application can take into account the key point detection of both discrete frames and continuous frames, while improving the stability of the key point detection results of adjacent frames of continuous frame type images.
[0074] In one embodiment, after the above step S310 and before step S320, the image key point detection method further includes: performing contour enhancement processing on the image to be detected to generate the image to be detected after contour enhancement.
[0075] Furthermore, contour enhancement processing is performed on the image to be detected to generate an image to be detected after contour enhancement, including: using a high-contrast retention algorithm to perform contour texture extraction processing on the image to be detected to obtain a contour texture map; using a strong light overlay algorithm to overlay the contour texture map on the image to be detected to obtain a contour texture deepening map; and using the strong light overlay algorithm again to overlay the contour texture map on the contour texture deepening map, thereby generating an image to be detected after contour enhancement.
[0076] For example, the operation of the first strong light superposition algorithm is as follows:
[0077]
[0078] In the above formula, src3 represents the pixel value of each pixel in the contour texture deepening image, src1 represents the pixel value of the corresponding pixel in the image to be detected, and src2 represents the pixel value of the corresponding pixel in the contour texture image.
[0079] For example, the calculation of the second strong light superposition algorithm is as follows:
[0080]
[0081] In the above formula, Dst represents the pixel value of each pixel in the image to be detected after contour enhancement, src3 represents the pixel value of the corresponding pixel in the contour texture deepening image, and src2 represents the pixel value of the corresponding pixel in the contour texture image.
[0082] When performing key point detection of objects such as faces, if the contours of the key points of the object are not obvious, it will also affect the accuracy of key point detection, and may cause large deviations in the key points detected near the contours. In the above embodiments of the present application, the boundary of the object contour is improved by contrast enhancement and back-addition. First, the contour texture of, for example, the facial features is extracted by a high-contrast retention algorithm, and then the contour texture is further deepened by two strong light superposition algorithms and linearly superimposed back on the image to be detected, thereby obtaining the image to be detected after contour enhancement. Using the image to be detected after contour enhancement to perform key point detection can make the key point results detected in the contour area more accurate and stable.
[0083] The above describes the method steps that the terminal 100 may execute in the detection stage, wherein in step S306 of the detection stage, a pre-trained key point detection model is used, which requires that in a training stage before the detection stage, the architecture of the key point detection model is pre-constructed, and the constructed key point detection model is trained to optimize and determine the layer weights of the key point detection model to obtain a trained key point detection model.
[0084] In one embodiment, Figure 7 As shown, in the training stage, the key point detection model includes a regression branch, an inter-frame residual branch and a type classification branch set in parallel. The regression branch, the inter-frame residual branch and the type classification branch respectively receive the multi-channel image obtained by preprocessing, and output the regression prediction value, the residual prediction value and the classification prediction value respectively. By setting the regression branch, the inter-frame residual branch and the type classification branch in parallel in the key point detection model in the training stage, the trained key point detection model can take into account the regression of the image key point position, the inter-frame residual between adjacent frames in continuous frames and the characteristics of different types of images. As described above, in the detection stage after the training stage, the key point detection model may only include the regression branch and the inter-frame residual branch set in parallel, or may also include the regression branch, the inter-frame residual branch and the type classification branch but only the output values of the regression branch and the inter-frame residual branch need to be used.
[0085] In one embodiment, Figure 8As shown, in the training stage, the key point detection model also includes a backbone network, the backbone network receives the multi-channel image obtained by preprocessing, the backbone network outputs an intermediate feature map, and feeds the intermediate feature map into the regression branch, the inter-frame residual branch and the type classification branch respectively. In this embodiment, by setting a backbone network shared before the regression branch, the inter-frame residual branch and the type classification branch, the multi-channel image can be pre-extracted by the backbone network to obtain an intermediate feature map, and then the intermediate feature map is input into different branches to perform different prediction operations, which can effectively improve the feature extraction efficiency of the multi-channel image and improve the overall computing efficiency of the model. It can be understood that if the key point detection model also includes a backbone network in the training stage, then the key point detection model also includes the backbone network in the detection stage after the training stage.
[0086] In one embodiment, the regression branch includes a plurality of first convolutional neural network layers (Conv), a flattening layer (Flatten), and a first fully connected layer (FC) connected in series; the inter-frame residual branch includes a plurality of second convolutional neural network layers, a flattening layer, and a second fully connected layer connected in series; the type classification branch includes a plurality of third convolutional neural network layers, a flattening layer, and a third fully connected layer connected in series.
[0087] For example, Fig. 9 As shown, each of the regression branch, the inter-frame residual branch, and the type classification branch may have two serial convolutional neural network layers; however, in other embodiments, each branch may also have a different number of convolutional neural network layers.
[0088] In one embodiment, the first fully connected layer and the second fully connected layer each have a first number of output neurons, where the first number is determined according to a total number of the plurality of key points.
[0089] Specifically, the first number may be equal to the total number of key points multiplied by the parameter amount of the coordinates of each key point.
[0090] For example, Fig. 9 As shown in , when the defined key points include 106 key points, the coordinates (x, y) of each key point have two parameters x and y, so the coordinates of the 106 key points need to be determined by 106×2=212 parameters. Therefore, it can be designed that the first fully connected layer has 212 output neurons, and the second fully connected layer also has 212 output neurons.
[0091] In one embodiment, the third fully connected layer has a second number of output neurons, and the second number depends on the loss function category of the type classification branch.
[0092] For example, Fig. 9As shown in , when the activation function of the type classification branch is the Sigmoid activation function and the loss function is the Binary Cross Entropy Loss function, the third fully connected layer has a single output neuron. Thus, the classification prediction value H output by the third fully connected layer is cls Including the output value of the single output neuron, the output value of the single output neuron represents the probability y1 that the image to be detected belongs to the discrete frame type, 0≤y1≤1. By comparing the probability y1 with the given threshold T, the type of the image to be detected can be determined. When y1≥T, it can be determined that the image to be detected belongs to the discrete frame type. Conversely, when y1﹤T, it can be determined that the image to be detected belongs to the continuous frame type. Alternatively, the output value of the single output neuron can also represent the probability y1 that the image to be detected belongs to the continuous frame type. Then, when y1≥T, it can be determined that the image to be detected belongs to the continuous frame type. Conversely, when y1﹤T, it can be determined that the image to be detected belongs to the discrete frame type. T can be 0.5 for example.
[0093] In another example, when the activation function of the type classification branch is the Softmax activation function and the loss function is the Categorical Cross Entropy Loss function, the third fully connected layer may also have two output neurons. Thus, the classification prediction value H output by the third fully connected layer is cls It includes two output values of two output neurons, where the output value of one output neuron represents the probability y2 that the image to be detected belongs to the discrete frame type, 0≤y2≤1. At the same time, the output value of the other output neuron represents the probability y3 that the image to be detected belongs to the continuous frame type, 0≤y3≤1. The sum of these two probabilities (y2+y3) is 1. When y2≥y3, it can be determined that the image to be detected belongs to the discrete frame type. Conversely, when y2﹤y3, it can be determined that the image to be detected belongs to the continuous frame type.
[0094] In a specific example, see Fig. 9 As shown in FIG, the key point detection model includes a backbone network, and a regression branch, an inter-frame residual branch, and a type classification branch that are serially connected to the backbone network and parallel to each other. The backbone network receives the image to be detected and outputs an intermediate feature map, which is fed into the regression branch, the inter-frame residual branch, and the type classification branch respectively. The regression branch, the inter-frame residual branch, and the type classification branch output regression prediction values, residual prediction values, and classification prediction values respectively.
[0095] Among them, see Fig. 9As shown, the backbone network is used to perform preliminary feature extraction on the input image to be detected. The backbone network may include several serial and / or parallel convolutional neural network layers, and the backbone network may also include several pooling layers, batch normalization layers, and other layers. For example, the backbone network may include a lightweight network V2 (MobileNetV2, see MobileNetV2: Inverted Residuals and Linear Bottlenecks, Mark Sandler, et.al., The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 4510-4520). It can be understood that the specific structure of the backbone network can have a variety of different structures, as long as it can achieve preliminary feature extraction of the image to be detected, and the size and number of channels of the output feature map match the input end of the subsequent regression branch, inter-frame residual branch, and type classification branch, and the present application is not limited thereto. For example, the image to be detected is a four-channel image with a resolution of 112×112 (which can be expressed as "1×112×112×4"). After the four-channel image with a resolution of 112×112 is input into the backbone network, a 256-channel image with a resolution of 7×7 (which can be expressed as "1×7×7×256") can be output. The 256-channel image with a resolution of 7×7 is the intermediate feature map.
[0096] For example, see Fig. 9As shown, the regression branch includes two first convolutional network layers, a flattening layer, and a first fully connected layer in series. The dimension of each first convolutional network layer is "256×3×3×256", which means that the filter depth of each first convolutional network layer is 256 (the filter depth is determined by the number of channels of the input graph of the layer), the convolution kernel size is 3×3, the number of filters is 256 (the number of filters will determine the number of channels of the output graph of the layer), and the stride is 2. A 7×7 resolution 256-channel image is input into the first first convolutional network layer. After the convolution operation of the first first convolutional network layer, a 4×4 resolution 256-channel image is obtained. The 4×4 resolution 256-channel image is input into the second first convolutional network layer. After the convolution operation of the second first convolutional network layer, a 1×1 resolution 256-channel image is obtained. Then the 1×1 resolution 256-channel image is flattened by the flattening layer to obtain a 1×256 one-dimensional vector. The dimension of the first fully connected layer is "256×212", which means that the input parameter of the first fully connected layer is 256, and it has 212 neurons (the number of neurons determines the output parameter of this layer). Therefore, after a 1×256 one-dimensional vector is input into the first fully connected layer, a 1×212 one-dimensional vector is output. This 1×212 one-dimensional vector is the regression prediction value H loc .
[0097] The inter-frame residual branch has a similar structure to the regression branch. The inter-frame residual branch includes two second convolutional network layers, a flattening layer, and a second fully connected layer in series. Similarly, the dimension of each second convolutional network layer is also "256×3×3×256", the step size is also 2, and the dimension of the second fully connected layer is also "256×212". Therefore, similarly to the above regression branch, a 7×7 resolution 256-channel image is input into the inter-frame residual branch, and after the operation of two second convolutional network layers, a flattening layer, and a second fully connected layer, a 1×212 one-dimensional vector is output. The 1×212 one-dimensional vector is the residual prediction value H. diff .
[0098] Similarly, the type classification branch includes two third convolutional network layers, a flattening layer and a third fully connected layer in sequence. The dimension of each third convolutional network layer is also "256×3×3×256", the step size is also 2, and the dimension of the third fully connected layer is "256×1", which means that the input parameter amount of the third fully connected layer is 256, and it has 1 neuron (the number of neurons determines the output parameter amount of this layer). The activation function of the third fully connected layer is the Sigmoid activation function, so that a 7×7 resolution 256-channel image is input into the inter-frame residual branch, and then passes through two second convolutional network layers and a flattening layer to obtain a 1×256 one-dimensional vector. After the 1×256 one-dimensional vector is input into the third fully connected layer, a 1×1 one-dimensional vector (single parameter, i.e., probability y1) is output. The 1×1 one-dimensional vector is the classification prediction value H cls .
[0099] Forward propagation and backward propagation are the basic processes that neural networks may perform during the testing and training processes. In the model training phase, the model is repeatedly subjected to forward propagation and backward propagation using the training set data to optimize the layer weight values and determine the trained model; in the model testing phase, the model performs forward propagation on the input data to output the prediction results. Among them, forward propagation describes how the input of the neural network is transmitted and processed forward along each layer of the neural network to obtain the output; while backward propagation describes how, after the loss function is determined through forward propagation, it is transmitted backward along each layer of the neural network and the error of each layer is derived to adjust the layer weight values of each layer accordingly.
[0100] For example, the process described above of inputting a multi-channel image into a key point detection model, obtaining an intermediate feature map through a backbone network, respectively inputting the intermediate feature map into a regression branch, an inter-frame residual branch, and a type classification branch, and respectively outputting a regression prediction value, a residual prediction value, and a classification prediction value, is the forward propagation process of the key point detection model. The basic principle of the back propagation process of the key point detection model of the present application is similar to the back propagation process of the existing convolutional neural network. Those skilled in the art can understand the back propagation process of the key point detection model of the present application by combining the back propagation principle of the existing convolutional neural network and the above-mentioned structure of the key point detection model of the present application, so it will not be repeated.
[0101] In the training phase, by training the key point detection model constructed as above, a trained key point detection model can be determined.
[0102] In one embodiment, Fig.10 and Fig.11As shown, before acquiring the image to be detected, in the training stage, the training process of the key point detection model includes the following steps S1010-S1040:
[0103] Step S1010, obtaining a training set; the training set includes a plurality of multi-channel images of discrete frame types and the true value corresponding to each multi-channel image of the discrete frame type (i.e., ground truth, Ground Truth, GT), and a plurality of multi-channel images of continuous frame types and the true value corresponding to each multi-channel image of the continuous frame type.
[0104] In one embodiment, step 1010 includes: step S1011, obtaining a discrete frame set, the discrete frame set includes multiple discrete frame images; generating multi-channel images of multiple discrete frame types corresponding to the multiple discrete frame images, and determining the true value corresponding to the multi-channel image of each discrete frame type; and step S1012, obtaining a continuous frame set, the continuous frame set includes multiple continuous frame images; generating multi-channel images of multiple continuous frame types corresponding to the multiple continuous frame images, and determining the true value corresponding to the multi-channel image of each continuous frame type.
[0105] It can be understood that the multi-channel images of discrete frame type and the multi-channel images of continuous frame type in the training set in this step can have the same resolution, image mode and number of channels as the multi-channel images of discrete frame type and the multi-channel images of continuous frame type in the above-mentioned detection stage, and can be obtained by performing the same preprocessing as in the above-mentioned detection stage on the discrete frame images and the continuous frame images.
[0106] In one embodiment, the true value corresponding to the multi-channel image of each discrete frame type in the training set includes the regression true value of the regression branch and the classification true value of the type classification branch, and the true value corresponding to the multi-channel image of each continuous frame type in the training set includes the regression true value of the regression branch, the residual true value of the inter-frame residual branch, and the classification true value of the type classification branch.
[0107] In one embodiment, step S1011 may include: obtaining a discrete frame set, wherein the discrete frame set includes multiple discrete frame images that are not related to each other, each of the multiple discrete frame images includes an object of interest and has the same resolution and image mode as the above-mentioned image to be detected; for example, each discrete frame image can also be determined by performing object detection on the acquired original image, intercepting a sub-image, and resetting the sub-image size to the target size. Then, extract a discrete frame image from the discrete frame set, and generate a zero-value image with all pixel values of 0, and combine the zero-value image with the extracted discrete frame image into a discrete frame pair, and repeat the operation until all discrete frame images in the discrete frame set are extracted, thereby obtaining N1 discrete frame pairs, wherein the discrete frame image actually captured in each discrete frame pair is defined as F cur , the generated zero-value image is defined as F his For each discrete frame pair, a single-channel first grayscale image with all pixel values zero is generated as the grayscale residual image, and F cur The multi-channel image of the color channel is combined with the gray residual image to generate a multi-channel image, see Figure 5 As shown, N1 multi-channel images are generated. For each multi-channel image, the true value GT of the multi-channel image is calibrated by manual calibration or recognition using known tools, including the regression true value GT of the regression branch. loc And the classification truth value GT of the type classification branch cls .
[0108] For example, each discrete frame image in the discrete frame set is a 112×112 resolution red, green, and blue three-channel image, and each discrete frame type multi-channel image is a 112×112 resolution red, green, blue, and grayscale residual four-channel image. The defined key points include 106 key points. Then, the frame image F can be calibrated manually or through tools. cur The 106 coordinates (x, y) of the 106 key points in the image are arranged in sequence to form a 1×212 one-dimensional vector as the regression true value GT of the multi-channel image of the discrete frame type. loc ;exist Fig. 9 The classification prediction value H output by the third fully connected layer in cls Including the output value of a single output neuron, and the output value of the single output neuron represents the probability y1 that the input image belongs to a discrete frame type, the classification true value GT of the multi-channel image of the discrete frame type cls y1=1.
[0109] In one embodiment, step S1012 may include: acquiring a continuous frame set, wherein the continuous frame set includes a plurality of continuous frame images with a front-to-back association, each of the plurality of continuous frame images includes an object of interest, and has the same resolution and image mode as the above-mentioned image to be detected; for example, each continuous frame image may also be determined by performing object detection on the acquired original image, intercepting a sub-image, and resetting the sub-image size to a target size. Then, extracting two adjacent continuous frame images from the continuous frame set to form a continuous frame pair, repeating the operation until all continuous frame images in the continuous frame set are extracted, thereby obtaining N2 continuous frame pairs, wherein the latter continuous frame image in each continuous frame pair is defined as F cur , the previous continuous frame image is defined as F his For each pair of consecutive frames, F cur Take the grayscale to generate the second grayscale image, and convert F his Take the grayscale to generate the third grayscale image, determine the difference between the second grayscale image and the third grayscale image as the grayscale residual image, and convert F cur The multi-channel image of the color channel is combined with the gray residual image to generate a multi-channel image, see Figure 6 As shown, N2 multi-channel images are generated. For each multi-channel image, the true value GT of the multi-channel image is calibrated by manual calibration or recognition using known tools, including the regression true value GT of the regression branch. loc , the residual true value GT of the gray residual image diff And the classification truth value GT of the type classification branch cls .
[0110] For example, each continuous frame image in the continuous frame set is a 112×112 resolution red, green, and blue three-channel image, and each continuous frame type multi-channel image is a 112×112 resolution red, green, blue, and grayscale residual four-channel image, and the defined key points include 106 key points. Then, a continuous frame pair corresponding to the multi-channel image of the continuous frame type can be calibrated manually or through a tool, and the next continuous frame image F cur The 106 coordinates (x, y) of the 106 key points in the image are arranged in sequence to form a 1×212 one-dimensional vector as the regression true value GT of the multi-channel image of the continuous frame type. loc ; Mark the same continuous frame pair, the previous continuous frame image F his The 106 coordinates (x, y) of the 106 key points in the image are arranged in sequence to form a 1×212 one-dimensional vector as the historical regression true value GT of the multi-channel image of the continuous frame type. his, thereby determining the residual true value GT of the gray residual image corresponding to the multi-channel image of the continuous frame type diff =GT loc -GT his ;exist Fig. 9 The classification prediction value H output by the third fully connected layer in cls Including the output value of a single output neuron, and the output value of the single output neuron represents the probability y1 that the input image belongs to the discrete frame type, the classification true value GT of the multi-channel image of the continuous frame type cls y1=0.
[0111] Furthermore, in one embodiment, efficiency can be improved by performing changes on a single-frame image to simulate the generation of continuous frames. Accordingly, step S1012 may include: acquiring a single-frame image, and determining a true value corresponding to a multi-channel image generated from the single-frame image; performing continuous illumination, color and / or deformation processing on the single-frame image to simulate the generation of a continuous plurality of continuous frame images; preprocessing each of the continuous plurality of continuous frame images to generate a multi-channel image of a plurality of continuous frame types; and determining a true value corresponding to a multi-channel image of a continuous frame type generated from each of the plurality of continuous frame images based on the deformation amount between the plurality of continuous frame images and the single-frame image and based on the true value corresponding to the multi-channel image generated from the single-frame image.
[0112] For determining the true value in the present embodiment, specifically, the position of each key point in the original single-frame image can be calibrated, and according to the deformation amount between the remaining simulated continuous frame images and the original single-frame image, based on the position of each key point in the original single-frame image, the position of each key point in each simulated continuous frame image can be determined, and then according to the method described in the above embodiment, multi-channel images of multiple continuous frame types and the true value corresponding to the multi-channel image of each continuous frame type can be determined from the continuous frame image set consisting of the original single-frame image and the remaining simulated continuous frame images.
[0113] Compared to using traditional videos to obtain continuous frame images and determining the true value of each frame image one by one, in this embodiment, by performing changes on a single frame image to simulate the generation of a continuous frame set, the true value corresponding to the multi-channel image of other continuous frame images can be quickly determined based on the deformation amount between it and the original single frame image, thereby enabling more rapid and efficient generation of multi-channel images of multiple continuous frame types and the true value corresponding to the multi-channel image of each continuous frame type.
[0114] Step S1020, initializing the layer weights in the key point detection model.
[0115] In this step, the layer weights of each layer in the keypoint detection model are initialized.
[0116] Step S1030, using the training set, repeatedly execute the forward propagation to calculate the total loss function and the back propagation to update the layer weights of the key point detection model until the value of the total loss function converges to determine the trained layer weights; wherein the multi-channel image of the discrete frame type acts on the layer weight training of the regression branch and the type classification branch, and the multi-channel image of the continuous frame type acts on the layer weight training of the regression branch, the inter-frame residual branch and the type classification branch.
[0117] Among them, the value convergence of the total loss function can refer to that the loss value calculated by the total loss function falls within the expected loss value range or the number of iterations reaches a predetermined number.
[0118] In one embodiment, the regression branch has a first loss function, which is used to characterize the difference between the regression prediction value output by the regression branch and the regression true value; the inter-frame residual branch has a second loss function, which is used to characterize the difference between the residual prediction value output by the inter-frame residual branch and the residual true value; and the type classification branch has a third loss function, which is used to characterize the difference between the classification prediction value output by the type classification branch and the classification true value.
[0119] For example, the training data in the training set can be divided into multiple batches (Batch), and the training data of each batch is used to perform forward propagation training on the key point detection model to calculate the total loss function and back propagate to update the layer weights until the value of the total loss function converges.
[0120] For example, for a single batch, the first loss function L of the regression branch is loc It can be a smooth L1 loss function (Smooth L1 Loss), the first loss function L loc The definition is as follows:
[0121] x=GT loc -H loc
[0122]
[0123] In the above formula, n is the total number of multi-channel images in the batch, GT loc is the true regression value of each multi-channel image in the batch, H loc is the regression prediction value for each multi-channel image in the batch.
[0124] For a single batch, the second loss function L of the inter-frame residual branch is diff It can also be a smooth L1 loss function (Smooth L1 Loss), the second loss function L diff The definition is as follows:
[0125] x=GT diff -H diff
[0126]
[0127] In the above formula, n is the total number of multi-channel images in the batch, GT diff is the residual true value of each multi-channel image in the batch, H diff is the residual prediction value for each multi-channel image in the batch.
[0128] For a single batch, the third loss function L of the type classification branch is cls It can be a binary cross entropy loss function (Binary Cross Entropy Loss), the third loss function L cls The definition is as follows:
[0129]
[0130] In the above formula, n is the total number of multi-channel images in the batch, GT cls is the true classification value of each multi-channel image in the batch, H cls is the classification prediction value for each multi-channel image in the batch.
[0131] In one embodiment, when a key point detection model is trained using a multi-channel image of a discrete frame type, the total loss function of the key point detection model is determined based on a first loss function and a third loss function; when a key point detection model is trained using a multi-channel image of a continuous frame type, the total loss function of the key point detection model is determined based on the first loss function, the second loss function and the third loss function.
[0132] The total loss function L of the key point detection model total The definition is as follows:
[0133]
[0134] Where L loc Used to ensure the accuracy of key point detection model prediction, L diff It is used to ensure the stability of the key point detection model when predicting the images to be detected of continuous frame types. cls Used to enable the key point detection model to automatically determine the type of the current input image.
[0135] In the above embodiments, for example, in step S1030, the key point detection model performs training using different total loss functions according to different types of training data. This may require the key point detection model to know the type of the input training data before, for example, step S1030. In one embodiment, the key point detection model may determine the type to which the multi-channel image used for training belongs based on the classification prediction value output by the type classification branch. Exemplarily, when the activation function of the type classification branch is the Sigmoid activation function and the loss function is the Binary Cross Entropy Loss, the classification prediction value H output by the type classification branch cls may be the probability y1 that the image to be detected belongs to the discrete frame type, where 0 ≤ y1 ≤ 1. When y1 ≥ T, it can be determined that the multi-channel image used for training belongs to the discrete frame type. When y1 < T, it can be determined that the multi-channel image used for training belongs to the continuous frame type. T can be taken as 0.5 exemplarily. In any step of the above other embodiments, the key point detection model or the terminal 100 may also determine the type to which the multi-channel image input to the model belongs based on the classification prediction value output by the type classification branch when needed.
[0136] Step S1040: Determine the pre-trained key point detection model based on the trained layer weights.
[0137] Through the training of the foregoing steps, the optimized layer weights can be determined. The key point detection model defined by the optimized layer weights is the trained key point detection model. The trained key point detection model can be stored in the terminal 100 for use in detecting the key point positions in the image to be detected in the subsequent detection stage.
[0138] After the above step S310 and before step S320 in the detection stage, when the image key point detection method further includes performing contour enhancement processing on the image to be detected to generate the image to be detected with enhanced contours, correspondingly, in step S1010 of the training stage, contour enhancement processing can also be performed on each image in the training set to further improve the key point detection accuracy of the model.
[0139] In one embodiment, after obtaining the discrete frame set in the above step S1011 and before generating the multi-channel image, the training process further includes a step of performing contour enhancement processing on each discrete frame image to generate the discrete frame image with enhanced contours; after obtaining the continuous frame set in the above step S1012 and before generating the multi-channel image, the training process further includes a step of performing contour enhancement processing on each continuous frame image to generate the continuous frame image with enhanced contours.
[0140] The step of performing contour enhancement processing on each discrete frame image in the training process to generate a discrete frame image after contour enhancement, and the step of performing contour enhancement processing on each continuous frame image to generate a continuous frame image after contour enhancement in this embodiment are similar to the step of performing contour enhancement processing on the image to be detected in the above-mentioned detection stage to generate the image to be detected after contour enhancement, and therefore will not be repeated here.
[0141] It should be understood that although Figure 3 and Fig.10 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 3 and Fig.10 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0142] In one embodiment, Fig.12 As shown, an image key point detection device 1200 is provided, comprising: an image acquisition module 1210, a preprocessing module 1220, a model prediction module 1230 and a key point determination module 1240, wherein:
[0143] An image acquisition module 1210 is used to acquire an image to be detected;
[0144] A preprocessing module 1220 is used to perform corresponding preprocessing on the image to be detected according to the type of the image to be detected, so as to generate a multi-channel image of the corresponding type; wherein the type includes a discrete frame type or a continuous frame type;
[0145] A model prediction module 1230 is used to input the multi-channel image into a pre-trained key point detection model; wherein the pre-trained key point detection model includes a regression branch and an inter-frame residual branch set in parallel; and
[0146] The key point determination module 1240 is used to obtain the regression prediction value output by the regression branch when the image to be detected belongs to a discrete frame type, and determine the coordinates of multiple key points in the image to be detected based on the regression prediction value; and when the image to be detected belongs to a continuous frame type, obtain the regression prediction value output by the regression branch and the residual prediction value output by the inter-frame residual branch, and determine the coordinates of multiple key points in the image to be detected based on the regression prediction value and the residual prediction value.
[0147] For the specific definition of the image key point detection device 1200, please refer to the definition of the image key point detection method above, which will not be repeated here. The various modules in the above-mentioned image key point detection device 1200 can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0148] In one embodiment, a computer device is provided. The computer device may be a terminal 100, and its internal structure diagram may be as shown in FIG. Fig.13 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for detecting key points of an image is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a key, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.
[0149] Those skilled in the art will understand that Fig.13 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0150] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:
[0151] Acquire the image to be detected;
[0152] According to the type of the image to be detected, the image to be detected is preprocessed accordingly to generate a multi-channel image of the corresponding type; wherein the type includes a discrete frame type or a continuous frame type;
[0153] Inputting the multi-channel image into a pre-trained key point detection model; wherein the pre-trained key point detection model includes a regression branch and an inter-frame residual branch set in parallel;
[0154] When the image to be detected belongs to a discrete frame type, a regression prediction value output by the regression branch is obtained, and based on the regression prediction value, coordinates of multiple key points in the image to be detected are determined;
[0155] When the image to be detected is of a continuous frame type, the regression prediction value output by the regression branch and the residual prediction value output by the inter-frame residual branch are obtained, and the coordinates of multiple key points in the image to be detected are determined based on the regression prediction value and the residual prediction value.
[0156] In other embodiments, when the processor executes the computer program, it also implements the steps of the image key point detection method of any of the above embodiments.
[0157] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0158] Acquire the image to be detected;
[0159] According to the type of the image to be detected, the image to be detected is preprocessed accordingly to generate a multi-channel image of the corresponding type; wherein the type includes a discrete frame type or a continuous frame type;
[0160] Inputting the multi-channel image into a pre-trained key point detection model; wherein the pre-trained key point detection model includes a regression branch and an inter-frame residual branch set in parallel;
[0161] When the image to be detected belongs to a discrete frame type, a regression prediction value output by the regression branch is obtained, and based on the regression prediction value, coordinates of multiple key points in the image to be detected are determined;
[0162] When the image to be detected is of a continuous frame type, the regression prediction value output by the regression branch and the residual prediction value output by the inter-frame residual branch are obtained, and the coordinates of multiple key points in the image to be detected are determined based on the regression prediction value and the residual prediction value.
[0163] In other embodiments, when the computer program is executed by a processor, the steps of the image key point detection method in any of the above embodiments are also implemented.
[0164] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0165] The technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0166] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A method for detecting key points of an image, the method include: Acquire the image to be detected; According to the type of the image to be detected, the image to be detected is preprocessed accordingly to generate a multi-channel image of the corresponding type; wherein the type includes a discrete frame type or a continuous frame type, and wherein the image to be detected is a multi-color channel image, the multi-channel image of the discrete frame type includes the multi-color channel image and a first grayscale image with the same resolution as the image to be detected, and the multi-channel image of the continuous frame type includes the multi-color channel image and a grayscale residual image generated based on the difference between the image to be detected and a grayscale image of a previous frame image adjacent to the image to be detected; Inputting the multi-channel image into a pre-trained key point detection model; wherein the pre-trained key point detection model includes a regression branch and an inter-frame residual branch set in parallel; When the image to be detected belongs to the discrete frame type, obtaining a regression prediction value output by the regression branch, and determining the coordinates of a plurality of key points in the image to be detected based on the regression prediction value; When the image to be detected belongs to the continuous frame type, the regression prediction value output by the regression branch and the residual prediction value output by the inter-frame residual branch are obtained, and the coordinates of multiple key points in the image to be detected are determined based on the regression prediction value of the image to be detected and the inferred regression value which is the sum of the regression prediction value of the previous frame image and the residual prediction value of the image to be detected.
2. The method according to claim 1, It is characterized in that The step of preprocessing the image to be detected according to the type of the image to be detected to generate a multi-channel image of a corresponding type includes: When the image to be detected belongs to the discrete frame type, a first grayscale image with the same resolution as the image to be detected and the pixel values are all zero is generated as a grayscale residual image, and the multi-color channel image is combined with the grayscale residual image to generate a multi-channel image of the discrete frame type; When the image to be detected belongs to the continuous frame type, a second grayscale image of the image to be detected is determined, a previous frame image adjacent to the image to be detected is acquired, a third grayscale image of the previous frame image is determined, a difference between the second grayscale image and the third grayscale image is determined as a grayscale residual image, and the multi-color channel image is combined with the grayscale residual image to generate a multi-channel image of a continuous frame type.
3. The method according to claim 1, It is characterized in that The multi-color channel image is a three-channel image of red, green and blue, and the generated multi-channel image is a four-channel image of red, green, blue and grayscale residual.
4. The method according to claim 1, It is characterized in that When the image to be detected belongs to the continuous frame type, the coordinates of multiple key points in the image to be detected are determined based on the regression prediction value of the image to be detected and the inferred regression value which is the sum of the regression prediction value of the previous frame image and the residual prediction value of the image to be detected, including: Using the regression prediction value of the image to be detected as the direct regression value of the image to be detected; Obtaining the regression prediction value of the previous frame image adjacent to the image to be detected, which is outputted by the regression branch corresponding to the key point detection model, and taking the sum of the regression prediction value of the previous frame image and the residual prediction value of the image to be detected as the inferred regression value of the image to be detected; Calculating a weighted sum of the direct regression value and the inferred regression value as a key point inference value of the image to be detected; Based on the key point inference values, coordinates of multiple key points in the image to be detected are determined.
5. The method according to claim 1, It is characterized in that According to the type of the image to be detected, before performing corresponding preprocessing on the image to be detected to generate a multi-channel image of a corresponding type, the method further includes: Performing contour enhancement processing on the image to be detected to generate a contour-enhanced image to be detected.
6. The method according to any one of claims 1 to 5, It is characterized in that In the training phase of the key point detection model, the key point detection model includes a regression branch, an inter-frame residual branch and a type classification branch which are arranged in parallel.
7. The method according to claim 6, It is characterized in that In the training stage, the key point detection model also includes a backbone network, which receives the multi-channel image, outputs an intermediate feature map, and feeds the intermediate feature map into the regression branch, the inter-frame residual branch and the type classification branch respectively.
8. The method according to claim 6, It is characterized in that The regression branch includes a plurality of first convolutional neural network layers, a flattening layer and a first fully connected layer in series; The inter-frame residual branch includes a plurality of second convolutional neural network layers, a flattening layer and a second fully connected layer in series; The type classification branch includes a plurality of third convolutional neural network layers, a flattening layer and a third fully connected layer connected in series.
9. The method according to claim 8, It is characterized in that The first fully connected layer and the second fully connected layer both have a first number of output neurons, where the first number is determined according to the total number of the plurality of key points; and / or The third fully connected layer has a second number of output neurons, and the second number depends on the loss function category of the type classification branch.
10. The method according to any one of claims 7 to 9, It is characterized in that Before acquiring the image to be detected, the training process of the key point detection model includes: Acquire a training set; the training set includes a plurality of multi-channel images of discrete frame types and a true value corresponding to each multi-channel image of the discrete frame type, and a plurality of multi-channel images of continuous frame types and a true value corresponding to each multi-channel image of the continuous frame type; Initializing layer weights in the keypoint detection model; Using the training set, the key point detection model is repeatedly executed with forward propagation to calculate the total loss function and with back propagation to update the layer weights until the value of the total loss function converges, so as to determine the trained layer weights; wherein the multi-channel image of the discrete frame type acts on the layer weight training of the regression branch and the type classification branch, and the multi-channel image of the continuous frame type acts on the layer weight training of the regression branch, the inter-frame residual branch and the type classification branch; Based on the trained layer weights, the pre-trained keypoint detection model is determined.
11. The method according to claim 10, It is characterized in that Obtaining a plurality of multi-channel images of consecutive frame types and the true value corresponding to each multi-channel image of consecutive frame type includes: Acquire a single-frame image, and determine a true value corresponding to a multi-channel image generated by the single-frame image; Performing continuous illumination, color and / or deformation processing on the single frame image to simulate and generate a plurality of continuous frame images; Preprocessing each of the plurality of continuous frame images to generate a plurality of multi-channel images of continuous frame types; According to the deformation amount between the multiple continuous frame images and the single frame image, and based on the real value corresponding to the multi-channel image generated by the single frame image, the real value corresponding to the multi-channel image of the continuous frame type generated by each continuous frame image in the multiple continuous frame images is determined accordingly.
12. The method according to claim 10, It is characterized in that The true value corresponding to each multi-channel image of the discrete frame type in the training set includes the regression true value of the regression branch and the classification true value of the type classification branch, and The true value corresponding to the multi-channel image of each continuous frame type in the training set includes the regression true value of the regression branch, the residual true value of the inter-frame residual branch, and the classification true value of the type classification branch.
13. The method according to claim 12, It is characterized in that The regression branch has a first loss function, and the first loss function is used to characterize the difference between the regression prediction value output by the regression branch and the regression true value; The inter-frame residual branch has a second loss function, and the second loss function is used to characterize the difference between the residual prediction value output by the inter-frame residual branch and the residual true value; and The type classification branch has a third loss function, and the third loss function is used to characterize the difference between the classification prediction value output by the type classification branch and the classification true value.
14. The method according to claim 13, when the key point detection model is trained using the multi-channel image of the discrete frame type, the total loss function of the key point detection model is determined based on the first loss function and the third loss function; When the key point detection model is trained using the multi-channel images of the continuous frame type, the total loss function of the key point detection model is determined based on the first loss function, the second loss function and the third loss function.
15. An image key point detection device, It is characterized in that The device comprises: An image acquisition module, used for acquiring an image to be detected; A preprocessing module, configured to perform corresponding preprocessing on the image to be detected according to the type of the image to be detected, so as to generate a multi-channel image of a corresponding type; wherein the type includes a discrete frame type or a continuous frame type, and wherein the image to be detected is a multi-color channel image, the multi-channel image of the discrete frame type includes the multi-color channel image and a first grayscale image of the same resolution as the image to be detected, and the multi-channel image of the continuous frame type includes the multi-color channel image and a grayscale residual image generated based on the difference between the image to be detected and a grayscale image of a previous frame image adjacent to the image to be detected; A model prediction module, used for inputting the multi-channel image into a pre-trained key point detection model; wherein the pre-trained key point detection model includes a regression branch and an inter-frame residual branch arranged in parallel; and A key point determination module is used to obtain the regression prediction value output by the regression branch when the image to be detected belongs to the discrete frame type, and determine the coordinates of multiple key points in the image to be detected based on the regression prediction value; and when the image to be detected belongs to the continuous frame type, obtain the regression prediction value output by the regression branch and the residual prediction value output by the inter-frame residual branch, and determine the coordinates of multiple key points in the image to be detected based on the regression prediction value of the image to be detected and the inferred regression value which is the sum of the regression prediction value of the previous frame image and the residual prediction value of the image to be detected.
16. A computer device comprising a memory and a processor, wherein the memory stores a computer program. It is characterized in that When the processor executes the computer program, the steps of the method according to any one of claims 1 to 14 are implemented.
17. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 14 are implemented.
Citation Information
Patent Citations
Single-stage face detection and key point positioning method and system
CN111079686A
Resnet-3D convolution cattle video target detection method based on balance loss
CN112613428A