Key point detection method and device, computer readable medium and electronic device

By using the continuous frame key point detection model in video combined with current and historical key point information, the real-time and robustness of face key point detection is solved, and the detection efficiency and accuracy are improved.

CN114612976BActive Publication Date: 2025-08-26GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210228398.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-08
Publication Date
2025-08-26
Estimated Expiration
2042-03-08

AI Technical Summary

Technical Problem

In the prior art, face key point detection cannot be realized in real-time and accurate detection in video, and is poorly robust and inefficient.

Method used

By obtaining the key image information of the current image frame and the historical key point information of the preorder image frame, it is input into the pretrained continuous frame key point detection model for detection, and the self-attention mechanism is used to improve detection accuracy and robustness.

Benefits of technology

It improves the accuracy and robustness of face key points detection, and realizes the stability of face recognition and tracking in videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114612976B_ABST
    Figure CN114612976B_ABST
Patent Text Reader

Abstract

The present disclosure provides a key point detection method and device, a computer-readable medium, and an electronic device, relating to the field of artificial intelligence technology. The method includes: obtaining a current image frame and determining key image information corresponding to the current image frame; determining a previous image frame corresponding to the current image frame and obtaining historical key point information corresponding to the previous image frame; inputting the current image frame, key image information, and historical key point information into a pre-trained continuous frame key point detection model, and outputting the current key point information corresponding to the current image frame. The present disclosure can better utilize the spatial and temporal information of video image frames to perform more stable facial key point detection, effectively improving the accuracy and robustness of facial key point detection results in videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a key point detection method, a key point detection device, a computer-readable medium, and an electronic device. Background Art

[0002] With the rapid development of science and technology, face recognition and face tracking technologies have attracted more and more attention. Tracking a face target in a video image mainly relies on determining the position of the key points of the face in the video image.

[0003] Currently, related facial key point detection schemes either cannot accurately detect facial key points in videos in real time, resulting in poor robustness of the detection results, or require post-processing of the detected facial key points, which cannot guarantee real-time performance, is labor-intensive, and has low output efficiency. Summary of the Invention

[0004] The purpose of the present disclosure is to provide a key point detection method, a key point detection device, a computer-readable medium and an electronic device, thereby improving the detection efficiency of facial key points in videos, at least to a certain extent, and improving the accuracy and robustness of the output facial key points.

[0005] According to a first aspect of the present disclosure, a key point detection method is provided, comprising:

[0006] Acquire a current image frame, and determine key image information corresponding to the current image frame;

[0007] Determining a previous image frame corresponding to the current image frame, and obtaining historical key point information corresponding to the previous image frame;

[0008] The current image frame, the key image information, the previous image frame and the historical key point information are input into a pre-trained continuous frame key point detection model, and the current key point information corresponding to the current image frame is output.

[0009] According to a second aspect of the present disclosure, there is provided a key point detection device, comprising:

[0010] A key image information determination module is used to obtain a current image frame and determine key image information corresponding to the current image frame;

[0011] A historical key point information acquisition module, configured to determine a previous image frame corresponding to the current image frame, and acquire historical key point information corresponding to the previous image frame;

[0012] The current key point information output module is used to input the current image frame, the key image information, the previous image frame and the historical key point information into the pre-trained continuous frame key point detection model, and output the current key point information corresponding to the current image frame.

[0013] According to a third aspect of the present disclosure, a computer-readable medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the above method is implemented.

[0014] According to a fourth aspect of the present disclosure, there is provided an electronic device, comprising:

[0015] processor; and

[0016] The memory is used to store one or more programs. When the one or more programs are executed by one or more processors, the one or more processors implement the above method.

[0017] The key point detection method provided by an embodiment of the present disclosure can determine the key image information corresponding to the current image frame, then determine the previous image frame corresponding to the current image frame, and obtain the historical key point information corresponding to the previous image frame. Finally, the current image frame, key image information, previous image frame and historical key point information are input into a pre-trained continuous frame key point detection model, and the current key point information corresponding to the current image frame is output. On the one hand, the input of the model includes the current image frame and key image information in the spatial dimension, which can effectively improve the detection accuracy of facial key points and improve the effect of face recognition and face tracking in the video. On the other hand, the input of the model includes the historical key point information of the previous image frame in the temporal dimension, which can effectively improve the robustness of the detection results and achieve more stable key point prediction for the video.

[0018] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort. In the drawings:

[0020] Figure 1 A schematic diagram showing an exemplary system architecture to which embodiments of the present disclosure may be applied;

[0021] Figure 2A schematic diagram schematically illustrates a flow chart of a key point detection method in an exemplary embodiment of the present disclosure;

[0022] Figure 3 A schematic diagram of a process for determining a key image block in an exemplary embodiment of the present disclosure is schematically shown;

[0023] Figure 4 A schematic diagram of a process for performing block processing on a current image frame in an exemplary embodiment of the present disclosure is shown schematically;

[0024] Figure 5 A schematic diagram of a process for determining a sequence of continuous image frames in an exemplary embodiment of the present disclosure is schematically shown;

[0025] Figure 6 A schematic diagram schematically illustrates a principle of realizing key point detection in a current image frame in an exemplary embodiment of the present disclosure;

[0026] Figure 7 A schematic diagram of a process for detecting key points on a current image frame in an exemplary embodiment of the present disclosure is shown schematically;

[0027] Figure 8 A schematic diagram schematically illustrates the composition of a key point detection device in an exemplary embodiment of the present disclosure;

[0028] Figure 9 A schematic diagram of an electronic device to which the embodiments of the present disclosure can be applied is shown. DETAILED DESCRIPTION

[0029] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0030] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0031] Figure 1A schematic diagram shows a system architecture of an exemplary application environment in which a key point detection method and apparatus according to an embodiment of the present disclosure can be applied.

[0032] like Figure 1 As shown, the system architecture 100 may include one or more of terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc. The terminal devices 101, 102, 103 may be various electronic devices with image processing functions, including but not limited to desktop computers, portable computers, smart phones, and tablet computers, etc. It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as needed. For example, the server 105 may be a server cluster consisting of multiple servers.

[0033] The key point detection method provided in the embodiments of the present disclosure is generally executed by the terminal devices 101, 102, and 103, and accordingly, the key point detection device is generally provided in the terminal devices 101, 102, and 103. However, it is readily understood by those skilled in the art that the key point detection method provided in the embodiments of the present disclosure may also be executed by the server 105, and accordingly, the key point detection device may also be provided in the server 105, and this is not particularly limited in this exemplary embodiment.

[0034] For example, in an exemplary embodiment, a user may upload the captured video data to a server 105 through terminal devices 101, 102, and 103. After the server generates key point information through the key point detection method provided by the embodiment of the present disclosure, the server transmits the key point information to the terminal devices 101, 102, 103, etc.

[0035] One technical solution proposes locating facial keypoints from the current video image frame and determining whether the face in the current video image frame has moved relative to the face in the previous video image frame. If there has been movement, the facial keypoints located in the current video image frame are determined as valid facial keypoints for the current video image frame. If there has been no movement, the weighted sum of the facial keypoints located in the current video image frame and the valid facial keypoints at the corresponding positions in the previous image frame is determined as the valid facial keypoints for the current video image frame. However, in this solution, keypoint smoothing by determining whether there has been facial movement requires manual setting of strategies and thresholds. These strategies and thresholds may vary depending on the video scene, resulting in a high workload and low detection efficiency. Furthermore, the need to set strategies and thresholds may result in poor robustness of the detection results.

[0036] Another technical solution proposes calculating the original facial keypoints from a facial image; smoothing these original facial keypoints and replacing their original coordinates to obtain preliminary stable facial keypoints; tracking the movement direction and distance of the facial keypoints in the preceding and following frames, and performing outlier detection and correction on these preliminary stable facial keypoints. However, due to the diversity of facial features, this solution requires tailored strategies for each face, resulting in low efficiency. Furthermore, since this is a post-processing method, it cannot achieve real-time detection and output of facial keypoints in the video.

[0037] Based on one or more problems that may exist in the related art, the present disclosure first provides a key point detection method. The key point detection method of an exemplary embodiment of the present disclosure is described in detail below.

[0038] Figure 2 A schematic flow chart of a key point detection method in this exemplary embodiment is shown, which may include at least steps S210 to S230:

[0039] In step S210 , a current image frame is acquired, and key image information corresponding to the current image frame is determined.

[0040] In one exemplary embodiment, the current image frame refers to the image frame in the video that requires facial landmark recognition at the current moment. For example, assuming a video includes image frame A at time t-1, image frame B at time t, and image frame C at time t+1, then at time t-1, image frame A is the current image frame. Similarly, at time t or time t+1, image frame B or image frame C is the current image frame.

[0041] Key image information refers to image information related to key points of the face in the current image frame. For example, the key image information can be the facial image block corresponding to the face image in the current image frame, or it can be the distribution of pixel brightness values ​​(or grayscale values) in the current image frame. Of course, it can also be other related image information that can characterize the key points in the current image frame. This example embodiment does not make any special limitations on this.

[0042] In step S220 , the previous image frame corresponding to the current image frame is determined, and historical key point information corresponding to the previous image frame is obtained.

[0043] In an exemplary embodiment, a preceding image frame refers to an image frame that precedes the current image frame in the time dimension. For example, the preceding image frame may be an image frame at the previous moment relative to the current moment, or may be multiple or all image frames before the current moment. The specific range of preceding image frames may be customized according to actual conditions. For example, when the computing power of the computing device is greater than the computing power threshold, multiple or all image frames before the current moment may be used as the preceding image frames corresponding to the current image frame. When the computing power of the computing device is less than or equal to the computing power threshold, an image frame before the current moment may be used as the preceding image frame corresponding to the current image frame. This exemplary embodiment does not impose any special limitation on the number of preceding image frames.

[0044] For example, assuming that for a video, including image frame A at time t-1, image frame B at time t, and image frame C at time t+1, if the current time is time t+1, then the current image frame is image frame C. At this time, image frame A at time t-1 or image frame B at time t can be used as the preceding image frame of image frame C. Of course, the preceding image frame of image frame C can also include all image frames before time t+1, such as image frame A at time t-1 and image frame B at time t.

[0045] It is easy for those skilled in the art to understand that if the current image frame is the first frame of the video, then the preceding image frame of the current image frame can be an empty frame, that is, the information related to the preceding image frame can be set to 0; of course, a default image frame can also be made according to the content of the video, and the default image frame can be used as the preceding image frame of the current image frame, which is also within the protection scope of this embodiment.

[0046] Historical key point information refers to the detection result data obtained by performing key point detection on the previous image frame. The historical key point information may include the coordinates of the facial key points in the previous image frame, and may also include the facial posture, facial occlusion area, lighting, facial expression and other image information related to the key points in the previous image frame. This example embodiment does not make any special limitations on this.

[0047] In step S230, the current image frame, the key image information and the historical key point information are input into a pre-trained continuous frame key point detection model, and the current key point information corresponding to the current image frame is output.

[0048] In an exemplary embodiment, the continuous frame key point detection model refers to a deep learning model used to perform key point detection on the input current image frame. For example, the continuous frame key point detection model can be a Transformer model. Of course, it can also be other models based on the self-attention mechanism. This exemplary embodiment does not make any special limitations on this.

[0049] The current key point information refers to the detection result data obtained by performing key point detection on the current image frame through the continuous frame key point detection model. The current key point information may include the coordinates of the facial key points in the current image frame, and may also include the facial posture, facial occlusion area, lighting, facial expression and other image information related to the key points in the current image frame. This example embodiment does not make any special limitations on this.

[0050] It is easy to understand that after obtaining the current key point information corresponding to the current image frame, the current key point information can be stored in the cache space, and when performing key point prediction of the image frame at the next moment, the current image frame can be used as the previous image frame of the image frame at the next moment, and the current key point information can be used as the model input information of the image frame at the next moment to continue predicting the key points of the face.

[0051] By using the key image information of the current image frame and the historical key point information of the previous image frame as auxiliary information for key point detection of the current image frame, and combining it with the continuous frame key point detection model based on the self-attention mechanism, we can better utilize the feature information of continuous video frames in the time and spatial dimensions, effectively improve the accuracy of the face key points of the current image frame, and stably identify and output the face key points of the face area in the continuous frames of the video, thereby improving the robustness of the detection results.

[0052] Steps S210 to S230 are described in detail below.

[0053] In an exemplary embodiment, the key image information may include a key image block, which refers to an image area in the current image frame that contains key content. For example, the key image block may be an image block corresponding to facial features in the face area, such as image blocks at positions such as the eyes, nose, and mouth in the face area, or an image area with a large gradient change in the key image block. Of course, the key image block may also be an image area determined based on actual conditions and where image key points may exist in the current image frame. This exemplary embodiment does not specifically limit this.

[0054] Optionally, you can pass Figure 3 The steps in the above are used to determine the key image information corresponding to the current image frame. Figure 3 Specifically, it may include:

[0055] Step S310, detecting the current image frame using a feature point detection model to determine target feature points;

[0056] Step S320 : performing block processing on the current image frame based on the target feature points to obtain key image blocks corresponding to the current image frame.

[0057] Among them, the feature point detection model refers to a model used to detect all feature points in the current image frame. For example, the feature point detection model can be a simple five-point facial key point detection model, which is used to detect the facial feature points in the face area. Of course, the feature point detection model can also be a feature detection model based on a feature description operator, such as a scale invariant feature transform (SIFT) operator, a histogram of oriented gradients (HOG) operator, a Harris corner operator, etc. This example embodiment does not make any special restrictions on this.

[0058] The target feature points can be all the feature points identified by the feature point detection model in the current image frame, such as the target feature points can be the five facial feature points detected in the current image frame by the five-point facial feature point detection model, or they can be some of the feature points identified by the feature point detection model in the current image frame, such as the target feature points can be the five facial feature key points obtained by screening after detecting multiple facial key points in the current image frame by the feature point detection model. This example embodiment does not make any special limitations on this.

[0059] Block processing refers to the process of segmenting the current image frame according to the target feature points. For example, block processing can be a process of treating the image content within the target distance around the target feature point as one image block, or it can be a process of dividing the current image frame into multiple irregular image blocks according to the distribution position of the target feature point in the current image frame. This example embodiment does not make any special restrictions on this.

[0060] The target feature points are determined through the feature point detection model, and the current image frame is divided into blocks according to the target feature points to obtain key image blocks of different sizes in the current image frame. The current image frame and key image blocks of different sizes are used as the model input of the continuous frame key point detection model. This can effectively improve the continuous frame key point detection model's ability to recognize key points in the current image frame and improve the detection accuracy of facial key points in continuous frames.

[0061] In an exemplary embodiment, before determining the key image information corresponding to the current image frame or detecting the current image frame through a feature point detection model to determine the target feature points, face detection can be performed on the current image frame to determine the face area in the current image frame.

[0062] Optionally, face detection can be performed on the current image frame using a face detection model YOLO (You Only Look Once, an object detection algorithm based on a deep convolutional neural network), or by using a multi-task convolutional neural network model (MTCNN). This example embodiment does not impose any special limitations on this.

[0063] By performing face detection on the current image frame, determining the face area in the current image frame, and then only determining key image information or target feature points in the face area, the system's computational load can be effectively reduced and the system's detection speed can be improved.

[0064] Optionally, a face area threshold may be pre-set to delete face areas smaller than the face area threshold, thereby avoiding failure in facial key point detection due to a face area that is too small.

[0065] In an exemplary embodiment, the Figure 4 The steps in the block processing of the current image frame are implemented, refer to Figure 4 Specifically, it may include:

[0066] Step S410, obtaining preset bounding box parameters;

[0067] Step S420 , constructing a face bounding box according to the coordinates corresponding to the target feature points and the bounding box parameters, so as to perform block processing on the current image frame through the face bounding box.

[0068] Among them, the bounding box parameters refer to the relevant parameters used to determine the size of the bounding box. For example, the bounding box parameters can be the shape of the bounding box, such as a square, a circle, etc., or the area or size of the bounding box. This example embodiment does not make any special restrictions on this.

[0069] The coordinates corresponding to the target feature point can be obtained, and the coordinates corresponding to the target feature point can be used as the center to construct a face bounding box based on the bounding box parameters. For example, assuming that the coordinates of the target feature point in the current image frame are (2, 2), the bounding box parameters can include that the bounding box shape is a square with a side length of 2 units. Then, the center coordinates of the constructed face bounding box are (2, 2), and the coordinates of the four vertices are (1, 1), (1, 3), (3, 3), and (3, 1). Of course, this is only a schematic example. According to actual needs, the coordinates corresponding to the target feature point can be used as a vertex of the face bounding box. It is only necessary to include the target feature point in the face bounding box. This example embodiment is not limited to this.

[0070] Optionally, the current image frame can be detected by a feature point detection model that can detect more feature points, and the bounding box can be inferred by the detected multiple feature points. For example, the current image frame can be detected by a 63-point facial feature point detection model to obtain 63 facial feature points, and then the facial bounding box that can enclose the facial features area can be inferred by the 63 facial feature points. Compared with the method of constructing a facial bounding box according to the coordinates corresponding to the target feature points and the bounding box parameters, the determined facial bounding box can be made more precise and a more accurate key image block can be obtained.

[0071] By constructing a face bounding box that meets the requirements and dividing the current image frame into blocks according to the face bounding box, the accuracy of the key image blocks obtained after blocking is effectively guaranteed, and the accuracy of facial key point detection is further guaranteed.

[0072] In one exemplary embodiment, the key image information may include estimated key points, which are estimated values ​​of key points in the current image frame. Specifically, target feature points detected by a feature point detection model on the current image frame may be used as the estimated key points corresponding to the current image frame.

[0073] By taking the estimated key points as the key image information of the current image frame and as the input of the continuous frame key point detection model, the continuous frame key point detection model can be further restricted from detecting key points in the current image frame, ensuring that the output results of the continuous frame key point detection model will not be affected by abnormal key points, thereby improving the robustness of the continuous frame key point detection model.

[0074] In an exemplary embodiment, the current image frame and the previous image frame belong to the same continuous image frame sequence. Specifically, Figure 5 The steps in the above are used to determine the sequence of continuous image frames. Figure 5 Specifically, it may include:

[0075] Step S510, obtaining a target video;

[0076] Step S520 , extracting frames from the target video according to a preset frame extraction interval to obtain the continuous image frame sequence.

[0077] The target video refers to the video data for which key point detection is to be performed. The frame extraction interval refers to a preset frame extraction frequency. For example, the frame extraction interval can be 5 times per second or 20 times per second. The specific setting can be based on the frame rate of the target video and is not specifically limited in this exemplary embodiment.

[0078] After obtaining the current key point information of the current image frame, the key point can be rendered and displayed in the current image frame according to the key point coordinates in the current key point information, thereby realizing continuous facial key point detection of the target video.

[0079] Optionally, since the image frames for facial key point detection are all image frames in a continuous image frame sequence obtained by frame extraction, the key points of the current image frame can be rendered as the key points of the graphic frames after the current image frame in the target video and before the next frame in the continuous image frame sequence, thereby realizing real-time rendering and display of the key points of all image frames of the target video.

[0080] Figure 6 A schematic diagram schematically illustrates a principle for implementing key point detection in a current image frame in an exemplary embodiment of the present disclosure.

[0081] refer to Figure 6 As shown, assuming that for the current image frame T at time T, the previous image frames of the current image frame T may include the previous image frame T-1 at time T-1, the previous image frame T-2 at time T-2, and the previous image frame T-3 at time T-3.

[0082] First, face detection can be performed on the current image frame T to obtain a face region 601. Face region 601 is then detected using a feature point detection model. For example, a five-point face key point detection model is used to detect face region 601 and determine key points at locations such as the nose, eyes, and mouth. The current image frame T is then segmented using a constructed face bounding box to obtain key image blocks 602. Historical key point information corresponding to the previous image frame is then obtained, namely, historical key point information 603 corresponding to the previous image frame T-1, historical key point information 604 corresponding to the previous image frame T-2, and historical key point information 605 corresponding to the previous image frame T-3. This historical key point information may include the previous image frame, key image blocks corresponding to the previous image frame, key point coordinates corresponding to the previous image frame, facial pose output by the model, occlusion areas, lighting, facial expression, and other information. Other types of key point information may also be used, and this exemplary embodiment is not limited thereto.

[0083] Then, the previous image frame, the historical key point information corresponding to the previous image frame, the current image frame T, the key image block 602, the estimated key points, and other information can be input into the continuous frame key point detection model 606 to obtain the current key point information 607 corresponding to the current image frame T. The current key point information 607 can then be stored and used as the historical key point information for subsequent image frames.

[0084] Figure 7 A schematic diagram of a process of performing key point detection on a current image frame in an exemplary embodiment of the present disclosure is schematically shown.

[0085] refer to Figure 7 As shown, in step S710, a current image frame is obtained from a continuous image frame sequence corresponding to a target video;

[0086] Step S720, determining key image information corresponding to the current image frame;

[0087] Step S730, determining whether the current image frame is the first frame of the target video, if so, executing step S740, otherwise executing step S750;

[0088] Step S740, setting the previous image frame to 0, or using a default image frame as the previous image frame;

[0089] Step S750, determining a previous image frame and obtaining historical key point information corresponding to the previous image frame;

[0090] Step S760: input the current image frame, key image information, previous image frame and historical key point information into a pre-trained continuous frame key point detection model;

[0091] Step S770: Generate current key point information, render and display the current key point information to the target video, and end the current process.

[0092] In summary, in this exemplary embodiment, the key image information corresponding to the current image frame can be determined, and then the previous image frame corresponding to the current image frame can be determined, and the historical key point information corresponding to the previous image frame can be obtained. Finally, the current image frame, key image information, previous image frame, and historical key point information are input into a pre-trained continuous frame key point detection model, and the current key point information corresponding to the current image frame is output. On the one hand, the input of the model includes the current image frame and key image information in the spatial dimension, which can effectively improve the detection accuracy of facial key points and improve the effectiveness of face recognition and face tracking in videos. On the other hand, the input of the model includes the historical key point information of the previous image frame in the temporal dimension, which can effectively improve the robustness of the detection results and achieve more stable key point prediction for the video.

[0093] It should be noted that the above figures are merely illustrative of the processes included in the methods according to exemplary embodiments of the present disclosure and are not intended to be limiting. It is readily understood that the processes illustrated in the above figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0094] For further reference, Figure 8 As shown, in the embodiment of this example, a key point detection device 800 is further provided, which may include a key image information determination module 810, a historical key point information acquisition module 820, and a current key point information output module 830.

[0095] The key image information determination module 810 is used to obtain a current image frame and determine the key image information corresponding to the current image frame;

[0096] The historical key point information acquisition module 820 is used to determine the previous image frame corresponding to the current image frame, and obtain the historical key point information corresponding to the previous image frame;

[0097] The current key point information output module 830 is used to input the current image frame, the key image information and the historical key point information into a pre-trained continuous frame key point detection model, and output the current key point information corresponding to the current image frame.

[0098] In an exemplary embodiment, the key image information may include key image blocks, and the key image information determination module 810 may be configured to:

[0099] Detecting the current image frame using a feature point detection model to determine target feature points;

[0100] The current image frame is divided into blocks based on the target feature points to obtain key image blocks corresponding to the current image frame.

[0101] In an exemplary embodiment, the key image information determination module 810 may be configured to:

[0102] Get the preset bounding box parameters;

[0103] A face bounding box is constructed according to the coordinates corresponding to the target feature points and the bounding box parameters, so as to perform block processing on the current image frame through the face bounding box.

[0104] In an exemplary embodiment, the key image information may include estimated key points, and the key point detection apparatus 800 may further include an estimated key point determination unit, which may be configured to:

[0105] The estimated key points corresponding to the current image frame are determined using the target feature points.

[0106] In an exemplary embodiment, the key point detection apparatus 800 may further include a face region detection unit, which may be configured to:

[0107] Perform face detection on the current image frame to determine the face area.

[0108] In an exemplary embodiment, the key point detection apparatus 800 may further include a frame extraction unit, which may be configured to:

[0109] Get the target video;

[0110] The target video is frame-sampled according to a preset frame-sample interval to obtain the continuous image frame sequence.

[0111] In an exemplary embodiment, the continuous frame key point detection model may be a model based on a self-attention mechanism, and the model based on the self-attention mechanism may include a Transformer model.

[0112] The specific details of each module in the above device have been described in detail in the implementation method part. The undisclosed details can be found in the implementation method part, so they will not be repeated here.

[0113] Those skilled in the art will appreciate that various aspects of the present disclosure may be implemented as systems, methods, or program products. Therefore, various aspects of the present disclosure may be implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be collectively referred to herein as "circuits," "modules," or "systems."

[0114] An exemplary embodiment of the present disclosure provides an electronic device for implementing a key point detection method, which may be Figure 1 The terminal device 101, 102, 103 or the server 105 in the embodiment of the present invention comprises at least a processor and a memory, wherein the memory is used to store executable instructions of the processor, and the processor is configured to perform the key point detection method by executing the executable instructions.

[0115] Below is Figure 9 Taking the electronic device 900 in FIG. 1 as an example, the structure of the electronic device in the present disclosure is exemplarily described. Figure 9 The electronic device 900 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0116] like Figure 9As shown, electronic device 900 is implemented as a general-purpose computing device. Components of electronic device 900 may include, but are not limited to, at least one processing unit 910, at least one storage unit 920, a bus 930 connecting various system components (including storage unit 920 and processing unit 910), and a display unit 940.

[0117] The storage unit 920 stores program codes, and the program codes can be executed by the processing unit 910, so that the processing unit 910 performs the key point detection method in this specification.

[0118] The storage unit 920 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 921 and / or a cache memory unit 922 , and may further include a read-only memory unit (ROM) 923 .

[0119] The storage unit 920 may also include a program / utility 924 having a set (at least one) of program modules 925, such program modules 925 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0120] Bus 930 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0121] The electronic device 900 can also communicate with one or more external devices 970 (e.g., sensor devices, Bluetooth devices, etc.), one or more devices that enable a user to interact with the electronic device 900, and / or any device that enables the electronic device 900 to communicate with one or more other computing devices (e.g., routers, modems, etc.). This communication can occur via an input / output (I / O) interface 950. Furthermore, the electronic device 900 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 960. As shown, the network adapter 960 communicates with other modules of the electronic device 900 via a bus 930. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 900, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, data backup storage systems, and sensor modules (e.g., gyroscope sensors, magnetic sensors, accelerometers, distance sensors, proximity sensors, etc.).

[0122] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes a number of instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0123] The exemplary embodiments of the present disclosure further provide a computer-readable storage medium on which a program product capable of implementing the above-mentioned method of the present specification is stored. In some possible implementations, various aspects of the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section of the present disclosure, for example, Figures 2 to 7 Any one or more steps in .

[0124] It should be noted that the computer-readable medium shown in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0125] In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the foregoing.

[0126] In addition, the program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and the like, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0127] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow from the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the claims.

[0128] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A key point detection method, characterized in that: include: Acquire a current image frame, and determine key image information corresponding to the current image frame; Determining a previous image frame corresponding to the current image frame, and obtaining historical key point information corresponding to the previous image frame; Inputting the current image frame, the key image information, the previous image frame, and the historical key point information into a pre-trained continuous frame key point detection model, and outputting the current key point information corresponding to the current image frame; The key image information includes key image blocks and estimated key points, and determining the key image information corresponding to the current image frame includes: Detecting the current image frame using a feature point detection model to determine target feature points; Performing block processing on the current image frame based on the target feature points to obtain key image blocks corresponding to the current image frame; The method further comprises: The estimated key points corresponding to the current image frame are determined using the target feature points.

2. The method according to claim 1, characterized in that The performing block processing on the current image frame based on the target feature points includes: Get the preset bounding box parameters; A face bounding box is constructed according to the coordinates corresponding to the target feature points and the bounding box parameters, so as to perform block processing on the current image frame through the face bounding box.

3. The method according to any one of claims 1 to 2, characterized in that Before determining the key image information corresponding to the current image frame, the method further includes: Perform face detection on the current image frame to determine the face area.

4. The method according to claim 1, wherein The current image frame and the previous image frame belong to the same continuous image frame sequence, and the method further includes: Get the target video; The target video is frame-sampled according to a preset frame-sample interval to obtain the continuous image frame sequence.

5. The method according to claim 1, wherein The continuous frame key point detection model is a model based on the self-attention mechanism, and the model based on the self-attention mechanism includes a Transformer model.

6. A key point detection device, characterized in that: include: A key image information determination module is used to obtain a current image frame and determine key image information corresponding to the current image frame; The key image information includes key image blocks and estimated key points; A historical key point information acquisition module, configured to determine a previous image frame corresponding to the current image frame, and acquire historical key point information corresponding to the previous image frame; A current key point information output module is used to input the current image frame, the key image information, the previous image frame and the historical key point information into a pre-trained continuous frame key point detection model, and output the current key point information corresponding to the current image frame; The determining of the key image information corresponding to the current image frame includes: detecting the current image frame using a feature point detection model to determine target feature points; and performing block processing on the current image frame based on the target feature points to obtain key image blocks corresponding to the current image frame; The device is further configured to determine estimated key points corresponding to the current image frame using the target feature points.

7. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

8. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to perform the method according to any one of claims 1 to 5 by executing the executable instructions.

Citation Information

Patent Citations

  • method and device for generating information

    CN109829432A