Data processing method and device, electronic equipment and storage medium

By combining key point recognition models and tracking predictions from historical frame images in video data, and comprehensively analyzing the first and second position information, the problem of low positioning accuracy in single-frame images is solved, achieving high-accuracy positioning and stability of key points, and improving the video data processing effect.

CN114627519BActive Publication Date: 2025-12-05ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011469640.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-14
Publication Date
2025-12-05
Estimated Expiration
2040-12-14

AI Technical Summary

Technical Problem

In existing technologies, when key point localization models analyze single-frame images, the accuracy of key point localization is low, resulting in poor effects in facial beautification and makeup processing.

Method used

By acquiring frame images from video data, a key point recognition model is used to determine the first location information of the target object. Combined with tracking and prediction of historical frame images, the second location information of the target object is obtained. The first and second location information are analyzed together to determine the attribute information of the key points, thereby improving the positioning accuracy.

Benefits of technology

It improves the accuracy and stability of key point positioning in video data, and enhances the effects of facial beautification, makeup, and special effects processing, especially in live videos and e-commerce live streams, improving the continuity and accuracy of beautification, makeup, and special effects processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114627519B_ABST
    Figure CN114627519B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a kind of data processing method, device, electronic equipment and storage medium, the described method comprises: obtaining video data, and determining frame image in video data;Frame image is identified, and the detection image corresponding to detection frame is extracted, and the detection image contains the target object positioned by detection frame;Determine the first position information and the second position information of the key point of target object in target frame image, the first position information is determined according to the target detection image of target frame image, and the second position information is determined by tracking and predicting the position of the key point of target object in historical frame image, and the historical frame image includes at least one frame of frame image before the target frame image;According to the first position information and the second position information, determine the attribute information of the key point of target object in target frame image;It can improve the accuracy of key point positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a data processing method, a data processing device, an electronic device, and a storage medium. Background Technology

[0002] With the increasing popularity of applications such as live video streaming and video shopping, beautification, makeup, and special effects are important components in many video applications. Most of the time, beautification, makeup, special effects, and enhancement of facial features require first accurately locating the key points of the face, and then using the location of the key points to process the face.

[0003] One existing key point localization model is to perform image recognition on a single frame of a face image, determine the position information of each key point on the face in the single frame image, and use the position information of the key points to perform beautification and makeup processing on the face.

[0004] However, the method of using a key point localization model to analyze a single frame image to locate key points has low accuracy. Summary of the Invention

[0005] This application provides a data processing method to improve the accuracy of key point positioning.

[0006] Accordingly, embodiments of this application also provide a data processing device, an electronic device, and a storage medium to ensure the implementation and application of the above system.

[0007] To address the aforementioned problems, this application discloses a data processing method, comprising: acquiring video data and determining frame images in the video data; recognizing the frame images and extracting detection images corresponding to detection boxes, wherein the detection images contain target objects located by the detection boxes; determining first position information and second position information of key points of the target objects in the target frame images, wherein the first position information is determined based on the target detection image of the target frame images, and the second position information is determined by tracking and predicting the positions of key points of the target objects in historical frame images, wherein the historical frame images include frame images of at least one frame preceding the target frame images; and determining attribute information of key points of the target objects in the target frame images based on the first position information and the second position information.

[0008] To address the aforementioned issues, this application discloses a data processing method, comprising: training a keypoint recognition model using a training input image and training input annotations on a single-frame image; inputting the training image from the training video data into the keypoint recognition model that has completed single-frame image training to obtain training annotations representing the positions of keypoints of the target object in the training image; tracking and predicting historical training images of the target training image to determine the predicted training annotations of the training image at the target time; verifying the reliability of the training annotations of the target training image based on the difference between the predicted training annotations and the training annotations of the training image at the target time; and training the keypoint recognition model that has completed single-frame image training using the verified training annotations and the corresponding training images.

[0009] To address the aforementioned problems, this application discloses a data processing method, comprising: acquiring video data and determining frame images in the video data; recognizing the frame images and extracting detection images corresponding to detection boxes, wherein the detection images contain target objects located by the detection boxes; determining first and second position information of key points of the target objects in the target frame images, wherein the first position information is determined based on the target detection image of the target frame images, and the second position information is determined by tracking and predicting the positions of key points of the target objects in historical frame images, wherein the historical frame images include frame images of at least one frame preceding the target frame images; determining attribute information of key points of the target objects in the target frame images based on the first and second position information; and performing special effects processing on the frame images in the video data based on the attribute information of the key points of the target objects.

[0010] To address the aforementioned issues, this application discloses a data processing method, comprising: acquiring video data and determining frame images within the video data; identifying the frame images and extracting detection images corresponding to detection boxes, the detection images containing facial data located by the detection boxes; determining first and second position information of key points of the facial data in a target frame image, the first position information being determined based on a target detection image of the target frame image, the second position information being determined by tracking and predicting the positions of key points of the facial data in historical frame images, the historical frame images including at least one frame image preceding the target frame image; determining attribute information of key points of the facial data in the target frame image based on the first and second position information; and performing special effects processing on the frame images in the video data based on the attribute information of the key points of the facial data, the special effects processing including at least one of beautification processing, special effects addition processing, makeup processing, and beauty enhancement processing.

[0011] To address the aforementioned issues, this application discloses a data processing method, comprising: acquiring live video data and identifying frame images in the live video data; extracting detection images corresponding to detection boxes, wherein the detection images contain facial data located by the detection boxes; determining first and second position information of key points of the facial data in a target frame image, wherein the first position information is determined based on a target detection image of the target frame image, and the second position information is determined by tracking and predicting the positions of key points of the facial data in historical frame images, wherein the historical frame images include at least one frame image preceding the target frame image; determining attribute information of key points of the facial data in the target frame image based on the first and second position information; and performing special effects processing on the frame images in the live video data based on the attribute information of the key points of the facial data, wherein the special effects processing includes at least one of beautification processing, special effects addition processing, makeup processing, and beauty enhancement processing.

[0012] To address the aforementioned issues, this application discloses a data processing method, comprising: providing a first interface to acquire relevant video data and determine frame images in the video data; recognizing the frame images and extracting detection images corresponding to detection boxes, wherein the detection images contain target objects located by the detection boxes; determining first and second position information of key points of the target objects in the target frame image, wherein the first position information is determined based on the target detection image of the target frame image, and the second position information is determined by tracking and predicting the positions of key points of the target objects in historical frame images, wherein the historical frame images include frame images of at least one frame preceding the target frame image; determining attribute information of key points of the target objects in the target frame image based on the first and second position information; and feeding back the attribute information through a second interface.

[0013] To address the aforementioned issues, this application discloses a data processing apparatus, comprising: a video data acquisition module for acquiring video data and determining frame images within the video data; a detection image acquisition module for recognizing the frame images and extracting detection images corresponding to detection boxes, wherein the detection images contain target objects located by the detection boxes; a location information acquisition module for determining first location information and second location information of key points of the target object in a target frame image, wherein the first location information is determined based on a target detection image of the target frame image, and the second location information is determined by tracking and predicting the positions of key points of the target object in historical frame images, wherein the historical frame images include frame images of at least one frame preceding the target frame image; and an attribute information acquisition module for determining attribute information of key points of the target object in the target frame image based on the first location information and the second location information.

[0014] To address the aforementioned issues, this application discloses a data processing apparatus, comprising: a single-frame training module for training a keypoint recognition model using training input images and training input annotations; a training annotation acquisition module for inputting training images from training video data into the single-frame image-trained keypoint recognition model to obtain training annotations representing the positions of keypoints of a target object in the training images; a prediction annotation acquisition module for tracking and predicting historical training images of the target training image to determine the predicted training annotations for the training image at the target time; a training annotation verification module for verifying the reliability of the training annotations of the target training image based on the difference between the predicted training annotations and the training annotations of the training image at the target time; and a secondary training module for training the single-frame image-trained keypoint recognition model using the verified training annotations and corresponding training images.

[0015] To address the aforementioned issues, this application discloses an electronic device comprising: a processor; and a memory storing executable code thereon, wherein when the executable code is executed, the processor performs one or more of the methods described in the above method embodiments.

[0016] To address the aforementioned issues, embodiments of this application disclose one or more machine-readable media storing executable code thereon, which, when executed, causes a processor to perform one or more of the methods described in the above method embodiments.

[0017] Compared with the prior art, the embodiments of this application have the following advantages:

[0018] In this embodiment, frame images in video data can be identified, and detection images containing target objects can be extracted. Then, on the one hand, the target detection image of the target frame image can be input into a keypoint recognition model to determine the first position information of the keypoints of the target object in the target frame image; on the other hand, based on the continuity of the target object's movement in the video data, the position of the keypoints of the target object in historical frame images can be tracked and analyzed to obtain the second position information of the keypoints of the target object in the target frame image. Subsequently, based on the first and second position information, the attribute information of the keypoints of the target object in the target frame image is obtained. Compared with the method of determining the first position information of the keypoints of the target object using only a keypoint recognition model, this embodiment also considers the continuous movement of the keypoints of the target object in the video data to analyze the second position information, and determines the attribute information of the keypoints based on the first and second position information, which can improve the accuracy of keypoint localization of the target object. Attached Figure Description

[0019] Figure 1This is a schematic flowchart of a data processing method according to an embodiment of this application;

[0020] Figure 2 This is a schematic flowchart of a data processing method according to another embodiment of this application;

[0021] Figure 3 This is a schematic flowchart of a data processing method according to another embodiment of this application;

[0022] Figure 4 This is a schematic flowchart of a data processing method according to another embodiment of this application;

[0023] Figure 5 This is a schematic flowchart of a data processing method according to another embodiment of this application;

[0024] Figure 6 This is a schematic flowchart of a data processing method according to another embodiment of this application;

[0025] Figure 7 This is a schematic flowchart of a data processing method according to another embodiment of this application;

[0026] Figure 8 This is a schematic flowchart of a data processing method according to another embodiment of this application;

[0027] Figure 9 This is a schematic diagram of the structure of a data processing apparatus according to an embodiment of this application;

[0028] Figure 10 This is a schematic diagram of the structure of a data processing apparatus according to another embodiment of this application;

[0029] Figure 11 This is a schematic diagram of the structure of a data processing apparatus according to another embodiment of this application;

[0030] Figure 12 This is a schematic diagram of the structure of a data processing apparatus according to another embodiment of this application;

[0031] Figure 13 This is a schematic diagram of the structure of a data processing apparatus according to another embodiment of this application;

[0032] Figure 14 This is a schematic diagram of the structure of a data processing apparatus according to another embodiment of this application;

[0033] Figure 15 This is a schematic diagram of the structure of an apparatus provided in one embodiment of this application. Detailed Implementation

[0034] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0035] This application can be applied to the field of identifying key points of target objects in video data, in order to determine the location information of the key points of the target objects in the frame images of the video data, so as to perform image processing on the images in the video data based on the location information of the key points, such as beautifying, applying makeup, or adding special effects to facial images.

[0036] This application provides a data processing method. In this method, on one hand, a keypoint recognition model can be used to identify target frame images in video data to determine the first position information of the keypoints of the target object in the target frame image. On the other hand, based on the continuity of the target object's motion in the video data, historical frame images from the previous frame (or multiple frames) of the target frame image can be used for tracking analysis to predict the second position information of the keypoints of the target object in the target frame image. Then, the offset between the first and second position information is analyzed. When the offset between the first and second position information is large (exceeding an offset threshold), the first position information is used as the attribute information of the keypoints of the target object in the target frame image; when the offset between the first and second position information is small, the second position information is used as the attribute information of the keypoints of the target object in the target frame image. Afterwards, the image can be processed accordingly based on the attribute information of the keypoints of the target object, such as performing beautification or makeup on a face image.

[0037] In this embodiment, a keypoint recognition model can be used to determine the first position information of the keypoints of the target object in the target frame image. Based on the continuity of the target object's motion in the video data, the position of the target object in the historical frame images before the target frame image is tracked and analyzed to obtain the second position information of the keypoints of the target object in the target frame image. Then, based on the first and second position information, the attribute information of the keypoints of the target object in the target frame image is obtained. Compared with the method of only using a keypoint recognition model to determine the first position information of the keypoints of the target object, this embodiment also considers the continuous behavior of the keypoints of the target object in the video data to analyze the second position information, which can improve the accuracy of the keypoint positioning of the target object and improve the stability of the keypoints in the video data, providing a more stable keypoint position for subsequent image processing and improving the image processing effect.

[0038] Specifically, such as Figure 1As shown, after acquiring the video data, the frame images in the video data can be determined. Then, the first detection box and the second detection box of the target object in the target frame image can be determined respectively. Then, based on the first detection box and the second detection box, the target detection box of the target object in the target frame image can be determined so as to extract the detection image. Specifically, on the one hand, the object descriptor box used to locate the target object in the target frame image can be determined, and the image corresponding to the object descriptor box can be input into the object box recognition model. The object box recognition model can identify the key points of the target object in the image of the object descriptor box, and determine the angle of the target object based on the key points of the target object. Based on the angle, the target frame image is rotated and adjusted, and the first detection box of the target object in the target frame image is obtained based on the located key points. On the other hand, the position of the key points of the target object in the previous frame (or multiple frames) of the historical frame image can be tracked and predicted to predict the second detection box of the target object in the target frame image. After determining the first detection box and the second detection box, the deviation between the first detection box and the second detection box can be analyzed. Then, based on the deviation, the first detection box (when the deviation exceeds the deviation threshold) or the second detection box (when the deviation does not exceed the deviation threshold) can be determined as the target detection box of the target object, and the detection image corresponding to the target detection box can be extracted.

[0039] After extracting the detection image containing the target object from the target frame image, on the one hand, the target detection image of the target object can be input into a keypoint recognition model. The keypoint recognition model can identify the keypoints of the target object and obtain the first position information of the keypoints of the target object in the target frame image. On the other hand, based on the continuity of the target object's motion in the video data, the position of the target object in the previous historical frame images (one or more previous frames) can be tracked and analyzed to predict the second position information of the keypoints of the target object in the target frame image. Then, based on the offset between the first and second position information, the first position information (when the offset exceeds the offset threshold) or the second position information (when the offset does not exceed the offset threshold) is determined as the attribute information of the keypoints of the target object.

[0040] A common approach to video data recognition involves using a keypoint recognition model to identify keypoints in a single frame, then processing the image based on these keypoints. However, this method fails to consider the movement of the target object (e.g., a face) within consecutive frames of video data. During movement (e.g., a face changing from a frontal view to a side view), keypoints may not be identified in some frames, or the identified keypoints may be inaccurate. This can easily lead to interruptions or incompatibility in subsequent frame processing (e.g., beauty effects shifting in some frames, causing background distortion), resulting in low image quality. In contrast, the method described in this application utilizes a keypoint recognition model to obtain first position information, tracks and analyzes historical frame images to obtain second position information, and then determines the attribute information of the target object's keypoints in the image based on the first and second position information. This attribute information is then used to process the frame image. In this embodiment, the attribute information of key points can be determined by comprehensive analysis based on the first and second position information. This results in more stable key points, and more accurate and continuous attribute information can be used to perform special effects processing on the frame image, thereby improving the image processing effect and the data quality of the video data.

[0041] The method described in this application can be applied to scenarios involving the identification of facial key points of a live streamer. It can locate the first position information of the key points based on a key point recognition model, predict the second position information of the key points by tracking and analyzing historical frames in the live stream, and then determine the attribute information of the key points based on the first and second position information. This allows for continuous and stable beautification and makeup processing for the streamer, improving the beautification and makeup effects on the streamer's facial data. Furthermore, it can provide the streamer with virtual 3D structural data of wearable items (such as glasses) and adapt the wearable items to the streamer according to the attribute information of the key points. In this embodiment, the attribute information of the key points can include not only the position of the key points but also the spatial angle information of the streamer's face. This allows for the rotation of the virtual 3D structural data of the wearable items based on the spatial angle information of the streamer's face, and then adapting the rotated virtual 3D structural data to the streamer's face, improving the adaptation effect between the virtual 3D structural data and the streamer's face.

[0042] The method described in this application can also be applied to the scenario of identifying key points of products in e-commerce live streaming. It can locate the first position information of the key point based on the key point recognition model, predict the second position information of the key point by tracking and analyzing historical frame images in the e-commerce live streaming video, and then determine the attribute information of the key point based on the first and second position information. In order to accurately locate the position of the product in the image based on the attribute information, and perform corresponding processing on the image corresponding to the product, such as enlarging the image corresponding to the product, and obtaining the three-dimensional structure data of the product. Based on the three-dimensional structure data of the product, the clarity of the product in the e-commerce live streaming video can be improved so that users watching the live stream can understand the details of the product.

[0043] The method described in this application can also be applied to scenarios involving the identification of key points of target objects (such as athletes) in sports videos. Based on a key point identification model, it can locate the first position information of key points (such as key points on an athlete's head, arms, or legs). Then, based on tracking and analysis of historical frames in the sports video, it can predict the second position information of the key points. Finally, it can determine the attribute information of the key points based on the first and second position information. This allows for more accurate location of the target object, enabling the application of special effects to the sports video (such as magnifying the target object), resulting in clearer and more accurate playback data and improving the data quality of the playback data.

[0044] The method described in this application can also be applied to scenarios involving vehicle identification in road surveillance videos. It can locate the first location information of key points based on a key point recognition model, predict the second location information of key points by tracking and analyzing historical frame images in the road surveillance video, and then determine the attribute information of key points based on the first and second location information. This allows for the extraction of clearer and more accurate relevant information (such as vehicle license plate, driver status, vehicle speed, etc.) based on this attribute information.

[0045] This application provides a data processing method that can be executed by a processing terminal. It can perform image recognition on video data to locate the position of a target object within a frame of the video data. Based on the location information of key points, image processing can be performed on the video data, such as beautifying or applying makeup to a face. The processing terminal can be understood as a device that acquires, stores, and forwards data, such as a device that acquires, stores, or forwards live video or surveillance video. Figure 2 As shown, the method includes:

[0046] Step 202: Acquire video data and identify frame images within the video data. The video data can be live sports videos, live streamer videos, traffic monitoring videos, community monitoring videos, etc. The video data contains frame images. In this embodiment, the corresponding frame images can be extracted from the video data. In step 204, the frame images are identified, and the detection images corresponding to the detection boxes are extracted. The detection images contain the target object located by the detection boxes. In this scheme, frame images can be identified to distinguish between the target object and the background. Then, the detection boxes of the target objects in the frame images can be further identified, and the corresponding detection images can be extracted for subsequent analysis. In this embodiment, the detection boxes of the target objects can be located first, and then key points in the images within the detection boxes can be located, which can reduce the size of the images analyzed later and improve the speed of data processing. In this embodiment, target object recognition technology can be used to determine the position of key points of the target object in the frame image to distinguish the target object from the background. Then, based on the position of the key points of the target object in the frame image, a detection box for the target object is determined, and the detection image corresponding to the detection box is extracted for identification, thereby obtaining the positional information of the key points of the target object. The target object recognition technology can identify one or more features of the target object to determine whether the target object exists in the frame image, and thus determine the detection box of the target object. For example, when the target object is a face, this embodiment can use face recognition technology to determine the position of facial features (such as mouth, eyes, and nose) in the frame image to distinguish the face from the background, determine the detection box of the face in the frame image, and extract the image corresponding to the detection box to locate the key points of the face in the detection image.

[0047] This embodiment can further adjust the identified detection boxes to obtain more accurate detection boxes, thereby making the location of key points more accurate. Specifically, as an optional embodiment, the step of recognizing the frame image and extracting the detection image corresponding to the detection box includes: analyzing the target frame image to determine the object descriptor box of the target object in the target frame image; analyzing the image within the object descriptor box using an object box recognition model to determine the first detection box of the target object in the target frame image; tracking and predicting the historical frame images of the target frame image to predict the second detection box of the target object in the target frame image, wherein the historical frame images include at least one frame image preceding the target frame image; and determining the target detection box of the target object in the target frame image based on the first and second detection boxes, and extracting the detection image corresponding to the target detection box. This embodiment can use target object recognition technology to identify the features of the target object in the target frame image and determine the object descriptor box of the target object in the target frame image. After determining the object descriptor box, on the one hand, a pre-trained object box recognition model can be used to analyze and adjust the image corresponding to the object descriptor box to obtain the first detection box; on the other hand, based on the continuity of the target object's motion in consecutive frame images, historical frame images can be used for tracking and prediction to analyze and obtain the second detection box of the target object in the target frame image. Then, based on the difference between the first and second detection boxes, the target detection box of the target object in the target frame image is determined, and the corresponding detection image is extracted for key point localization.

[0048] Specifically, on the one hand, the image corresponding to the object descriptor box can be input into a pre-trained object box recognition model. The object box recognition model can identify the key points of the target object and determine the position of each key point of the target object. The key points of the target object are key points that can characterize the features of the target object, such as key points representing features of a face, such as the nose, chin, forehead, and eyes. After determining the position of each key point, the target frame image can be adjusted to obtain the first detection box. Specifically, as an optional embodiment, the step of analyzing the image within the object descriptor box using the object box recognition model to determine the first detection box of the target object in the target frame image includes: determining the positioning information of the key points of the target object in the image of the object descriptor box using the object box recognition model; determining the object angle information of the target object in the target frame image based on the positioning information; rotating and adjusting the target frame image based on the object angle information, and combining this with the positioning information to determine the first detection box. In this scheme, the object bounding box recognition model is used to locate the key points of the target object in the object description box, obtain the positioning information, and determine the object angle information of the target object in the target frame image based on the positioning information. Then, the target frame image is adjusted based on the object angle information, and the first detection box of each key point of the external target object is determined based on the positioning information. This embodiment can use the object bounding box recognition model to more accurately locate the first detection box of the target object in the frame image, improving the accuracy of subsequent key point positioning.

[0049] On the other hand, based on the continuity of the target object's motion in consecutive frame images, the detection boxes corresponding to the previous (or more) historical frame images of the target frame image can be used for tracking and prediction to predict the second detection box of the target object within the target object. Combining the deviation between the first and second detection boxes, the first or second detection box is determined as the target detection box of the target object in the target frame image. Specifically, as an optional embodiment, determining the target detection box of the target object in the target frame image based on the first and second detection boxes includes: determining the deviation between the first and second detection boxes; if the deviation exceeds a deviation threshold, using the first detection box as the target detection box; if the deviation does not exceed the deviation threshold, using the second detection box as the target detection box. The deviation between the first and second detection boxes can include the deviation between the sides of the detection boxes. For example, if the first and second detection boxes are boxes containing four sides (top, bottom, left, and right), this embodiment can analyze the deviations of the four sides of the first and second detection boxes as the deviation amount, and compare the deviation amount of each side with a preset deviation threshold to obtain the target detection box.

[0050] In the training process of object bounding box recognition models, detection boxes of training images are usually labeled manually. However, manually labeled detection boxes may not be very accurate. Therefore, using the first detection box obtained by the model analysis alone as the target object detection box may result in significant fluctuations (inaccuracy). Therefore, in this embodiment, the deviation between the first and second detection boxes can be utilized. When the deviation between the first and second detection boxes is small, the second detection box, which is continuous in consecutive frames, is used as the target object detection box. When the deviation between the first and second detection boxes is large, the first detection box obtained by the object bounding box analysis model is used as the target object detection box. This can improve the accuracy of the target detection box.

[0051] After the detection box is determined, in step 206, the first position information and the second position information of the key points of the target object in the target frame image can be determined.

[0052] This embodiment utilizes a keypoint recognition model to identify the target frame image and determine the first position information. It then uses historical frame images of the target frame image for tracking and prediction to determine the second position information. Specifically, as an optional embodiment, determining the first and second position information of the target object's key points in the target frame image includes: inputting the target detection image corresponding to the target frame image into the keypoint recognition model to determine the first position information of the target object's key points in the target frame image; and tracking and predicting the positions of the target object's key points in historical frame images to predict the second position information of the target object's key points in the target frame image. On one hand, this embodiment can pre-train the keypoint recognition model, which is used to identify the input image (detection image), determine the first position information of each key point of the target object, and output it. The first position information can be the coordinates of the key points in the target frame image. On the other hand, this embodiment can track and predict the positions of the target object's key points in historical frame images based on the continuity of the target object's actions in consecutive frame images, predicting the second position information of the target object's key points in the target frame image. The second position information can be understood as the coordinates of the target object's key points in the target frame image.

[0053] In this embodiment, a keypoint recognition model can be trained using a single-frame image and training video, which can improve the accuracy of the trained recognition model. Specifically, as an optional embodiment, the method further includes the following steps for training the keypoint recognition model: training the keypoint recognition model using a single-frame image with training input images and training input annotations; inputting training images from the training video data into the keypoint recognition model that has completed single-frame image training to obtain training annotations representing the positions of keypoints of the target object in the training images; tracking and predicting historical training images of the target training image to determine the predicted training annotations of the target training image, wherein the historical training images include training images at least one frame prior to the target training image; verifying the credibility of the training annotations of the target training image based on the difference between the predicted training annotations and the training annotations; and training the keypoint recognition model that has completed single-frame image training using the verified training annotations and the corresponding training images.

[0054] Existing keypoint recognition models are typically trained using single-frame images (and their annotations). However, training with a large number of annotated single-frame images results in excessively high manual annotation costs, while training with a small number of annotated single-frame images can lead to low accuracy. Therefore, this embodiment utilizes a small number of annotated single-frame images for initial training (or single-frame image training). The pre-trained keypoint recognition model is then used to identify consecutive training images in the training video data to determine the corresponding training annotations. After determining the training annotations for each training image, the training annotations of the previous (or several previous) historical training images can be used for tracking and prediction to predict the predicted training annotations for the target training image. The predicted training annotations and the training annotations are then matched to determine the difference (the difference can be determined based on the differences between the positions of each keypoint). The difference is then compared with a preset difference threshold to verify the reliability of the training annotations for the target training image. For example, if the difference does not exceed the difference threshold, the training annotation verification is considered successful; if the difference exceeds the difference threshold, the training annotation verification is considered unsuccessful. After verifying the training labels, the key point recognition model can be trained using the verified training labels and corresponding training images.

[0055] In this embodiment, a small number of labeled single-frame images can be used to initially train the key point recognition model. Then, the pre-trained key point recognition model can be used to determine the training labels of the training images in the training video data. The continuity of the key points of the target object in the training video data can be used to verify whether the training labels are accurate and reliable. After that, the verified training images and training labels can be used to further train the key point recognition model, thereby increasing the amount of data for training the key point recognition model and improving the accuracy of the key point recognition model.

[0056] After determining the first and second location information of the key points of the target object, in step 208, the attribute information of the key points of the target object in the target frame image can be determined based on the first and second location information. In an optional embodiment, the first and second location information can be combined and analyzed to determine the third location information corresponding to the first and second location information, which serves as the attribute information of the key points of the target object in the target frame image. In another optional embodiment, one of the first and second location information can be used as the attribute information of the key points of the target object in the target frame image. Specifically, determining the attribute information of the key points of the target object in the target frame image based on the first and second location information includes: determining the offset between the first and second location information; if the offset exceeds an offset threshold, determining the attribute information of the key points of the target object in the target frame image based on the first location information; if the offset does not exceed the offset threshold, determining the attribute information of the key points of the target object in the target frame image based on the second location information. The attribute information of key points of a target object in a target frame image may include at least one of the following: position information (e.g., coordinates), spatial structure information (e.g., spatial relationships between key points), and depth information of the key points. This embodiment can determine the offset between the first and second position information, and then compare the offset with an offset threshold to determine the reliability of the second position information. When the offset is small (not exceeding the offset threshold), the second position information determined based on the continuity of the target object is used as the attribute information of the key points; when the offset is large (exceeding the offset threshold), the first position information determined based on the key point recognition model is used as the attribute information of the key points.

[0057] After determining the attribute information of the key points of the target object, special effects processing can be applied to the frame image according to the attribute information of the key points of the target object. Specifically, as an optional embodiment, the method further includes: applying special effects processing to the frame image in the video data based on the attribute information of the key points of the target object. Depending on the target object, different special effects processing methods can be adopted to process the frame image. For example, when the target object is a face, beauty and makeup effects can be applied to the face based on the attribute information of the key points of the face.

[0058] In this embodiment, frame images in video data can be identified to determine the detection image of the target object within the frame image. Then, on one hand, the target detection image of the target frame image can be input into a keypoint recognition model to determine the first position information of the keypoints of the target object in the target frame image. On the other hand, based on the continuity of the target object's movement in the video data, the positions of the keypoints of the target object in historical frame images preceding the target frame image can be tracked and analyzed to obtain the second position information of the keypoints of the target object in the target frame image. Subsequently, based on the first and second position information, the attribute information of the keypoints of the target object in the target frame image is obtained. Compared to the method of determining the first position information of the keypoints of the target object using only a keypoint recognition model, this embodiment also considers the continuous movement of the keypoints of the target object in the video data to analyze the second position information, and uses the first and second position information to determine the attribute information of the keypoints, which can improve the accuracy of keypoint localization of the target object.

[0059] Based on the above embodiments, this application also provides a data processing method, such as... Figure 3 As shown, the method includes:

[0060] Step 302: Acquire video data and identify the frame images in the video data.

[0061] Step 304: Analyze the target frame image to determine the object descriptor box of the target object in the target frame image.

[0062] Step 306: Determine the location information of the key points of the target object in the image of the object description box through the object box recognition model.

[0063] Step 308: Based on the positioning information, determine the object angle information of the target object in the target frame image.

[0064] Step 310: Based on the object angle information, rotate and adjust the target frame image, and combine it with the positioning information to determine the first detection box.

[0065] Step 312: Track and predict the historical frame images of the target frame image to predict the second detection box of the target object in the target frame image. The historical frame images include at least one frame image before the target frame image.

[0066] Step 314: Determine the deviation between the first detection frame and the second detection frame, and determine whether the deviation exceeds the deviation threshold.

[0067] Step 316: If the deviation exceeds the deviation threshold, the first detection box is used as the target detection box in order to extract the detection image.

[0068] Step 318: If the deviation amount does not exceed the deviation threshold, the second detection box is used as the target detection box in order to extract the detection image.

[0069] Step 320: Input the target detection image corresponding to the target frame image into the keypoint recognition model to determine the first position information of the keypoints of the target object in the target frame image. As an optional embodiment, the method further includes a keypoint recognition model training step: using training input images and training input annotations, train the keypoint recognition model on a single frame image; input the training images from the training video data into the keypoint recognition model that has completed single-frame image training to obtain training annotations representing the positions of the keypoints of the target object in the training images; track and predict the historical training images of the target training images to determine the predicted training annotations of the target training images, wherein the historical training images include training images at least one frame prior to the target training image; verify the credibility of the training annotations of the target training images based on the difference between the predicted training annotations and the training annotations; and train the keypoint recognition model that has completed single-frame image training using the verified training annotations and corresponding training images.

[0070] Step 322: Track and predict the position of the key points of the target object in the historical frame image, and predict the second position information of the key points of the target object in the target frame image, wherein the historical frame image includes at least one frame image preceding the target frame image.

[0071] Step 324: Determine the offset between the first location information and the second location information, and determine whether the offset exceeds the deviation threshold.

[0072] Step 326: If the offset exceeds the offset threshold, determine the attribute information of the key points of the target object in the target frame image based on the first position information.

[0073] Step 328: If the offset does not exceed the offset threshold, determine the attribute information of the key points of the target object in the target frame image based on the second position information.

[0074] Step 330: Based on the attribute information of the key points of the target object, perform special effects processing on the frame images in the video data.

[0075] In this embodiment, based on frame images in video data, an object descriptor box for the target object in the target frame image can be determined. Then, based on an object box recognition model, a first detection box can be determined. Historical frame images can be tracked and predicted to determine a second detection box. The first or second detection box is then used as the target object's detection box, and the corresponding detection image is extracted for subsequent keypoint localization. After determining the detection image containing the target object, on one hand, the target detection image can be input into a keypoint recognition model to determine the first position information of the target object's keypoints in the target frame image; on the other hand, based on the continuity of the target object's motion in the video data, historical frame images preceding the target frame image can be tracked and analyzed to obtain the second position information of the target object's keypoints in the target frame image. Then, the first or second position information is determined as the attribute information of the target object's keypoints, and special effects processing is applied to the frame image based on this attribute information.

[0076] Based on the above embodiments, this application also provides a data processing method. This method can train a keypoint recognition model with high accuracy using a small amount of labeled single-frame image and video data. Specifically, for example... Figure 4 As shown, the method includes:

[0077] Step 402: Use the training input image and training input annotation to train the key point recognition model using a single frame image.

[0078] Step 404: Input the training images from the training video data into the key point recognition model that has completed single-frame image training to obtain training annotations that represent the positions of key points of the target object in the training images.

[0079] Step 406: Track and predict the historical training images of the target training image to determine the predicted training labels for the training image at the target time.

[0080] Step 408: Verify the credibility of the training labels of the target training image based on the difference between the predicted training labels and the training labels of the training image at the target time.

[0081] Step 410: Using the validated training annotations and corresponding training images, the key point recognition model for single-frame image training was successfully trained.

[0082] The implementation methods of this application are similar to those of the above embodiments. For specific implementation methods, please refer to the specific implementation methods of the above embodiments, which will not be repeated here.

[0083] Existing keypoint recognition models are typically trained using single-frame images (and their annotations). However, training with a large number of annotated single-frame images results in excessively high manual annotation costs, while training with a small number of annotated single-frame images can lead to low accuracy. Therefore, this embodiment utilizes a small number of annotated single-frame images for initial training (or single-frame image training). The pre-trained keypoint recognition model is then used to identify consecutive training images in the training video data to determine the corresponding training annotations. After determining the training annotations for each training image, the training annotations of the previous (or several previous) historical training images can be used for tracking and prediction to predict the predicted training annotations for the target training image. The predicted training annotations and the training annotations are then matched to determine the difference (the difference can be determined based on the differences between the positions of each keypoint). The difference is then compared with a preset difference threshold to verify the reliability of the training annotations for the target training image. For example, if the difference does not exceed the difference threshold, the training annotation verification is considered successful; if the difference exceeds the difference threshold, the training annotation verification is considered unsuccessful. After verifying the training labels, the key point recognition model can be trained using the verified training labels and corresponding training images.

[0084] In this embodiment, a small number of labeled single-frame images can be used to initially train the key point recognition model. Then, the pre-trained key point recognition model can be used to determine the training labels of the training images in the training video data. The continuity of the key points of the target object in the training video data can be used to verify whether the training labels are accurate and reliable. After that, the verified training images and training labels can be used to further train the key point recognition model, thereby increasing the amount of data for training the key point recognition model and improving the accuracy of the key point recognition model.

[0085] Based on the above embodiments, this application also provides a data processing method. This method allows for more accurate and stable localization of key points of a target object using consecutive frames in video data. This allows for the application of special effects processing to the frame images based on the location of the key points. For example, it can be applied to scenarios involving the localization of key points of competitors in sports videos to extract more accurate competitor images and apply special effects to these images to create playback data. Specifically, as... Figure 5 As shown, the method includes:

[0086] Step 502: Acquire video data and identify the frame images in the video data.

[0087] Step 504: Recognize the frame image and extract the detection image corresponding to the detection box. The detection image contains the target object located by the detection box.

[0088] Step 506: Determine the first position information and the second position information of the key points of the target object in the target frame image. The first position information is determined based on the target detection image of the target frame image. The second position information is determined by tracking and predicting the position of the key points of the target object in historical frame images. The historical frame images include at least one frame image before the target frame image.

[0089] Step 508: Based on the first location information and the second location information, determine the attribute information of the key points of the target object in the target frame image.

[0090] Step 510: Based on the attribute information of the key points of the target object, perform special effects processing on the frame images in the video data.

[0091] The implementation methods of this application are similar to those of the above embodiments. For specific implementation methods, please refer to the specific implementation methods of the above embodiments, which will not be repeated here.

[0092] In this embodiment, the video data can be road surveillance video, sports event video, or live video. This embodiment can identify frame images in the video data, determine the detection box of the target object in the frame image, and extract the detection image corresponding to the detection box. Then, on the one hand, the target detection image of the target frame image can be input into the key point recognition model to determine the first position information of the key points of the target object in the target frame image; on the other hand, based on the continuity of the target object's movement in the video data, the position of the key points of the target object in the historical frame images before the target frame image can be tracked and analyzed to obtain the second position information of the key points of the target object in the target frame image. Then, based on the first and second position information, the attribute information of the key points of the target object in the target frame image is obtained. Then, this attribute information can be used to perform special effects processing on the frame images in the video data. Depending on the target object, different special effects processing methods can be adopted to process the frame images. For example, when the target object is a face, the face can be processed with special effects such as beautification and makeup based on the attribute information of the key points of the face.

[0093] Based on the above embodiments, this application also provides a data processing method. This method can perform key point localization on frame images containing facial data in video data, which can improve the accuracy of key points in facial data. This allows for the use of attribute information from the key points in facial data to perform beautification and makeup processing on the facial data in the frame images, thereby improving the beautification and makeup effects and enhancing the image quality of the video data. Specifically, as shown... Figure 6 As shown, the method includes:

[0094] Step 602: Acquire video data and determine the frame images in the video data.

[0095] Step 604: Recognize the frame image and extract the detection image corresponding to the detection box. The detection image contains facial data located by the detection box.

[0096] Step 606: Determine the first and second position information of the key points of the facial data in the target frame image. The first position information is determined based on the target detection image of the target frame image. The second position information is determined by tracking and predicting the position of the key points of the facial data in historical frame images. The historical frame images include at least one frame image before the target frame image.

[0097] Step 608: Based on the first location information and the second location information, determine the attribute information of the key points of the facial data in the target frame image.

[0098] Step 610: Based on the attribute information of key points in the facial data, perform special effects processing on the frame images in the video data. The special effects processing includes at least one of beautification processing, special effects addition processing, makeup processing, and beauty processing.

[0099] The implementation methods of this application are similar to those of the above embodiments. For specific implementation methods, please refer to the specific implementation methods of the above embodiments, which will not be repeated here.

[0100] In this embodiment, frame images in video data can be identified to determine the detection boxes for facial data within the frame images, and the corresponding detection images can be extracted. Then, on one hand, the target detection image of the target frame image can be input into a keypoint recognition model to determine the first position information of the keypoints of the facial data in the target frame image; on the other hand, based on the continuity of facial data movement in the video data, the positions of the keypoints of the facial data in historical frame images preceding the target frame image can be tracked and analyzed to obtain the second position information of the keypoints of the facial data in the target frame image. Subsequently, based on the first and second position information, the attribute information of the keypoints of the facial data in the target frame image is obtained. This attribute information can then be used to perform special effects processing such as beautification, enhancement, makeup, and adding special effects to the face.

[0101] Based on the above embodiments, this application also provides a data processing method. This method allows for more accurate and stable localization of key points on the anchor's face based on consecutive frame images in live video data. This allows for the use of these key point locations to apply special effects to the facial data in the frame images, such as beautification and makeup effects, thereby improving the beautification and makeup effects on the anchor's face and enhancing the overall image quality of the live video data. Specifically, for example... Figure 7 As shown, the method includes:

[0102] Step 702: Acquire live video data, identify the frame images in the live video data, and extract the detection images corresponding to the detection boxes. The detection images contain facial data located by the detection boxes.

[0103] Step 704: Determine the first position information and the second position information of the key points of the facial data in the target frame image. The first position information is determined based on the target detection image of the target frame image. The second position information is determined by tracking and predicting the position of the key points of the facial data in historical frame images. The historical frame images include at least one frame image before the target frame image.

[0104] Step 706: Based on the first location information and the second location information, determine the attribute information of the key points of the facial data in the target frame image.

[0105] Step 708: Based on the attribute information of key points in the facial data, perform special effects processing on the frame images in the live video data. The special effects processing includes at least one of beautification processing, special effects addition processing, makeup processing, and beauty processing.

[0106] The implementation methods of this application are similar to those of the above embodiments. For specific implementation methods, please refer to the specific implementation methods of the above embodiments, which will not be repeated here.

[0107] In this embodiment, frame images in live video data can be identified to determine the detection box of the anchor's facial data in the frame image, and the detection image corresponding to the detection box can be extracted. Then, on the one hand, the target detection image of the target frame image can be input into the key point recognition model to determine the first position information of the anchor's facial key points in the target frame image; on the other hand, based on the continuity of facial data movement in the video data, the positions of the facial key points in historical frame images before the target frame image can be tracked and analyzed to obtain the second position information of the facial key points in the target frame image. Then, based on the first and second position information, the attribute information of the facial key points in the target frame image is obtained. Then, this attribute information can be used to perform special effects processing such as beautification, beauty enhancement, makeup, and special effects addition on the anchor.

[0108] Based on the above embodiments, this application also provides a data processing method that can be applied to Software-as-a-Service (SaaS) scenarios. The above process is encapsulated as a key point recognition service, thereby providing key point recognition services for video images to users (such as enterprise users).

[0109] Specifically, the processing end (or server) provides users with video data upload interfaces and result distribution interfaces. Enterprise users can upload corresponding video data through the upload interface. Then, the processing end can more accurately and stably locate key points of the target object based on continuous frame images in the video data, obtaining recognition results (attribute information of key points in the image). The processing end can then distribute the recognition results of the video data to the enterprise user through the distribution interface. The enterprise user can then optimize the video data based on the recognition results. Additionally, the processing end can also perform corresponding processing on the video data based on the recognition results (such as adding special effects), and then distribute the processed video data to the user through the distribution interface. Specifically, for example... Figure 8 As shown, the method includes:

[0110] Step 802: Provide a first interface to obtain relevant video data and determine the frame images in the video data.

[0111] Step 804: Recognize the frame image and extract the detection image corresponding to the detection box. The detection image contains the target object located by the detection box.

[0112] Step 806: Determine the first position information and the second position information of the key points of the target object in the target frame image. The first position information is determined based on the target detection image of the target frame image. The second position information is determined by tracking and predicting the position of the key points of the target object in historical frame images. The historical frame images include at least one frame image before the target frame image.

[0113] Step 808: Based on the first location information and the second location information, determine the attribute information of the key points of the target object in the target frame image.

[0114] Step 810: Feed back the attribute information through the second interface.

[0115] The implementation methods of this application are similar to those of the above embodiments. For specific implementation methods, please refer to the specific implementation methods of the above embodiments, which will not be repeated here.

[0116] The method in this embodiment can be applied to a processing end, which can be understood as a platform. The processing end can provide multiple interfaces to offer corresponding services. Users (such as enterprises) can upload video data through the first interface. It should be noted that the first interface can receive video data relayed by the enterprise or directly connect to a video data acquisition device (such as a camera) to obtain video data. In this embodiment, the interface represents the connection between two devices (such as hardware connection, network connection, etc.). After obtaining video data through the first interface, this embodiment can identify the frame images in the video data, determine the detection box of the target object in the frame image, and extract the detection image corresponding to the detection box. Then, on the one hand, the target detection image of the target frame image can be input into the key point recognition model to determine the first position information of the key points of the target object in the target frame image; on the other hand, based on the continuity of the target object's movement in the video data, the position of the key points of the target object in historical frame images before the target frame image can be tracked and analyzed to obtain the second position information of the key points of the target object in the target frame image. Then, based on the first and second position information, the attribute information of the key points of the target object in the target frame image is obtained. Then, on the one hand... The processing unit can send the attribute information of key points to the user through a second interface. On the user's side, the video data can be optimized based on the recognition results, such as applying beauty filters or makeup to the people in the video data. In addition, the processing unit can also perform corresponding processing on the video data based on the attribute information of key points (such as adding special effects), and then send the processed video data to the user through the second interface.

[0117] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of this application are not limited to the described order of actions, because according to the embodiments of this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of this application.

[0118] Based on the above embodiments, this embodiment also provides a data processing apparatus, referring to... Figure 9 Specifically, it can include the following modules:

[0119] The video data acquisition module 902 is used to acquire video data and determine the frame images in the video data.

[0120] The detection image acquisition module 904 is used to identify the frame image and extract the detection image corresponding to the detection box, wherein the detection image contains the target object located by the detection box.

[0121] The location information acquisition module 906 is used to determine the first location information and the second location information of the key points of the target object in the target frame image. The first location information is determined based on the target detection image of the target frame image, and the second location information is determined by tracking and predicting the position of the key points of the target object in historical frame images. The historical frame images include at least one frame image before the target frame image.

[0122] The attribute information acquisition module 908 is used to determine the attribute information of the key points of the target object in the target frame image based on the first position information and the second position information.

[0123] In summary, in this embodiment, frame images in video data can be identified, and detection images containing target objects can be extracted. Then, on the one hand, the target detection image of the target frame image can be input into a keypoint recognition model to determine the first position information of the keypoints of the target object in the target frame image; on the other hand, based on the continuity of the target object's movement in the video data, the position of the keypoints of the target object in historical frame images can be tracked and analyzed to obtain the second position information of the keypoints of the target object in the target frame image. Subsequently, based on the first and second position information, the attribute information of the keypoints of the target object in the target frame image is obtained. Compared with the method of determining the first position information of the keypoints of the target object using only a keypoint recognition model, this embodiment also considers the continuous movement of the keypoints of the target object in the video data to analyze the second position information, and determines the attribute information of the keypoints based on the first and second position information, which can improve the accuracy of keypoint localization of the target object.

[0124] Based on the above embodiments, this embodiment also provides a data processing device, which may specifically include the following modules:

[0125] The video data acquisition and processing module is used to acquire video data and determine the frame images in the video data.

[0126] The object descriptor box acquisition and processing module is used to analyze the target frame image and determine the object descriptor box of the target object in the target frame image.

[0127] The positioning information acquisition and processing module is used to determine the positioning information of key points of the target object in the image of the object description box through the object box recognition model.

[0128] The object angle acquisition and processing module is used to determine the object angle information of the target object in the target frame image based on the positioning information.

[0129] The first detection box acquisition and processing module is used to rotate and adjust the target frame image based on the object angle information, and determine the first detection box in combination with the positioning information.

[0130] The second detection box acquisition and processing module is used to track and predict the historical frame images of the target frame image and predict the second detection box of the target object in the target frame image. The historical frame images include at least one frame image before the target frame image.

[0131] The deviation acquisition and processing module is used to determine the deviation between the first detection frame and the second detection frame, and to determine whether the deviation exceeds the deviation threshold.

[0132] The first deviation processing module is used to treat the first detection box as the target detection box when the deviation exceeds the deviation threshold.

[0133] The second deviation processing module is used to use the second detection box as the target detection box when the deviation amount does not exceed the deviation threshold.

[0134] The first position acquisition and processing module is used to input the target detection image corresponding to the target frame image into the key point recognition model to determine the first position information of the key points of the target object in the target frame image.

[0135] The second position acquisition and processing module is used to track and predict the position of key points of the target object in historical frame images, and predict the second position information of the key points of the target object in the target frame image, wherein the historical frame image includes at least one frame image preceding the target frame image.

[0136] The offset acquisition and processing module is used to determine the offset between the first position information and the second position information, and to determine whether the offset exceeds the deviation threshold.

[0137] The first offset processing module is used to determine the attribute information of the key points of the target object in the target frame image based on the first position information when the offset exceeds the offset threshold.

[0138] The second offset processing module is used to determine the attribute information of the key points of the target object in the target frame image based on the second position information, provided that the offset amount does not exceed the offset threshold.

[0139] The special effects image acquisition and processing module is used to perform special effects processing on frame images in video data based on the attribute information of key points of the target object.

[0140] In this embodiment, based on frame images in video data, an object descriptor box for the target object in the target frame image can be determined. Then, based on an object box recognition model, a first detection box can be determined. Historical frame images can be tracked and predicted to determine a second detection box. The first or second detection box is then used as the target object's detection box, and the corresponding detection image is extracted for subsequent keypoint localization. After determining the detection image containing the target object, on one hand, the target detection image can be input into a keypoint recognition model to determine the first position information of the target object's keypoints in the target frame image; on the other hand, based on the continuity of the target object's motion in the video data, historical frame images preceding the target frame image can be tracked and analyzed to obtain the second position information of the target object's keypoints in the target frame image. Then, the first or second position information is determined as the attribute information of the target object's keypoints, and special effects processing is applied to the frame image based on this attribute information.

[0141] Based on the above embodiments, this embodiment also provides a data processing apparatus, referring to... Figure 10 Specifically, it can include the following modules:

[0142] The single-frame training module 1002 is used to train the key point recognition model using a single-frame image with the training input image and the training input annotation.

[0143] The training annotation acquisition module 1004 is used to input the training images in the training video data into the key point recognition model that has completed single-frame image training, and obtain training annotations that represent the positions of key points of the target object in the training images.

[0144] The prediction annotation acquisition module 1006 is used to track and predict the historical training images of the target training image and determine the predicted training annotations of the training image at the target time.

[0145] The training annotation verification module 1008 is used to verify the credibility of the training annotation of the target training image based on the difference between the predicted training annotation and the training annotation of the training image at the target time.

[0146] The model secondary training module 1010 is used to train a key point recognition model for a single frame image by using the validated training annotations and the corresponding training images.

[0147] In summary, existing keypoint recognition models are typically trained using single-frame images (and their annotations). However, training with a large number of annotated single-frame images results in excessively high manual annotation costs, while training with a small number of annotated single-frame images can lead to low accuracy. Therefore, this embodiment utilizes a small number of annotated single-frame images for initial training (or single-frame image training). The pre-trained keypoint recognition model is then used to identify consecutive training images in the training video data to determine the corresponding training annotations. After determining the training annotations for each training image, the training annotations of the previous (or previous multiple) historical training images can be used for tracking and prediction to predict the predicted training annotations for the target training image. The predicted training annotations and the training annotations are then matched to determine the difference (the difference can be determined based on the differences between the positions of each keypoint). The difference is then compared with a preset difference threshold to verify the reliability of the training annotations for the target training image. For example, if the difference does not exceed the difference threshold, the training annotation verification is considered successful; if the difference exceeds the difference threshold, the training annotation verification is considered unsuccessful. After verifying the training labels, the key point recognition model can be trained using the verified training labels and corresponding training images.

[0148] In this embodiment, a small number of labeled single-frame images can be used to initially train the key point recognition model. Then, the pre-trained key point recognition model can be used to determine the training labels of the training images in the training video data. The continuity of the key points of the target object in the training video data can be used to verify whether the training labels are accurate and reliable. After that, the verified training images and training labels can be used to further train the key point recognition model, thereby increasing the amount of data for training the key point recognition model and improving the accuracy of the key point recognition model.

[0149] Based on the above embodiments, this embodiment also provides a data processing apparatus, referring to... Figure 11 Specifically, it can include the following modules:

[0150] The video data determination module 1102 is used to acquire video data and determine the frame images in the video data.

[0151] The detection image determination module 1104 is used to identify the frame image and extract the detection image corresponding to the detection box, wherein the detection image contains the target object located by the detection box.

[0152] The location information determination module 1106 is used to determine the first location information and the second location information of the key points of the target object in the target frame image. The first location information is determined based on the target detection image of the target frame image, and the second location information is determined by tracking and predicting the position of the key points of the target object in historical frame images. The historical frame images include at least one frame image before the target frame image.

[0153] The attribute information determination module 1108 is used to determine the attribute information of the key points of the target object in the target frame image based on the first position information and the second position information.

[0154] The special effects image determination module 1110 is used to perform special effects processing on frame images in video data based on the attribute information of key points of the target object.

[0155] In summary, in this embodiment, frame images in video data can be identified to determine the detection box of the target object in the frame image and extract the detection image corresponding to the detection box. Then, on the one hand, the target detection image of the target frame image can be input into the key point recognition model to determine the first position information of the key points of the target object in the target frame image; on the other hand, based on the continuity of the target object's movement in the video data, the position of the key points of the target object in historical frame images before the target frame image can be tracked and analyzed to obtain the second position information of the key points of the target object in the target frame image. Then, based on the first and second position information, the attribute information of the key points of the target object in the target frame image is obtained. Then, this attribute information can be used to perform special effects processing on the frame images in the video data. Depending on the target object, different special effects processing methods can be adopted to process the frame images. For example, when the target object is a face, the face can be processed with beautification, makeup, and other special effects based on the attribute information of the key points of the face.

[0156] Based on the above embodiments, this embodiment also provides a data processing apparatus, referring to... Figure 12 Specifically, it can include the following modules:

[0157] The video data acquisition module 1202 is used to acquire video data and determine the frame images in the video data.

[0158] The detection image acquisition module 1204 is used to identify the frame image and extract the detection image corresponding to the detection box. The detection image contains facial data of the detection box location.

[0159] The location information acquisition module 1206 is used to determine the first location information and the second location information of the key points of the facial data in the target frame image. The first location information is determined based on the target detection image of the target frame image, and the second location information is determined by tracking and predicting the position of the key points of the facial data in historical frame images. The historical frame images include at least one frame image before the target frame image.

[0160] The attribute information acquisition module 1208 is used to determine the attribute information of key points of facial data in the target frame image based on the first location information and the second location information.

[0161] The special effects image acquisition module 1210 is used to perform special effects processing on frame images in video data based on the attribute information of key points in facial data. The special effects processing includes at least one of beautification processing, special effects addition processing, makeup processing, and beauty processing.

[0162] In summary, in this embodiment, frame images in video data can be identified to determine the detection boxes for facial data within the frame images, and the detection images corresponding to the detection boxes can be extracted. Then, on one hand, the target detection image of the target frame image can be input into a keypoint recognition model to determine the first position information of the keypoints of the facial data in the target frame image; on the other hand, based on the continuity of facial data movement in the video data, the positions of the keypoints of the facial data in historical frame images preceding the target frame image can be tracked and analyzed to obtain the second position information of the keypoints of the facial data in the target frame image. Subsequently, based on the first and second position information, the attribute information of the keypoints of the facial data in the target frame image is obtained. This attribute information can then be used to perform special effects processing such as beautification, enhancement, makeup, and adding special effects to the face.

[0163] Based on the above embodiments, this embodiment also provides a data processing apparatus, referring to... Figure 13 Specifically, it can include the following modules:

[0164] The video data acquisition module 1302 is used to acquire live video data, identify the frame images in the live video data, and extract the detection image corresponding to the detection box. The detection image contains facial data located by the detection box.

[0165] The location information acquisition module 1304 is used to determine the first location information and the second location information of the key points of the facial data in the target frame image. The first location information is determined based on the target detection image of the target frame image, and the second location information is determined by tracking and predicting the position of the key points of the facial data in historical frame images. The historical frame images include at least one frame image before the target frame image.

[0166] The attribute information acquisition module 1306 is used to determine the attribute information of the key points of the facial data in the target frame image based on the first location information and the second location information.

[0167] The special effects image acquisition module 1308 is used to perform special effects processing on frame images in live video data based on the attribute information of key points in facial data. The special effects processing includes at least one of beautification processing, special effects addition processing, makeup processing, and beauty processing.

[0168] In summary, in this embodiment, frame images in live video data can be identified to determine the detection box of the anchor's facial data in the frame image, and the detection image corresponding to the detection box can be extracted. Then, on the one hand, the target detection image of the target frame image can be input into the key point recognition model to determine the first position information of the anchor's facial key points in the target frame image; on the other hand, based on the continuity of facial data movement in the video data, the positions of the facial key points in historical frame images before the target frame image can be tracked and analyzed to obtain the second position information of the facial key points in the target frame image. Then, based on the first and second position information, the attribute information of the facial key points in the target frame image is obtained. Then, this attribute information can be used to perform special effects processing such as beautification, beauty enhancement, makeup, and special effects addition on the anchor.

[0169] Based on the above embodiments, this embodiment also provides a data processing apparatus, referring to... Figure 14 Specifically, it can include the following modules:

[0170] Service providing module 1402 is used to provide a first interface to obtain relevant video data and determine the frame images in the video data.

[0171] The service processing module 1404 is used to identify the frame image, extract the detection image corresponding to the detection box, the detection image containing the target object located by the detection box; determine the first position information and the second position information of the key points of the target object in the target frame image, the first position information is determined based on the target detection image of the target frame image, and the second position information is determined by tracking and predicting the position of the key points of the target object in historical frame images, the historical frame images including at least one frame image before the target frame image; and determine the attribute information of the key points of the target object in the target frame image based on the first position information and the second position information.

[0172] The result feedback module 1406 is used to feed back the attribute information through the second interface.

[0173] In this embodiment, the application can be applied to a processing end, which can be understood as a platform. The processing end can provide multiple interfaces to offer corresponding services. Users (such as enterprises) can upload video data through a first interface. It should be noted that the first interface can receive video data relayed by the enterprise or directly connect to a video data acquisition device (such as a camera) to obtain video data. After obtaining the video data through the first interface, this embodiment can identify the frame images in the video data, determine the detection box of the target object in the frame image, and extract the detection image corresponding to the detection box. Then, on the one hand, the target detection image of the target frame image can be input into a key point recognition model to determine the first position information of the key points of the target object in the target frame image; on the other hand, based on the continuity of the target object's movement in the video data, the position of the key points of the target object in historical frame images before the target frame image can be tracked and analyzed to obtain the second position information of the key points of the target object in the target frame image. Then, based on the first and second position information, the attribute information of the key points of the target object in the target frame image is obtained. Then, on the one hand... The processing unit can send the attribute information of key points to the user through a second interface. On the user's side, the video data can be optimized based on the recognition results, such as applying beauty filters or makeup to the people in the video data. In addition, the processing unit can also perform corresponding processing on the video data based on the attribute information of key points (such as adding special effects), and then send the processed video data to the user through the second interface.

[0174] This application also provides a non-volatile readable storage medium storing one or more modules (programs). When these modules are applied to a device, they enable the device to execute the instructions for the method steps in this application.

[0175] This application provides one or more machine-readable media storing instructions that, when executed by one or more processors, cause an electronic device to perform one or more of the methods described in the above embodiments. In this application, the electronic device includes devices such as servers and terminal devices.

[0176] Embodiments of this disclosure can be implemented as an apparatus with any suitable hardware, firmware, software, or any combination thereof, configured as desired, and the apparatus may include electronic devices such as servers (clusters) and terminals. Figure 15 An exemplary apparatus 1500 is schematically shown that can be used to implement the various embodiments described in this application.

[0177] In one embodiment, Figure 15 An exemplary device 1500 is shown, which includes one or more processors 1502, a control module (chipset) 1504 coupled to at least one of the processors 1502, a memory 1506 coupled to the control module 1504, a non-volatile memory (NVM) / storage device 1508 coupled to the control module 1504, one or more input / output devices 1510 coupled to the control module 1504, and a network interface 1512 coupled to the control module 1504.

[0178] Processor 1502 may include one or more single-core or multi-core processors, and processor 1502 may include any combination of general-purpose processors or special-purpose processors (e.g., graphics processors, application processors, baseband processors, etc.). In some embodiments, device 1500 can serve as a server, terminal, or other device as described in the embodiments of this application.

[0179] In some embodiments, apparatus 1500 may include one or more computer-readable media (e.g., memory 1506 or NVM / storage device 1508) having instructions 1514 and one or more processors 1502 that are combined with the one or more computer-readable media and configured to execute instructions 1514 to implement a module thereby performing the actions described in this disclosure.

[0180] In one embodiment, the control module 1504 may include any suitable interface controller to provide any suitable interface to at least one of the processors 1502 and / or any suitable device or component communicating with the control module 1504.

[0181] The control module 1504 may include a memory controller module to provide an interface to the memory 1506. The memory controller module may be a hardware module, a software module, and / or a firmware module.

[0182] Memory 1506 may be used, for example, to load and store data and / or instructions 1514 for device 1500. In one embodiment, memory 1506 may include any suitable volatile memory, such as suitable DRAM. In some embodiments, memory 1506 may include double data rate type quad synchronous dynamic random access memory (DDR4 SDRAM).

[0183] In one embodiment, the control module 1504 may include one or more input / output controllers to provide interfaces to the NVM / storage device 1508 and (one or more) input / output devices 1510.

[0184] For example, NVM / storage device 1508 may be used to store data and / or instructions 1514. NVM / storage device 1508 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drives (HDDs), one or more optical disc drives (CDs), and / or one or more digital universal optical disc (DVD) drives).

[0185] NVM / storage device 1508 may include storage resources that are part of a device on which device 1500 is mounted, or that are accessible by the device but do not necessarily have to be part of the device. For example, NVM / storage device 1508 may be accessed via a network via one or more input / output devices 1510.

[0186] One or more input / output devices 1510 may provide an interface for device 1500 to communicate with any other suitable device. Input / output devices 1510 may include communication components, audio components, sensor components, etc. Network interface 1512 may provide an interface for device 1500 to communicate via one or more networks. Device 1500 may wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, such as accessing wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G, 5G, etc., or combinations thereof.

[0187] In one embodiment, at least one of the processors 1502 may be logically packaged with one or more controllers (e.g., memory controller modules) of the control module 1504. In one embodiment, at least one of the processors 1502 may be logically packaged with one or more controllers of the control module 1504 to form a system-in-package (SiP). In one embodiment, at least one of the processors 1502 may be integrated with the logic of one or more controllers of the control module 1504 on the same die. In one embodiment, at least one of the processors 1502 may be integrated with the logic of one or more controllers of the control module 1504 on the same die to form a system-on-a-chip (SoC).

[0188] In various embodiments, device 1500 may be, but is not limited to, a server, desktop computing device, or mobile computing device (e.g., laptop, handheld computing device, tablet, netbook, etc.). In various embodiments, device 1500 may have more or fewer components and / or a different architecture. For example, in some embodiments, device 1500 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.

[0189] The detection device can use a main control chip as a processor or control module, and sensor data, position information, etc. can be stored in a memory or NVM / storage device. The sensor group can be used as an input / output device, and the communication interface can include a network interface.

[0190] This application also provides an electronic device, including: a processor; and a memory storing executable code thereon, which, when executed, causes the processor to perform one or more methods as described in this application.

[0191] This application also provides one or more machine-readable media having executable code stored thereon, which, when executed, causes a processor to perform one or more of the methods described in this application.

[0192] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0193] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0194] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0195] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0196] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0197] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.

[0198] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0199] The above provides a detailed description of a data processing method, a data processing device, an electronic device, and a storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A data processing method, characterized by, The method comprises: acquiring video data and determining frame images in the video data; identifying the frame images, extracting a detection image corresponding to a detection box, the detection image containing a target object positioned by the detection box; determining first position information and second position information of a key point of the target object in a target frame image, the first position information being determined according to a target detection image of the target frame image, and the second position information being determined by tracking and predicting the position of the key point of the target object in a historical frame image, the historical frame image including at least one frame of frame image before the target frame image; determining an offset between the first position information and the second position information; in a case where the offset exceeds an offset threshold, determining attribute information of the key point of the target object in the target frame image according to the first position information; in a case where the offset does not exceed the offset threshold, determining the attribute information of the key point of the target object in the target frame image according to the second position information.

2. The method of claim 1, wherein, The determination of the first position information and the second position information of the key point of the target object in the target frame image comprises: inputting a target detection image corresponding to the target frame image into a key point identification model to determine the first position information of the key point of the target object in the target frame image; tracking and predicting the position of the key point of the target object in the historical frame image to predict the second position information of the key point of the target object in the target frame image.

3. The method of claim 1, wherein, The identification of the frame images and the extraction of the detection image corresponding to the detection box comprise: analyzing the target frame image to determine an object description box of the target object in the target frame image; analyzing the image in the object description box by an object box identification model to determine a first detection box of the target object in the target frame image; tracking and predicting the historical frame image of the target frame image to predict a second detection box of the target object in the target frame image, the historical frame image including at least one frame of frame image before the target frame image; determining a target detection box of the target object in the target frame image according to the first detection box and the second detection box, and extracting a detection image corresponding to the target detection box.

4. The method of claim 3, wherein, The analysis of the image in the object description box by the object box identification model to determine the first detection box of the target object in the target frame image comprises: determining positioning information of the key point of the target object in the image of the object description box by the object box identification model; determining object angle information of the target object in the target frame image according to the positioning information; rotating and adjusting the target frame image according to the object angle information, and determining the first detection box in combination with the positioning information.

5. The method of claim 3, wherein, The determination of the target detection box of the target object in the target frame image according to the first detection box and the second detection box comprises: determining a deviation between the first detection box and the second detection box; in a case where the deviation exceeds a deviation threshold, taking the first detection box as the target detection box; in a case where the deviation does not exceed the deviation threshold, taking the second detection box as the target detection box.

6. The method of claim 1, wherein, The method further comprises a training step of the key point identification model: The key point recognition model is trained by using the training input image and the training input label; The training image in the training video data is input into the key point recognition model which has completed the single-frame image training, so as to obtain the training label representing the position of the key point of the target object in the training image; The historical training image of the target training image is tracked and predicted to determine the predicted training label of the target training image, wherein the historical training image includes at least one frame of training image before the target training image; The credibility of the training label of the target training image is verified according to the difference between the predicted training label of the target training image and the training label; The key point recognition model which has completed the single-frame image training is trained by using the training label which passes the verification and the corresponding training image.

7. The method of claim 1, wherein, Further comprising: According to the attribute information of the key point of the target object, the frame image in the video data is processed with special effects.

8. A data processing method, characterized by, It includes: The key point recognition model is trained by using the training input image and the training input label; The training image in the training video data is input into the key point recognition model which has completed the single-frame image training, so as to obtain the training label representing the position of the key point of the target object in the training image; The historical training image of the target training image is tracked and predicted to determine the predicted training label of the target training image, wherein the historical training image includes at least one frame of training image before the target training image; The credibility of the training label of the target training image is verified according to the difference between the predicted training label of the target training image and the training label; The key point recognition model which has completed the single-frame image training is trained by using the training label which passes the verification and the corresponding training image.

9. A data processing method, characterized by, It includes: Obtain video data and determine frame images in the video data; The frame image is identified to extract a detection image corresponding to a detection box, and the detection image contains a target object positioned by the detection box; Determine the first position information and the second position information of the key point of the target object in the target frame image, wherein the first position information is determined according to the target detection image of the target frame image, and the second position information is determined by tracking and predicting the position of the key point of the target object in the historical frame image, and the historical frame image includes at least one frame of frame image before the target frame image; Determine the offset between the first position information and the second position information; In the case where the offset exceeds the offset threshold, the attribute information of the key point of the target object in the target frame image is determined according to the first position information; In the case where the offset does not exceed the offset threshold, the attribute information of the key point of the target object in the target frame image is determined according to the second position information; According to the attribute information of the key point of the target object, the frame image in the video data is processed with special effects.

10. A data processing method, characterized by, It includes: Obtain video data and determine frame images in the video data; The frame image is identified to extract a detection image corresponding to a detection box, and the detection image contains face data positioned by the detection box; determining first position information and second position information of the key point of the face data in the target frame image, the first position information being determined according to a target detection image of the target frame image, and the second position information being determined by tracking and predicting the position of the key point of the face data in a historical frame image, the historical frame image including at least one frame image before the target frame image; determining an offset between the first position information and the second position information; in a case where the offset exceeds an offset threshold, determining attribute information of the key point of the face data in the target frame image according to the first position information; in a case where the offset does not exceed the offset threshold, determining the attribute information of the key point of the face data in the target frame image according to the second position information; performing special effect processing on the frame image in the video data according to the attribute information of the key point of the face data, the special effect processing including at least one of beautification processing, special effect adding processing, makeup processing, and beautifying processing.

11. A data processing method, characterized by, comprising: obtaining live video data, and identifying a frame image in the live video data to extract a detection image corresponding to a detection frame, the detection image containing face data positioned by the detection frame; determining first position information and second position information of a key point of the face data in a target frame image, the first position information being determined according to a target detection image of the target frame image, and the second position information being determined by tracking and predicting the position of the key point of the face data in a historical frame image, the historical frame image including at least one frame image before the target frame image; determining an offset between the first position information and the second position information; in a case where the offset exceeds an offset threshold, determining attribute information of the key point of the face data in the target frame image according to the first position information; in a case where the offset does not exceed the offset threshold, determining the attribute information of the key point of the face data in the target frame image according to the second position information; performing special effect processing on the frame image in the live video data according to the attribute information of the key point of the face data, the special effect processing including at least one of beautification processing, special effect adding processing, makeup processing, and beautifying processing.

12. A data processing method, characterized by, comprising: providing a first interface to obtain related video data through the first interface and determine a frame image in the video data; identifying the frame image to extract a detection image corresponding to a detection frame, the detection image containing a target object positioned by the detection frame; determining first position information and second position information of a key point of the target object in a target frame image, the first position information being determined according to a target detection image of the target frame image, and the second position information being determined by tracking and predicting the position of the key point of the target object in a historical frame image, the historical frame image including at least one frame image before the target frame image; determining an offset between the first position information and the second position information; in a case where the offset exceeds an offset threshold, determining attribute information of the key point of the target object in the target frame image according to the first position information; In a case where the offset does not exceed the offset threshold, attribute information of the key point of the target object in the target frame image is determined according to the second position information. The attribute information is fed back through the second interface.

13. A data processing apparatus, characterized by: The method comprises: The video data acquisition module is configured to acquire video data and determine a frame image in the video data. The detection image acquisition module is configured to identify the frame image, extract a detection image corresponding to a detection frame, and the detection image contains a target object positioned by the detection frame. The position information acquisition module is configured to determine first position information and second position information of a key point of a target object in a target frame image, the first position information is determined according to a target detection image of the target frame image, and the second position information is determined by tracking and predicting a position of the key point of the target object in a historical frame image, the historical frame image includes at least one frame of frame image before the target frame image. The attribute information acquisition module is configured to determine an offset between the first position information and the second position information, and in a case where the offset exceeds an offset threshold, attribute information of the key point of the target object in the target frame image is determined according to the first position information. In a case where the offset does not exceed the offset threshold, attribute information of the key point of the target object in the target frame image is determined according to the second position information.

14. A data processing apparatus, characterized by The method comprises: The model single-frame training module is configured to perform single-frame image training on the key point recognition model by using a training input image and a training input label. The training label acquisition module is configured to input a training image in training video data into the key point recognition model that has completed single-frame image training, to obtain a training label representing a position of a key point of a target object in the training image. The prediction label acquisition module is configured to track and predict a historical training image of a target training image, to determine a prediction training label of the target training image at a target time, the historical training image includes at least one frame of training image before the target training image. The training label verification module is configured to verify a credibility of the training label of the target training image according to a difference between the prediction training label and the training label of the target training image at the target time. The model secondary training module is configured to train the key point recognition model that has completed single-frame image training by using the training label that passes the verification and a corresponding training image.

15. An electronic device, comprising: The method comprises: A processor; and A memory having executable code stored thereon, when the executable code is executed, causing the processor to perform the method of one or more of claims 1-12.

16. One or more machine-readable media having stored thereon executable code that, when executed, cause a processor to perform the method of one or more of claims 1-12.

Citation Information

Patent Citations

  • method and device for generating information

    CN109829432A

  • Behavior action recognition method and device

    CN111783515A