A method, device and electronic device for human skeleton annotation
By obtaining the key points information of human body images and their skeletons, and using preset detection models to automatically mark and correct them, the problem of time-consuming, labor-intensive and accurate marking of human body skeletons is solved, achieving a fast and accurate marking effect.
Patent Information
- Application Number
- CN202210141107.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-16
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2042-02-16
AI Technical Summary
In the prior art, the key point labeling method of human skeleton is time-consuming and labor-intensive, has high cost and high error rate. The automatic labeling method has poor effect on labeling complex actions, and its accuracy depends on the detection model.
By obtaining the key points information of human body images and their skeletons, using preset detection models for automatic labeling, combining the labeling tool to perform missed and mis-checking judgments on the initial labeling image, and correcting it to achieve fast and accurate labeling.
It improves the efficiency and accuracy of human skeleton labeling, reduces the time and cost of manual labeling, and obtains a data set with higher confidence.
Smart Images

Figure CN114519804B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of deep learning technology, and in particular, to a human skeleton annotation method, device, and electronic device. Background Art
[0002] Human skeleton key points include joints, facial features, etc. Describing human bone information through these key points is crucial for describing human postures and predicting human behaviors. In recent years, great progress has been made in the detection schemes of human skeleton key points. Especially with the rise of deep learning, for example, the proposed detection models such as openpose and alphapose, the application of human skeleton key point detection schemes in actual scenarios has become more and more extensive, such as entertainment fitness, rehabilitation training, action recognition, etc.
[0003] However, training a detection model for human skeleton key points requires a large number of images annotated with human skeleton key points, and the annotation of human skeleton key points is a very cumbersome, delicate, and time-consuming task.
[0004] Currently, the annotation methods for human skeleton key points are mainly divided into two types: one is the fully manual annotation method, where annotators annotate each human bone key point in a large number of images to be annotated one by one; the other is the fully automatic annotation method, which usually uses a ready-made human skeleton key point detection model with relatively high accuracy to detect images and obtain the skeleton key points.
[0005] However, for the former method, it inevitably requires a large amount of human and time costs, has a relatively high error rate, and reduces the efficiency of dataset optimization and model establishment; for the latter method, the annotation effect for complex actions is poor, and the accuracy of the annotated data completely depends on the used skeleton key point detection model. Summary of the Invention
[0006] In view of this, embodiments of this application provide a human skeleton annotation method, device, and electronic device, which can solve one or more technical problems in the related art.
[0007] In a first aspect, an embodiment of the present application provides a method for annotating a human skeleton, including: obtaining a current-frame human body image to be annotated and its corresponding current-frame skeleton key point information, and saving the image information of the current-frame human body image and the current-frame skeleton key point information as a current-frame annotation file, where the skeleton key point information includes the type information and position information of the skeleton key points; using the current-frame annotation file to annotate the current-frame human body image to obtain an initial annotation image, and determining whether there is any missed detection or false detection in the initial annotation image to obtain a determination result; and correcting the skeleton key points with missed detection or false detection in the initial annotation image corresponding to the current-frame human body image according to the determination result to obtain a current-frame annotation image.
[0008] In a second aspect, an embodiment of the present application provides a human skeleton annotation device, including: an acquisition module, configured to obtain a current-frame human body image to be annotated and its corresponding current-frame skeleton key point information, and save the image information of the current-frame human body image and the current-frame skeleton key point information as a current-frame annotation file, where the skeleton key point information includes the type information and position information of the skeleton key points; a determination module, configured to use the current-frame annotation file to annotate the current-frame human body image to obtain an initial annotation image, and determine whether there is any missed detection or false detection in the initial annotation image to obtain a determination result; and a correction module, configured to correct the skeleton key points with missed detection or false detection in the initial annotation image corresponding to the current-frame human body image according to the determination result to obtain a current-frame annotation image.
[0009] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where when the processor executes the computer program, the steps of the human skeleton annotation method according to any embodiment of the first aspect are implemented.
[0010] In a fourth aspect, an embodiment of the present application provides a computer storage medium, where the computer storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the human skeleton annotation method according to any embodiment of the first aspect are implemented.
[0011] In a fifth aspect, an embodiment of the present application provides a computer program product, which when running on an electronic device enables the electronic device to implement the steps of the human skeleton annotation method according to any embodiment of the first aspect.
[0012] Embodiments of the present application can quickly and accurately obtain a human skeleton annotation image.
[0013] In some embodiments, a preset skeleton detection model is used to detect a human body image, obtaining an automatic annotation result, reducing manual annotation, and greatly improving the annotation efficiency.
[0014] In some embodiments, the automatic annotation result is visualized, facilitating the re-inspection of the automatic annotation result, improving the re-inspection efficiency, and also improving the annotation accuracy, obtaining a dataset with higher confidence. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0016] Figure 1 is a schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0017] Figure 2A is a schematic flowchart of the implementation of a human skeleton annotation method provided by an embodiment of the present application;
[0018] Figure 2B is a schematic flowchart of the implementation of another human skeleton annotation method provided by an embodiment of the present application;
[0019] Figure 3 is a schematic diagram of the implementation process of step S100 in a human skeleton annotation method provided by an embodiment of the present application;
[0020] Figure 4 is a schematic diagram of human skeleton key points provided by an embodiment of the present application;
[0021] Figure 5 is a schematic diagram of a display interface including an initial annotation image provided by an embodiment of the present application;
[0022] Figure 6 is a schematic structural diagram of a human skeleton annotation device provided by an embodiment of the present application;
[0023] Figure 7 is a schematic structural diagram of another human skeleton annotation device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0025] The term "and / or" used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0026] The description of "an embodiment" or "some embodiments" in the specification of the present application means that specific features, structures, or characteristics described in connection with the embodiment are included in one or more embodiments of the present application. Thus, statements such as "in an embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc., which appear in different places in this specification, do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.
[0027] In addition, in the description of the present application, the meaning of "a plurality of" is two or more.
[0028] It should also be understood that unless otherwise clearly specified or limited, the term "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be directly connected, or indirectly connected through an intermediate medium, and can be the communication inside two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meaning of the above terms in the present application can be understood according to specific situations.
[0029] Currently, the method of manually annotating human skeletons is time-consuming, laborious, costly, inefficient, and has a high error rate. While the method of fully automated human skeleton annotation has poor annotation effects for complex actions, and the accuracy of the annotated data completely depends on the used skeleton key point detection model.
[0030] Therefore, the embodiments of the present application provide a human skeleton annotation method, which can achieve fast and accurate annotation of human skeletons in human images, thereby promoting the development and application of related technologies for skeleton detection.
[0031] In order to illustrate the technical solutions described in the present application, the following will be described through specific embodiments.
[0032] Figure 1 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device includes, but is not limited to, a computer, a tablet, a laptop, a netbook, a server, etc. The embodiments of the present application do not impose any restrictions on the specific type of the electronic device.
[0033] In some embodiments of the present application, as Figure 1 shown, the electronic device may include one or more processors 10 ( Figure 1 only one is shown in the figure), a memory 11, and a computer program 12 stored in the memory 11 and executable on one or more processors 10. For example, a program for performing human skeleton annotation on a human body image. When one or more processors 10 execute the computer program 12, each step in the following embodiments of the human skeleton annotation method can be implemented. Alternatively, when one or more processors 10 execute the computer program 12, the functions of each module / unit in the following embodiments of the human skeleton annotation device can be implemented.
[0034] Exemplarily, the computer program 12 may be divided into one or more modules / units. One or more modules / units are stored in the memory 11 and executed by the processor 10 to complete the present application. One or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 12 in the processing unit. For example, the computer program 12 may be divided into the following several modules. The specific functions of each module are as follows:
[0035] An acquisition module, configured to acquire the current frame human body image to be annotated and its corresponding current frame skeleton key point information, and save the image information of the current frame human body image and the current frame skeleton key point information as a current frame annotation file, where the skeleton key point information includes the type information and position information of the skeleton key points;
[0036] A judgment module, configured to perform annotation on the current frame human body image by using the current frame annotation file to obtain an initial annotation image, and judge whether there is a missed detection or a false detection in the initial annotation image to obtain a judgment result;
[0037] A correction module, configured to correct the skeleton key points with missed detection or false detection in the initial annotation image corresponding to the current frame human body image according to the judgment result to obtain a current frame annotation image.
[0038] Those skilled in the art can understand that Figure 1 merely examples of the electronic device do not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the electronic device may further include input / output devices, network access devices, a bus, etc.
[0039] The so-called processor 10 may be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.
[0040] The memory 11 may be an internal storage unit of the processing unit, such as the hard disk or memory of the processing unit. The memory 11 may also be an external storage device of the processing unit, such as a plug-in hard disk equipped on the processing unit, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 11 may also include both the internal storage unit of the processing unit and the external storage device. The memory 11 is used to store computer programs and other programs and data required by the processing unit. The memory 11 may also be used to temporarily store the data that has been output or will be output.
[0041] Another preferred embodiment of the electronic device is also provided in an embodiment of the present application. In this embodiment, the electronic device includes one or more processors, and the one or more processors are used to execute the following program modules stored in the memory:
[0042] An acquisition module, configured to acquire the current frame human body image to be labeled and its corresponding current frame skeleton key point information, and save the image information of the current frame human body image and the current frame skeleton key point information as the current frame annotation file, where the skeleton key point information includes the type information and position information of the skeleton key points;
[0043] A judgment module, configured to label the current frame human body image with the current frame annotation file to obtain an initial annotation image, and judge whether there is a missed detection or a false detection in the initial annotation image to obtain a judgment result;
[0044] A correction module, configured to correct the skeleton key points with missed detections or false detections in the initial annotation image corresponding to the current frame human body image according to the judgment result to obtain the current frame annotation image.
[0045] The embodiment of the present application also provides a human skeleton annotation method. The human skeleton annotation method in the embodiment of the present application is applicable to the situation where key points of the skeleton need to be annotated for a human body image. The human skeleton annotation method in the embodiment of the present application can be executed by an electronic device. As an example but not a limitation, the human skeleton annotation method can be executed by the electronic device in the Figure 1 embodiment shown.
[0046] Figure 2A FIG. 6 is a schematic flowchart of the implementation of a human skeleton annotation method provided by an embodiment of the present application. As shown in Figure 2A FIG. 7, the human skeleton annotation method may include: step S110 to step S130.
[0047] S110, obtain the current frame human body image to be annotated and its corresponding current frame skeleton key point information, and save the image information of the current frame human body image and the current frame skeleton key point information as the current frame annotation file, where the skeleton key point information includes the type information and position information of the skeleton key points.
[0048] Specifically, the current frame human body image can be a single-frame human body image or an image frame in a video. The video can be a multi-frame continuous image of different actions of a human body captured by a camera or other imaging device. It should be noted that the human body image to be annotated can be a color image, a gray image, an infrared image, etc., and there is no limitation here.
[0049] After the human skeleton annotation of the current frame human body image is completed according to the technical solution of the present application, the next frame human body image is obtained for human skeleton annotation, and at this time the next frame human body image is called the current frame human body image. It should be understood that in the embodiment of the present application, an exemplary description is made for the complete annotation process of the current frame human body image.
[0050] In some embodiments of the present application, each frame of the human body image to be annotated has corresponding skeleton key point information. Before step S110, as shown in Figure 2B FIG. 8, it further includes step S100, using a preset skeleton detection model to detect the human body image to obtain the skeleton key point information of the human body image.
[0051] More specifically, the skeleton key point information corresponding to each frame of the human body image to be annotated is obtained by using a preset skeleton detection model to detect the human body image of this frame. Skeleton detection is to detect the human body and the corresponding skeleton key point information from the input image. The skeleton key point information includes the type information and position information of the skeleton key points. It should be understood that the human body image to be annotated may also include a large number of different human bodies with different actions, and the preset skeleton detection model can also detect the skeleton key point information of different human bodies and assign different indexes to different human bodies for distinction.
[0052] In some possible implementations, the skeletal key points of each human body may include 14 types, namely: right shoulder, right elbow, right wrist, left shoulder, left elbow, left wrist, right hip (or right buttock), right knee, right ankle, left hip (or left buttock), left knee, left ankle, top of the head, and neck.
[0053] In some other possible implementations, the skeletal key points of each human body may include 17 types, namely: nose, right eye, left eye, right ear, left ear, right shoulder, left shoulder, right elbow, left elbow, right wrist, left wrist, right hip, left hip, right knee, left knee, right ankle, and left ankle.
[0054] In some other possible implementations, the skeletal key points of each human body may include 19 types, namely: nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, right ankle, left toe tip, and right toe tip.
[0055] In some other possible implementations, each human body's skeletal key points may include 18, as Figure 3 shown, through a skeletal detection model for human body skeletal detection, the type information of the human body skeletal key points and the position information of the human body skeletal key points (i.e., coordinates, the coordinate information is not shown) are obtained. Among them, the type information of the human body skeletal key points is represented as 18 joints, which can be represented by corresponding labels in sequence as 0 and integers from 1 to 17.
[0056] In some possible implementations, when there are multiple human bodies in a human body image, when obtaining the skeletal key point information of each human body, the skeletal key point label corresponding to each human body also includes the index corresponding to each human body. For example, 1-0, 1-1, ···, 1-17 represent the human body skeletal key points of one user, and 2-0, 2-1, ···, 2-17 represent the human body skeletal key points of another user, and so on. By assigning an index corresponding to each human body to distinguish the key points of different human bodies.
[0057] In the embodiments of the present application, the preset skeletal detection model can be deployed in advance in the memory of the electronic device and can be called when in use. Using the preset skeletal detection model to detect a human body image, that is, the process of automatically annotating the human body image. The preset skeletal detection model can be a trained skeletal detection model, and these models can be models obtained through deep learning training based on a publicly available image dataset. Preferably, it is a skeletal detection model with higher accuracy such as open pose or alpha pose.
[0058] It should be noted that, in addition to the above-mentioned model, any other skeleton detection model can be selected in this application to perform skeleton key point detection on the human body image to be labeled, so as to obtain the skeleton key point information corresponding to each frame of the human body image. The embodiments of this application do not specifically limit the skeleton detection model.
[0059] In a possible implementation manner, when using a preset skeleton detection model to perform human body skeleton detection, there may be a situation where the skeleton key points are not detected. In response to the above situation, on the basis of using labels to represent the skeleton key points, the position information of the undetected skeleton key points is initialized. In one embodiment, the position information of the undetected skeleton key points can be initialized with the position coordinates (0, 0); in another embodiment, the position information of the undetected skeleton key points can also be initialized with the position coordinates (B*n, 0), that is, the undetected skeleton key points of the current frame are initialized so that the undetected skeleton key points of the current frame are spaced apart. Through this setting, compared with the point stacking caused by initializing all undetected skeleton key points to (0, 0), spacing the undetected skeleton key points avoids point stacking, reduces the number of times of selecting the corresponding skeleton key points of the corresponding part during annotation in the subsequent process of correcting the missed detected skeleton key points, thereby reducing the number of times the user manually moves the skeleton key points, reducing the operation complexity, and improving the annotation efficiency.
[0060] More specifically, initializing the undetected skeleton key points of the current frame so that the undetected skeleton key points of the current frame are spaced apart includes: initializing the coordinates of the undetected skeleton key points of the current frame to (B*n, 0), where n is the sorting number of the undetected skeleton key points, and B is a preset distance interval. The sorting number n can adopt the serial number order of the undetected skeleton key points among all human body skeleton key points, such as an integer from 1 to 14; or, the sorting number n can adopt the quantity numbering of the undetected skeleton key points. For example, if there are 4 undetected skeleton key points, the initialization coordinates of the nth undetected skeleton key point are (B*n, 0), and the value of n is an integer from 1 to 4. The preset distance interval B can take an empirical value, and the embodiments of this application do not limit this.
[0061] In some embodiments, the image information of the human body image and the skeleton key point information corresponding to the human body image are saved as corresponding annotation files, so that when it is necessary to perform annotation on a certain frame of the human body image in subsequent steps, the annotation file corresponding to the frame of the human body image can be found through the image information of the frame of the human body image, and the corresponding skeleton key point information can be obtained.
[0062] In one embodiment, the image information of the human body image is the image name of the human body image. The image name of the human body image and the skeleton key point information corresponding to the human body image are written into the annotation file. Optionally, the annotation file can be in txt format. It should be understood that the type of the annotation file is not specifically limited in the embodiments of the present application.
[0063] In a possible implementation manner, for each frame of human body image, the image name of each frame of human body image and the skeleton key point information corresponding to the frame of human body image are saved in a txt document in a preset order. For example, first record the image name "imagename" of the human body image, and then sequentially record the type information of the skeleton key points of the frame of human body image in the order of numbers. For example, taking 14 skeleton key points as an example, the skeleton key point information is sequentially: 1 right shoulder, 2 right elbow, 3 right wrist, 4 left shoulder, 5 left elbow, 6 left wrist, 7 right hip, 8 right knee, 9 right ankle, 10 left hip, 11 left knee, 12 left ankle, 13 top of the head, and 14 neck. By arranging different skeleton key points in the order of numbers as described above, it is convenient to distinguish the types of skeleton key points; and when correcting the skeleton key points subsequently, the position corresponding to them on the human body skeleton can also be determined according to the order of the skeleton key points.
[0064] S120, use the current frame annotation file to annotate the current frame of human body image to obtain an initial annotated image, and determine whether there is a missed detection or a false detection in the initial annotated image to obtain a judgment result.
[0065] In some embodiments, use an annotation tool to display the current frame of human body image, read the current frame annotation file according to the image information of the current frame of human body image, and annotate the current frame of skeleton key points in the current frame of human body image according to the skeleton key point information included in the current frame annotation file to obtain an initial annotated image.
[0066] In a possible implementation manner, different types of skeleton key points can be annotated with different colors. As Figure 4 shown, it is a schematic diagram of a human body image annotated with skeleton key points, that is, an initial annotated image. In the Figure 4 shown example, taking 19 human body skeleton key points as an example, the 19 skeleton key points are annotated with different colors. In another possible implementation manner, different parts of the skeleton key points can be annotated with different colors, such as parts, for example, right hand, left hand, right leg, left leg, right foot, left foot, human head, etc.
[0067] In another possible implementation, for the case of multi-person detection, according to the indexes corresponding to different human bodies in the annotation file, different colors are used to label the skeleton key points of different human bodies. In some embodiments of the present application, by using different colors to label the skeleton key points of different types and / or different parts and / or different human bodies, it is more convenient for users to distinguish each skeleton key point and improve the accuracy of re-inspection. More generally, in some embodiments of the present application, the skeleton key points of different types and / or different parts and / or different human bodies are distinguished and labeled, and the styles of the distinguished labeling include but are not limited to different colors, different labeling patterns, different line thicknesses, etc.
[0068] It should be noted that after the automatic annotation of the human body image is completed through the preset skeleton detection model, the automatic annotation result may have inaccurate skeleton key point information, such as the position information of the skeleton key points, etc., that is, misdetection (obvious deviation from the original human body skeleton) or missed detection (that is, undetected points, there are overlapping skeleton points or the skeleton key points are distributed at preset intervals). Therefore, it is necessary to judge whether there are misdetected or missed detected skeleton key points in the initial annotation image and obtain the corresponding judgment result.
[0069] In one embodiment, if it is determined that there is no missed detection or misdetection of the skeleton key points of the current frame in the initial annotation image, then there is no need to correct the skeleton key points of the current frame. At this time, the initial annotation image of the current frame is defaulted to the annotation image of the current frame, and the above steps are executed for the next frame image; if it is determined that there is a missed detection or misdetection of the skeleton key points of the current frame in the initial annotation image, then it is necessary to correct the skeleton key points of the current frame, that is, step S130.
[0070] S130, according to the judgment result, correct the misdetected or missed detected skeleton key points in the initial annotation image corresponding to the current frame human body image to obtain the annotation image of the current frame.
[0071] In some embodiments of the present application, the annotation tool can be used to correct the missed detection or misdetection in the automatic annotation result, so as to improve the efficiency and accuracy of the annotation data. Specifically, the initial annotation image is visualized by using the annotation tool, which is convenient for users (i.e., annotators) to perform manual re-inspection. Users can judge the missed detection or misdetection in the automatic annotation result to improve the efficiency and accuracy of the re-inspection, so as to quickly obtain accurate annotation data.
[0072] In some embodiments, correcting the misdetected or missed detected skeleton key points in the initial annotation image corresponding to the current frame human body image according to the judgment result includes: for the missed detected skeleton key points of the current frame in the initial annotation image, relabel the missed detected skeleton key points of the current frame; for the misdetected skeleton key points of the current frame in the initial annotation image, adjust the misdetected skeleton key points of the current frame.
[0073] In a possible implementation, when the user discovers that there are undetected or misdetected current-frame skeleton key points in the initial annotated image, the user can identify the undetected or misdetected current-frame skeleton key points through user operations. Then, the user can input manual recheck data through an input device such as a mouse and / or keyboard of the electronic device, and the annotation tool corrects the undetected or misdetected skeleton key points according to the manual recheck data input by the user, so as to obtain a corrected annotated image. For example, the user can add the undetected key points by clicking the mouse, or use the drag-and-drop method to relabel the misdetected key points.
[0074] In another possible implementation, when the user discovers that there are undetected or misdetected current-frame skeleton key points in the initial annotated image, the user can identify the undetected or misdetected current-frame skeleton key points through user operations. Then, the electronic device determines whether the current-frame human body image and the historical-frame human body image, such as the previous-frame human body image, are consecutive-frame images, and the electronic device performs corresponding operations according to the determination result. Specifically, if it is determined that they are consecutive-frame images, the undetected or misdetected current-frame skeleton key points inherit the corresponding skeleton key point information in the historical-frame human body image. For example, if it is determined according to the identification input by the user that there is an undetected or misdetected right-wrist key point in the current-frame skeleton key points, the right-wrist key point in the current-frame human body image directly adopts the right-wrist key point information in the historical-frame human body image. If it is determined that they are not consecutive-frame images, the user can input manual recheck data through an input device such as a mouse and / or keyboard of the electronic device, and the annotation tool relabels or adjusts the undetected or misdetected skeleton key points according to the manual recheck data input by the user. On the basis of this implementation, when it is determined that they are consecutive-frame images and the undetected or misdetected current-frame skeleton key points inherit the corresponding skeleton key point information in the historical-frame human body image, the user can further check whether there are errors in the inheritance result. If there are errors, the user can correct the incorrect inheritance result through user operations to obtain a more accurate annotated image.
[0075] In another possible implementation, for misdetected skeleton key points, the user can first delete them using the one-key deletion method and then relabel them to the correct position.
[0076] In some embodiments, the electronic device determines whether the current-frame human body image and the historical-frame human body image are consecutive-frame images by determining whether the similarity between the current-frame human body image and the historical-frame human body image is greater than a threshold. If it is greater than the threshold, it is determined that the current-frame human body image and the historical-frame human body image are consecutive frames. If it is less than the threshold, it is determined that the current-frame human body image and the historical-frame human body image are not consecutive frames. The threshold can be an empirical value. When the similarity is equal to the threshold, it can be set that the current-frame human body image and the historical-frame human body image are consecutive frames or not consecutive frames, and it can be selectively set according to requirements.
[0077] As a possible implementation, calculate a first similarity between a preset number of current-frame skeleton key points in the current-frame human body image and the corresponding historical-frame skeleton key points in the historical-frame human body image. If the first similarities of the preset number meet the preset conditions, it is determined that the similarity between the current-frame human body image and the historical-frame human body image is greater than the threshold, that is, the two are consecutive frames; if the first similarities of the preset number do not meet the preset threshold, it is determined that the similarity between the current-frame human body image and the historical-frame human body image is less than the threshold, that is, the two are not consecutive frames. Among them, the preset number is the number of skeleton key points included in each human body in the automatic annotation result output by the skeleton detection model, and can be, for example, 14, 17, 18, or 19, etc. The preset conditions can be set as: the first similarities of a preset proportion in the preset number are greater than the preset threshold; or, the first similarities of a preset number in the preset number are greater than the preset threshold. The preset proportion can take an empirical value, for example, any value between 50% and 85%. The preset number can take an empirical value and can take any value less than the preset number. For example, when the preset number is 19, the preset number can be 15.
[0078] As a non-limiting example, the Euclidean distance is used to calculate the first similarity between any current-frame skeleton key point in the current-frame human body image and the corresponding historical-frame skeleton key point in the historical-frame human body image. Specifically, the Euclidean distance d = sqrt((x1 - x2)*(x1 - x2)+(y1 - y2)*(y1 - y2)), where the coordinates of the current-frame skeleton key point are (x1, y1), and the coordinates of the corresponding historical-frame skeleton key point are (x2, y2).
[0079] In some other embodiments, for the electronic device to determine whether the current-frame human body image and the historical-frame human body image are consecutive frame images, it can be achieved by determining whether the image names of the current-frame human body image and the historical-frame human body image are consecutively numbered. Usually, consecutive frame images usually use sequential numbers. Therefore, in this implementation, the determination is quickly completed in a simple way, further improving the image annotation efficiency. Specifically, when it is determined that the image names of the current-frame human body image and the historical-frame human body image are consecutively numbered, it is determined that the current-frame human body image and the historical-frame human body image are consecutive frame images; otherwise, they are not consecutive frame images.
[0080] It should be noted that if it is found through rechecking by the annotation tool and / or manual inspection by the user that there is no missing or incorrect annotation of the skeleton key point information, the original skeleton key point information is retained. Specifically, for the uncorrected skeleton key points, the original skeleton key point information is retained; for the corrected skeleton key points, the corrected skeleton key point information is retained.
[0081] In some embodiments, the initial labeled image and its skeleton key point types can be displayed on the same screen. In addition, the distinguishable labeling styles of the skeleton key points can also be displayed, as well as whether the skeleton key points are corrected, with different labeling styles before and after correction. For example, Figure 5 as shown, the initial labeled image, the labeling colors of each of the 19 skeleton key points, and the labeling styles before and after correction: square or circular are displayed simultaneously.
[0082] In some embodiments of the present application, a preset skeleton detection model is used to detect a human body image to obtain an automatic labeling result, reducing manual labeling and greatly improving the labeling efficiency. In some embodiments of the present application, the automatic labeling result is visualized, facilitating the re-inspection of the automatic labeling result, improving the re-inspection efficiency, and also improving the labeling accuracy, obtaining a dataset with higher confidence.
[0083] It should be noted that the step numbers in the embodiments cannot be construed as a limitation on the time sequence of each step. It should be understood that in some other embodiments, the front-back order between steps can be adjusted according to the logical relationship between steps without affecting the implementation of the present solution.
[0084] Corresponding to the above-mentioned human skeleton labeling method, an embodiment of the present application also provides a human skeleton labeling device. For details not described in detail in this human skeleton labeling device, please refer to the relevant descriptions of the foregoing method, and will not be elaborated here.
[0085] Figure 6 is a schematic structural diagram of a human skeleton labeling device provided by an embodiment of the present application. As an example, the human skeleton labeling device can be configured in Figure 1 the electronic device shown. The human skeleton labeling device includes: an acquisition module 61, a judgment module 62, and a correction module 63.
[0086] Among them, the acquisition module 61 acquires the current frame human body image to be labeled and its corresponding current frame skeleton key point information, and saves the image information of the current frame human body image and the current frame skeleton key point information as the current frame labeling file, where the skeleton key point information includes the type information and position information of the skeleton key points;
[0087] The judgment module 62 is used to label the current frame human body image with the current frame labeling file to obtain an initial labeled image, and judge whether there is a missed detection or false detection in the initial labeled image to obtain a judgment result;
[0088] The correction module 63 is used to correct the skeleton key points with missed detection or false detection in the initial labeled image corresponding to the current frame human body image according to the judgment result to obtain the current frame labeled image.
[0089] Figure 7It is a schematic structural diagram of a human skeleton annotation device provided by another embodiment of the present application. The human skeleton annotation device includes: an annotation module 60, an acquisition module 61, a judgment module 62, and a correction module 63. It should be noted that Figure 7 the illustrated embodiment and Figure 6 the same modules in the illustrated embodiment will not be described in detail here.
[0090] Among them, the annotation module 60 is used to detect a human body image by using a preset skeleton detection model to obtain skeleton key point information of the human body image.
[0091] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be described in detail here.
[0092] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments of each human skeleton annotation method can be implemented.
[0093] The embodiment of the present application provides a computer program product. When the computer program product runs on an electronic device, the electronic device can implement the steps in the above-mentioned method embodiments of each human skeleton annotation method.
[0094] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not described in detail or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0095] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0096] In the embodiments provided in the present application, it should be understood that the disclosed apparatus / electronic device and method can be implemented in other ways. For example, the apparatus / electronic device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the apparatus or unit can be in electrical, mechanical or other forms.
[0097] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0098] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0099] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0100] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for human skeleton annotation, characterized in that, Including: Obtain the current-frame human body image to be labeled and its corresponding current-frame skeleton key point information, and save the image information of the current-frame human body image and the current-frame skeleton key point information as a current-frame annotation file, where the skeleton key point information includes the type information and position information of the skeleton key points; Use the current-frame annotation file to label the current-frame human body image to obtain an initial annotation image, and determine whether there is any missed detection or misdetection in the initial annotation image to obtain a judgment result; According to the judgment result, correct the skeleton key points with missed detection or misdetection in the initial annotation image corresponding to the current-frame human body image to obtain a current-frame annotation image; Before the step of obtaining the current-frame human body image to be labeled and its corresponding current-frame skeleton key point information, it further includes: Use a preset skeleton detection model to detect the current-frame human body image to obtain the skeleton key point information of the current-frame human body image; If any skeleton key point in the current-frame human body image is not detected by using the preset skeleton detection model, initialize the position information of the non-detected skeleton key point, where the initialization includes initializing the position information of the non-detected skeleton key point with the available position coordinates (0, 0); or, initialize the non-detected skeleton key point so that the non-detected skeleton key points are distributed at intervals.
2. The human skeleton annotation method according to claim 1, wherein The step of initializing the non-detected skeleton key point so that the non-detected skeleton key points are distributed at intervals includes: initializing the coordinates of the non-detected skeleton key points as (B*n, 0), where n is the sorting number of the non-detected skeleton key points, and B is a preset distance interval.
3. The human skeleton annotation method according to claim 1, wherein, The step of correcting the skeleton key points with missed detection or misdetection in the initial annotation image corresponding to the current-frame human body image according to the judgment result includes: For the current-frame skeleton key points with missed detection in the initial annotation image, relabel the missed-detected current-frame skeleton key points; For the current-frame skeleton key points with misdetection in the initial annotation image, adjust the misdetected current-frame skeleton key points.
4. The human skeleton annotation method according to claim 3, characterized in that, The step of relabeling the missed-detected current-frame skeleton key points includes: Responding to the user's first operation, relabel the missed-detected current-frame skeleton key points; or, If it is determined that the similarity between the current-frame human body image and the historical-frame human body image is greater than a threshold, the missed-detected current-frame skeleton key points inherit the corresponding historical-frame skeleton key point information in the historical-frame human body image; if it is determined that the similarity between the current-frame human body image and the historical-frame human body image is less than the threshold, respond to the user's first operation and relabel the missed-detected skeleton key points; The step of adjusting the misdetected current-frame skeleton key points includes: Responding to the user's second operation, adjust the misdetected skeleton key points; or, If it is determined that the current frame human body image and the historical frame human body image are consecutive frames, the misdetected skeleton key points inherit the corresponding historical frame skeleton key point information in the historical frame human body image; if it is determined that the current frame human body image and the historical frame human body image are not consecutive frames, in response to the user's second operation, the undetected skeleton key points are adjusted.
5. The human skeleton annotation method according to claim 4, wherein The determination that the current frame human body image and the historical frame human body image are consecutive frames includes: Determining that the similarity between the current frame human body image and the historical frame human body image is greater than a threshold; or, determining that the consecutive serial numbers of the image names of the current frame human body image and the historical frame human body image are consecutive.
6. A human skeleton annotation device, characterized in that, It includes: An acquisition module, configured to acquire a current frame human body image to be labeled and its corresponding current frame skeleton key point information, and save the image information of the current frame human body image and the current frame skeleton key point information as a current frame labeling file, where the skeleton key point information includes the type information and position information of the skeleton key points; A judgment module, configured to label the current frame human body image using the current frame labeling file to obtain an initial labeled image, and judge whether there are undetected or misdetected parts in the initial labeled image to obtain a judgment result; A correction module, configured to correct the undetected or misdetected skeleton key points in the initial labeled image corresponding to the current frame human body image according to the judgment result to obtain a current frame labeled image; Before acquiring the current frame human body image to be labeled and its corresponding current frame skeleton key point information, it further includes: Using a preset skeleton detection model to detect the current frame human body image to obtain the skeleton key point information of the current frame human body image; If any skeleton key point in the current frame human body image is not detected by using the preset skeleton detection model, initialize the position information of the undetected any skeleton key point, where the initialization includes that the position information of the undetected any skeleton key point can be initialized with the position coordinates (0, 0); or, initialize the undetected any skeleton key point so that the undetected skeleton key points are distributed at intervals.
7. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, When the processor executes the computer program, it implements the human skeleton labeling method according to any one of claims 1 to 5.
8. A computer storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the human skeleton labeling method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Human body detection method and device
CN110705448A
Data labeling method and device
CN111126157A