Key point detection method and device, terminal equipment and computer readable storage medium
By cropping the original image in the keypoint detection model and directly inputting a local image, the problem of out-of-bounds keypoint detection is solved, achieving high-precision and high-efficiency keypoint detection, which is applicable to technical fields such as multi-person pose estimation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-06
- Publication Date
- 2026-03-10
AI Technical Summary
Existing keypoint detection technologies rely on the accuracy of human bounding boxes in multi-person pose estimation scenarios, which can lead to failure in detecting keypoints that go out of bounds. Furthermore, expanding the bounding box can introduce background noise interference, reducing detection accuracy and efficiency.
A keypoint detection model with the ability to detect out-of-bounds keypoints is adopted. By cropping the original image and directly inputting the local image according to the position of the target detection box, the detection is performed, avoiding the operation of expanding the detection box. The detection accuracy and efficiency are improved by utilizing the model's own detection capabilities.
It effectively detects out-of-bounds key points, reduces computing resource consumption, avoids background noise interference, and ensures high-precision and high-efficiency key point detection.
Smart Images

Figure CN121640002A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of key point detection technology, and in particular relates to a key point detection method, apparatus, terminal equipment and computer-readable storage medium. Background Technology
[0002] In recent years, with the continuous development of artificial intelligence technology, key point detection technology has played an important role in various technical fields such as human pose estimation and tracking, human-computer interaction, behavior recognition, motion analysis, 3D reconstruction, and security monitoring, bringing many conveniences to people's lives and work.
[0003] For example, in the field of human pose estimation, it is often necessary to identify the positional information of various key points of the human body from images or videos, and continuously track these key points in consecutive video frames. Then, the trajectory changes and pose changes of human movement can be obtained based on the positional information of these key points.
[0004] However, in some multi-person pose estimation scenarios, it is often necessary to first use a human detector to determine the bounding box of each human body region, known as a human detection box. Then, the corresponding human body thumbnail is fed into a keypoint detection model to obtain human keypoints. In this case, the accuracy of keypoint detection heavily depends on the accuracy of the human detection box. If the human detector's provided human detection box is inaccurate—for example, if some key parts of the human body do not fall within the detection box—it is very likely that keypoint detection for these key parts will fail; that is, the keypoint detection model will be unable to effectively detect the positions of keypoints that exceed the detection box. Summary of the Invention
[0005] This application provides a key point detection method, apparatus, terminal device, and computer-readable storage medium, which can effectively detect out-of-bounds key points in various scenarios and improve the overall accuracy and efficiency of key point detection.
[0006] The first aspect of this application provides a key point detection method, including:
[0007] Obtain the location of the target detection box in the original image, wherein the original image includes the imaging region of the target object, and the target detection box is the location box of at least a portion of the imaging region of the target object;
[0008] Based on the target detection bounding box, the original image is cropped to obtain the cropped local image;
[0009] The local image is input into the key point detection model to obtain the first detection result, which includes the location information of key points that extend beyond the boundary of the local image.
[0010] Based on the first detection result, the positions of each key point of the target object in the original image are determined.
[0011] A second aspect of this application provides a key point detection device, comprising:
[0012] The acquisition module is used to acquire the position of the target detection box in the original image, wherein the original image includes the imaging region of the target object, and the target detection box is the location box of at least part of the imaging region of the target object.
[0013] The cropping module is used to crop the original image based on the target detection box to obtain the cropped local image;
[0014] The detection module is used to input the local image into the key point detection model to obtain the first detection result, wherein the first detection result includes the position information of key points that extend beyond the boundary of the local image;
[0015] The determination module is used to determine the positions of each key point of the target object in the original image based on the first detection result.
[0016] A third aspect of this application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the steps of the aforementioned key point detection method.
[0017] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the aforementioned key point detection method.
[0018] The keypoint detection method provided in the first aspect of this application firstly crops the original image based on the position of the target detection box in the original image to obtain a cropped local image. Then, the local image is input into a keypoint detection model to obtain a first detection result. The keypoint detection model of this application has the ability to locate keypoints outside the boundary of the local image, so the first detection result includes the position information of keypoints beyond the boundary of the local image. Furthermore, based on the first detection result, the position of each keypoint of the target object in the original image can be accurately determined. This scheme does not require the operation of expanding the original target detection box, thus effectively reducing the computational resources consumed in preprocessing. Moreover, directly feeding the local image obtained by cropping the target detection box into the keypoint detection model for detection can also ensure high detection efficiency. In addition, compared with the prior art where expanding the detection box leads to the introduction of background noise, the scheme of this application relies on the detection capability of the model itself, which can effectively avoid the interference of surrounding noise, thereby ensuring the accuracy of keypoint detection.
[0019] It is understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart illustrating a key point detection method in the prior art;
[0022] Figure 2 This is a schematic flowchart of a key point detection method provided in one embodiment of this application;
[0023] Figure 3 This is a schematic diagram of the structure of a key point detection model provided in one embodiment of this application;
[0024] Figure 4 This is a flowchart illustrating a key point detection method provided in another embodiment of this application;
[0025] Figure 5 This is a schematic diagram of the structure of a key point detection device provided in one embodiment of this application;
[0026] Figure 6 This is a schematic diagram of the structure of a terminal device provided in one embodiment of this application. Detailed Implementation
[0027] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0028] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0029] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0030] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0031] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0032] As mentioned earlier, keypoint detection technology plays an important role in many technical fields. For example, in the field of human pose estimation technology, in order to accurately track changes in the trajectory and posture of human movement, it is necessary to first detect the positional information of various key points of the human body from images or videos, and then continuously track these key points in consecutive video frames.
[0033] There are various methods for keypoint detection. Among them, the two most widely used methods are top-down and bottom-up keypoint detection. Bottom-up keypoint detection does not require object detection; it can directly feed the original image into the keypoint detection model to output all keypoints. However, this approach is not suitable for multi-object scenarios. If there are multiple objects, this method cannot directly distinguish which points belong to which object. Further complex post-processing is required to group the keypoints, and it relies on the high resolution of the original image. Top-down keypoint detection, on the other hand, first performs object detection (e.g., using a human detector to generate a bounding box), extracts the target region, and then feeds it separately into the keypoint detection model, finally outputting the coordinates of the keypoints. While this approach is applicable to both single-person and multi-person scenarios, existing keypoint detection models predict keypoints based on the entire image range of the input image. Therefore, this approach is entirely limited by the results of object detection; inaccurate object detection results can significantly affect the accuracy of keypoint detection. For example, in densely populated scenes, the human body area outlined by the human body detection box is incomplete, and some key parts of the human body do not fall within the detection box, causing existing key point detection models to be unable to effectively detect key points that exceed the boundaries of the detection box.
[0034] Existing technologies typically address this type of problem by expanding the detection frame before keypoint detection. For example... Figure 1 As shown, for example, the original human bounding box output by the human detector is relatively small and does not include the ankle region within the box. The ankle region is a crucial locational basis for detecting ankle keypoints. If the human region within the bounding box is directly cropped into a smaller image and fed into the existing keypoint detection model, it will be difficult to detect the coordinates of all keypoints (especially the ankle points on both sides) because the existing keypoint detection model lacks the ability to detect keypoints that cross the boundary. (It should be noted that...) Figure 1 The first four keypoint heatmaps shown are for illustrative purposes only; in actual applications, these four images represent human scene images without keypoint information. To detect all keypoints, the human bounding box can be expanded to a larger bounding box. The original image can then be cropped along this expanded bounding box to obtain a smaller cropped human image. Next, based on the keypoint detection model's requirements for input image size, the smaller human image can be scaled while maintaining its aspect ratio and its edges padded to obtain an input image that meets the size requirements (e.g., image size W*H). Finally, the input image is fed into the trained keypoint detection model for keypoint detection, and the model outputs heatmaps containing the location information of each keypoint.
[0035] Easy to understand, although Figure 1 The keypoint detection method shown can effectively detect out-of-bounds keypoints in sparsely populated scenarios, but this approach still has several problems. Firstly, in densely populated scenarios, the expanded detection box easily includes non-subject persons, causing keypoints to be interchanged between people, which reduces the accuracy of keypoint detection. In this case, although the keypoint detection model provides complete keypoint location information, due to background noise interference, some out-of-bounds keypoint information may correspond to other people's keypoints, reducing the overall keypoint detection accuracy and affecting the accuracy of subsequent pose estimation, potentially leading to incorrect conclusions. Secondly, if the original human bounding box is too small, the expanded bounding box still cannot cover a relatively complete human body area, so this method still cannot effectively detect all keypoints. Furthermore, existing technologies do not distinguish whether the detection box does not cover the complete human body (i.e., they do not determine whether keypoints are out of bounds) before expanding the detection box, but expand all human bounding boxes indiscriminately. However, if the original human bounding box already contains the complete human body area, further expansion introduces more noise. Especially when the target human body is far from the image acquisition device, expanding the detection bounding box reduces the proportion of the human body region in the cropped human body thumbnail. The size of the human body thumbnail is effective; if the proportion is too small, the resolution of the human body region may be too low, causing the keypoint detection model to be unable to effectively detect the location of keypoints. Therefore, this situation can actually lead to a decrease in the accuracy of keypoint detection.
[0036] To at least partially address the aforementioned technical problems, embodiments of this application provide a keypoint detection method, apparatus, terminal device, and computer-readable storage medium. It is applicable to various fields requiring keypoint detection, including but not limited to human pose estimation and tracking, vehicle recognition and tracking, human-computer interaction, behavior recognition, motion analysis, 3D reconstruction, and security monitoring, and is particularly suitable for multi-target pose estimation and tracking in target-dense scenarios within these fields. This solution does not require expanding the target detection box; instead, it utilizes a keypoint detection model with boundary-crossing keypoint detection capabilities to detect the local image corresponding to the original target detection box. This solution not only effectively solves the problem of predicting boundary-crossing keypoints for multiple targets in target-dense scenarios, but also, compared to existing technologies that expand the detection box in the preprocessing stage, effectively avoids introducing noise interference to the detection, ensuring high detection accuracy and efficiency.
[0037] For simplicity, the following explanation uses key point detection in human pose estimation as an example. For instance, by detecting key points on the human body in surveillance videos of high-risk areas (such as school gates and hospital entrances), and estimating the human pose based on the detected key points, early warnings can be issued for potential security threats based on the estimated pose. For example, when suspicious individuals are detected loitering for an extended period or exhibiting abnormal behavior, an alarm mechanism can be automatically triggered to alert security personnel to take appropriate measures.
[0038] like Figure 2 As shown, the key point detection method provided in this application embodiment includes the following steps S210, S220 and S230.
[0039] Step S210: Obtain the location of the target detection box in the original image. The original image includes the imaging region of the target object, and the target detection box is the location bounding box of at least a portion of the imaging region of the target object.
[0040] In this embodiment of the application, the original image may be an image containing one or more target objects to be detected. The target object may be any suitable object, including but not limited to the human body or parts of the human body (e.g., hands, legs, face, etc.), animals such as cats or dogs, vehicles such as vehicles, airplanes, ships, etc., rigid objects such as robotic arms, robots, etc., specific objects such as parts on a production line, tools, etc., and virtual objects in augmented reality (AR) and virtual reality (VR).
[0041] The original image can be an image acquired using any suitable image acquisition device. It can be a black and white image or a color image. The original image can be an image of any suitable size and resolution. Of course, the original image can also be an image that meets preset requirements. For example, the original image can be a color RGB image that meets the resolution requirements. The original image can be a raw image directly acquired by the image acquisition device, or it can be an image obtained after preset processing of the raw image. For example, the raw image can be cropped to a preset size to obtain the original image. Alternatively, the raw image can be filtered according to a preset filtering method to obtain the original image. Or, the original color image can be processed into grayscale to obtain the image.
[0042] For simplicity, the following text will use an image of the area captured and processed by the surveillance equipment at the school gate as an example.
[0043] For example, the original image processed by the monitoring equipment can be fed into a trained object detection model in real time, and the model outputs object detection boxes for each human body in the original image. The object detection model can be a YOLO (You Only Look Once) series object detection model, an SSD (Single Shot MultiBox Detector) single-object multi-box detection model, Faster R-CNN, etc., and this application does not limit it. The object detection boxes can have any suitable shape. For example, the object detection box is a rectangular outer envelope. For example, if the monitoring area at the school gate includes 5 human bodies, the object detection model can output the positions of 5 rectangular outer envelopes corresponding to these 5 human bodies in the original image. For example, the position of the object detection box can be represented by (x, y, w, h), where (x, y) are the coordinates of the upper left corner of the rectangular outer envelope in the original image, and w and h represent the width and height of the rectangular outer envelope, respectively.
[0044] Step S220: Based on the target detection box, crop the original image to obtain the cropped local image.
[0045] For example, the original image can be cropped along the boundaries of the five rectangular bounding boxes to obtain five cropped images. These five cropped images can then be directly used as input images to the keypoint detection model. Alternatively, due to the requirements of the keypoint detection model on the input image, these five cropped images can be scaled while maintaining their original aspect ratios and then padded to obtain five local images of the same size. The scaling and padding strategies for the cropped images can be arbitrarily set according to actual needs, and this application does not impose any restrictions on them. The same size can be the maximum size limit of the input image for the keypoint detection model, or it can be a size obtained through testing that maximizes the balance between the detection accuracy and efficiency of the keypoint detection model while meeting the model's size requirements.
[0046] Step S230: Input the local image into the keypoint detection model to obtain the first detection result. The first detection result includes the positional information of keypoints that extend beyond the boundaries of the local image.
[0047] In the above example of human pose estimation, the keypoint detection model can be a machine learning model that can detect one or more human keypoints that are expected to be detected, such as the OpenPose model, MoveNet model, PoseNet model, DCPose model, etc., as long as these models have the ability to detect human keypoints.
[0048] It should be noted that the keypoint detection model in this step possesses the capability for out-of-bounds keypoint detection. That is, the model can search and locate the possible locations of keypoints not only within the local image area but also outside the boundaries of the local image. Various suitable methods can be used to endow the keypoint detection model with this capability. In one example, without changing the model structure, the ability to detect out-of-bounds keypoints can be achieved simply by improving the label encoding rules and the size of some network layers. The labels of the training images can be re-encoded. Specifically, the locations of keypoints can be re-encoded, not only labeling the locations of keypoints within the training image area but also labeling the locations of keypoints within the surrounding area of the training image. For example, the coordinates of the labeled keypoints can be based on the top-left corner of the extended area surrounding the training image, with the coordinate range covering both the training image and the surrounding extended area. Correspondingly, the output network layer of the model can be redefined, for example, by increasing the size of the output network layer to achieve additional output of keypoint prediction results within the extended area surrounding the training image. In another example, the structure and training method of an existing keypoint detection model can also be pre-optimized to enable out-of-bounds keypoint detection. For example, more complex network structures (such as deep residual networks, recurrent neural networks, etc.) can be used to enhance the model's representational capabilities. Simultaneously, more advanced training methods (such as transfer learning, adversarial training, etc.) can be employed to improve model performance. In other examples, external data or prior knowledge can be leveraged to enable the model to detect out-of-bounds keypoints. For instance, if it is known that certain keypoints are typically located in specific regions of an image or have specific relative positional relationships, these constraints can be incorporated into the model training.
[0049] For example, the human keypoints to be detected can be the 17 keypoints commonly used in human pose estimation: nose keypoint, left eye keypoint, right eye keypoint, left ear keypoint, right ear keypoint, left shoulder keypoint, right shoulder keypoint, left elbow keypoint, right elbow keypoint, left wrist keypoint, right wrist keypoint, left hip keypoint, right hip keypoint, left knee keypoint, right knee keypoint, left ankle keypoint, and right ankle keypoint. Of course, in examples where the target object is a local part of the human body, the keypoints to be detected can be keypoints specific to that local part. For example, if the target object is the hand, the keypoints to be detected could be the palm keypoints.
[0050] In this embodiment, the first detection result may include the location information of each key point to be detected, output by the model. This location information may be specific coordinates, or it may be the probability value or confidence level of a key point corresponding to a certain coordinate range.
[0051] In a specific example, the five local images obtained in the above steps can be input into a trained keypoint detection model, and the first detection result can be determined based on the model's output. For example, the first detection result can include the coordinates of the aforementioned 17 keypoints. The first detection result can be the model's output or a detection result directly determined based on the model's output. For example, the keypoint detection model outputs the probability value of each of the aforementioned 17 keypoints for each coordinate range. For each keypoint, the range of horizontal and vertical coordinates with the highest probability can be selected to obtain the coordinates of that keypoint, which can then be used as the first detection result.
[0052] Step S240: Based on the first detection result, determine the position of each key point of the target object in the original image.
[0053] It is understandable that the origin of the coordinate system of each keypoint in the first detection result differs from the origin of the original image. Therefore, in this step, the coordinates of each keypoint in the first detection result can be shifted based at least on the difference in the position of the two coordinate origins to obtain the coordinates of each keypoint of the target object in the original image. Furthermore, in the example above where a local image obtained by scaling and padding a cropped small image while maintaining its original aspect ratio is used as input to the keypoint detection model, the coordinates of each keypoint in the first detection result can also be scaled and shifted to accurately determine the position of each keypoint of the target object in the original image.
[0054] The keypoint detection method provided in this application first crops the original image based on the position of the target detection box in the original image to obtain a cropped local image. Then, the local image is input into the keypoint detection model to obtain a first detection result. The keypoint detection model in this application has the ability to locate keypoints outside the boundary of the local image, so the first detection result includes the position information of keypoints beyond the boundary of the local image. Furthermore, based on the first detection result, the position of each keypoint of the target object in the original image can be accurately determined. This scheme does not require the operation of expanding the original target detection box, thus effectively reducing the computational resources consumed in preprocessing. Moreover, directly feeding the local image obtained by cropping the target detection box into the keypoint detection model for detection can also ensure high detection efficiency and meet the needs of real-time processing. In addition, compared with the prior art where expanding the detection box leads to the introduction of background noise, the scheme in this application relies on the detection capability of the model itself, which can effectively avoid the interference of surrounding noise, thereby ensuring the accuracy of keypoint detection.
[0055] In one implementation, step S230 inputs the local image into the keypoint detection model to obtain a first detection result, including step S231. Step S231 involves inputting the local image into the keypoint detection model to determine the first coordinates of each keypoint in a first coordinate system. The first coordinate system has its origin at the top left corner of a first region, and the first region is obtained by expanding the local image to the corresponding dimensions in multiple surrounding directions.
[0056] Step S240 determines the location of key points of the target object in the original image based on the first detection result, including step S241. Step S241 involves transforming the first coordinates of each key point based at least on the size corresponding to multiple expansion directions and the position of the target detection box, to obtain the second coordinates of each key point in a second coordinate system. The second coordinate system uses the top-left corner of the original image as its origin.
[0057] The expansion direction can be arbitrarily set according to actual detection needs. For example, multiple expansion directions can include the image width direction (e.g., from left to right and / or from right to left) and the image height direction (e.g., from top to bottom and / or from bottom to top). The dimensions corresponding to different expansion directions can be the same or different. The dimensions corresponding to different expansion directions can be set to appropriate values according to actual needs. By performing keypoint detection on a local image and mapping it to the original image coordinate system, a flexible, accurate, and efficient keypoint localization method is provided.
[0058] In one implementation, the multiple extension directions include the image width direction and the image height direction, and the method further includes steps S201 and S202.
[0059] Step S201: Calculate the product of the first width and the first multiple to obtain the first size corresponding to the image width direction. Here, the first width is equal to the image width of each training thumbnail and the local image. The training thumbnail can be an image used to train the keypoint detection model.
[0060] Step S202: Calculate the product of the first height and the second multiple to obtain the second dimension corresponding to the image height direction. The first height is equal to the image height of each training thumbnail and local image. The first and second multiples can be expansion multiples in different directions and can be arbitrarily set according to actual needs. For example, for the case where the target detection box is a human detection box, the first and second multiples can be less than or equal to 1. Specifically, the first and second multiples can be the same or different.
[0061] In a specific example, such as Figure 4As shown, the first width and first height of the local image input to the keypoint detection model can be W and H, respectively. The first and second multiples can both be 1 / 2. Multiple expansion directions can include the top, bottom, left, and right directions of the local image. Specifically, the dimensions corresponding to the upward and downward expansion directions (i.e., the image height direction) can both be H / 2, i.e., the second dimension, and the dimensions corresponding to the left and right expansion directions (i.e., the image width direction) can both be W / 2. The first region can correspond to... Figure 4 The black rectangular region in the keypoint detection results. The first coordinate system can be defined with the top-left corner of the black rectangle as the origin, the positive x-axis pointing rightward along the image width, and the positive y-axis pointing downwards along the image height. Figure 4 In the example shown, in the first coordinate system, the coordinates of the top left corner of the first region are (0, 0), the coordinates of the bottom right corner of the first region are (2W, 2H), the coordinates of the top left corner of the local image are (W / 2, H / 2), and the coordinates of the bottom right corner of the local image are (3W / 2, 3H / 2).
[0062] In step S231, the five local images of size W*H obtained in step S210 can be input into the keypoint detection model to determine the first coordinates of the 17 keypoints of the human body represented by each local image in the first coordinate system corresponding to each local image. It can be understood that the first region and the first coordinate system corresponding to the five local images are also different because their positions are different.
[0063] In step S241, the first coordinates of each keypoint determined for each local image can be transformed based at least on the dimensions corresponding to the multiple expansion directions and the position of the target detection box, to obtain the second coordinates of each keypoint in the second coordinate system. The origin of the second coordinate system is different from that of the first coordinate system; the second coordinate system takes the upper left corner of the original image as its origin. In the example where five cropped images are directly used as input images (local images), the coordinates of each keypoint in the first detection result can be offset and transformed based on the positional difference of the origins of these two coordinate systems to obtain the coordinates of each keypoint of the target object in the original image. The positional difference of the origins of these two coordinate systems can be determined based on the coordinates of the target detection box in the original image and the positional difference between the origin of the input image and the origin of the first region. The positional difference between the origin of the input image and the origin of the first region can be easily determined based on the dimensions corresponding to the multiple expansion directions. Those skilled in the art will understand the specific implementation of this scheme, and for the sake of brevity, it will not be described in detail here. Furthermore, in the example above where the local image obtained by scaling and padding the cropped small image while maintaining its original aspect ratio is used as the input to the keypoint detection model, it is also necessary to consider that scaling and padding apply additional scaling and offset to the first coordinates of each keypoint to accurately determine the coordinates of each keypoint of the target object in the original image.
[0064] In the above scheme, the coordinates of each keypoint determined by the keypoint detection model are not based on the top-left corner of the local image, but rather on the top-left corner of the first region obtained by expanding the local image. This ensures that the model's predicted keypoint locations encompass the expanded range surrounding the local image, improving the model's adaptability to different image layouts and target object pose variations, and exhibiting better robustness to image noise and occlusion. Furthermore, this scheme is simpler to implement and produces more reasonable output results. Moreover, through simple coordinate transformation, the detected keypoint locations in the local image can be quickly converted to their absolute locations in the original image, facilitating subsequent processing and analysis. In summary, this keypoint localization method, which detects keypoints, including those crossing boundaries, in the local image and accurately maps their coordinates to the original image coordinate system, is more flexible, accurate, and efficient.
[0065] In one implementation, before inputting the local image into the keypoint detection model in step S230, the method further includes the following steps:
[0066] Step S221: Obtain multiple training images and the first annotation information corresponding to each training image. The first annotation information includes the second annotation coordinates of each key point of each first target object in the training image. The coordinate system of the second annotation coordinates has its origin at the top left corner of the training image. The first target object is any target object in the training image.
[0067] Step S222: For each large training image, multiple local image regions of the large training image are cropped to obtain multiple small training images, each of which corresponds to a local image region.
[0068] Step S223: For each training small image, based at least on the size corresponding to multiple expansion directions and the position of the local image region corresponding to the training small image, perform coordinate transformation on the second annotation coordinates of each key point of the first target object to obtain the first annotation coordinates of each key point, which serve as the second annotation information corresponding to the training small image. The coordinate system of the first annotation coordinates is based on the upper left corner of the second region as the origin, and the second region is obtained by expanding the training small image by the corresponding size in multiple expansion directions.
[0069] Step S224: Train the key point detection model using each training small image and the second annotation information corresponding to each training small image until the preset end training condition is reached.
[0070] Taking the aforementioned monitoring and early warning of suspicious individuals at school gates as an example, 1000 images of the captured area from the monitoring equipment can be collected and processed beforehand as training images. These images can all include human figures. The positions of the aforementioned 17 key points of each human figure in each image can be marked using manual annotation, machine annotation, or semi-automatic annotation methods. For human figures that are occluded, the coordinates of the key points of the occluded human figure in the entire image (with the upper left corner of the training image as the origin) can also be marked using various methods such as manual inference. Then, various suitable methods can be used to locally crop the training images. For example, each training image can include multiple human figures. Sliding windows of different sizes can be used to slide on the training images, and each time a window is slid to a position, the window area is used as a reference local image area. The image in the window area is then cropped into reference training images. In this way, for each training image, multiple reference training images can be obtained by cropping, proportional scaling of width and height, and edge padding. Afterwards, methods such as manual screening can be used to select training images containing human figures from these reference training images. For each training image, the coordinates of key points in the corresponding human body region can be found in the training image based on that region. A method corresponding to step S241 can be used to convert the coordinates of these key points in the original image to coordinates in a coordinate system with the top-left corner of the second region corresponding to the training image as the origin. This allows for quick and accurate labeling of the coordinates of each key point in the second region, even if the training image does not contain complete human body parts, serving as the location labels for these key points. The key point detection model can then be trained using each training image and its corresponding second labeling information until a preset end-of-training condition is met. This preset end-of-training condition can be arbitrarily set according to actual needs. For example, after multiple iterations of training, if the loss function calculated from the model's predicted key point locations and the actual values represented by the key point location labels no longer decreases significantly, the training can be considered converged, and the end-of-training condition can be considered met.
[0071] In the above scheme, during the keypoint detection model training, in the label encoding stage of keypoint positions in the training image, unlike existing methods that use the top-left corner of the training thumbnail as the origin for keypoint position labels, this embodiment uses the top-left corner of the extended region of the human body thumbnail as the origin, that is, adding the offset of the extended region to the keypoint position label. In this way, by using a coordinate transformation method that simulates the extension of the local image region to change the keypoint position labels, the model can retain more contextual information on the training thumbnail, thus endowing the model with the ability to detect out-of-bounds keypoints through training. Furthermore, the keypoint detection model trained by the above method can predict not only keypoints beyond the detection box boundary but also keypoints beyond the image boundary, resulting in higher keypoint detection accuracy and a wider range of applications.
[0072] In one implementation, the method further includes the following steps:
[0073] Step S221a: For each training large image, obtain the position of each labeled detection box in the training large image, wherein the labeled detection box is the bounding box of the complete imaging region of the first target object.
[0074] Step S221b involves shrinking each labeled detection box according to different shrinking ratios to obtain multiple training detection boxes. For example, the shrunken detection boxes can also be translated near the human body region to obtain even more training detection boxes.
[0075] Step S221c: The image regions in the multiple training detection boxes are used as at least a portion of the multiple local image regions.
[0076] It's understandable that shrinking transformations can increase the diversity of training samples without increasing the actual amount of data, which helps improve the model's generalization ability. Shrinking transformations can simulate different sizes of target objects in an image, allowing the model to adapt to changes in target object size. Shrinking operations may partially occlude the target object, similar to partial occlusion in real-world scenes, which helps train the model to handle occlusion issues. Because the model encounters targets of different sizes and partial occlusions during training, its robustness to occlusion and truncation phenomena is improved. Through shrinking transformations, the model can learn more compact target feature representations, helping to reduce false detections in complex backgrounds. Furthermore, shrinking transformations help the model better learn the key features of the target, thereby improving the accuracy and reliability of out-of-bounds keypoint detection.
[0077] In one implementation, the keypoint detection model includes a backbone network, a neck network, and a head network, wherein the head network includes convolutional layers and two fully connected layers. Step S231 inputs a local image into the keypoint detection model to determine the first coordinates of each keypoint in a first coordinate system, including the following steps:
[0078] Step S231.1: Use the backbone network to extract features from the local image to obtain a first feature map at at least one scale.
[0079] Step S231.2: The first feature map is fused using the neck network to obtain a fused feature map of at least one scale.
[0080] Step S231.3: Convolve the fused feature map using a convolutional layer to obtain a multi-channel feature map, wherein the number of channels in the multi-channel feature map is equal to the total number of key points for each target object.
[0081] Step S231.4: The two fully connected layers of the head network are used to process multiple one-dimensional arrays composed of pixel values of each channel at each pixel position in the multi-channel feature map, and output a first tensor and a second tensor. Each element in the first tensor represents the probability of the x-coordinate of each keypoint corresponding to each x-coordinate interval covered by the first region in the first coordinate system, and each element in the second tensor represents the probability of the y-coordinate of each keypoint corresponding to each y-coordinate interval covered by the first region in the first coordinate system.
[0082] Step S231.5: Determine the first coordinate of each key point in the first coordinate system based on the first tensor and the second tensor.
[0083] like Figure 3 As shown, a keypoint detection model can include a backbone network, a neck network, and a head network. The backbone network can be used to extract high-dimensional abstract features from images. For example, the backbone network can use CSPDarknet as its basic architecture, which extracts feature images at scales of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 in four stages. The neck network can use RepVGG, which learns multiple branches during the training phase and merges these branches during the deployment phase. Table 1 provides a specific example of a backbone network configuration. As shown in Table 1, the backbone network can consist of 10 modules, including 5 convolutional modules with a kernel radius of 3 and a stride of 2.
[0084] Table 1 Backbone Network Configuration Table
[0085]
[0086] For example, the neck network can use PAFPN as its basic architecture, fusing the three scale features output from the model backbone in a top-down and bottom-up manner, resulting in the output of only 1 / 32 scale features. The head network of the model can be constructed based on SimCC, receiving the single-scale features output from the neck, followed by a convolutional layer with kernel_size=7 and stride=1, changing the number of channels to the number of human keypoints (e.g., 17), and then flattening the H and W dimensions to output an array of size [n, 17, H / 32 * W / 32]. Then, two fully connected layers are used for horizontal and vertical coordinate classification. The coordinate classification uses a fixed size, corresponding to the size of the first region after the model's input image is expanded. Figure 4 The tensor in this embodiment is 2W*2H. Unlike existing technologies that output two tensors of shape [n,17,W] and [n,17,H] (where n represents the number of local images input to the model at one time, i.e., the sample size), the two tensors output in this embodiment are [n,17,2W] and [n,17,2H]. The extra tensor in the last dimension is used for the coordinate classification of the detection box when it exceeds the boundary by 50% in each direction.
[0087] In the above scheme, extracting feature maps of different scales through the backbone network can better capture target objects of different sizes, improving the model's accuracy in keypoint detection. The neck network fuses feature maps of different scales, which helps integrate information at different levels and enhances the model's ability to predict keypoint locations. The multi-channel feature maps output by the convolutional layers, with each channel corresponding to a keypoint, allow the model to independently predict the probability of each keypoint's presence at each pixel location, improving localization accuracy. Utilizing two fully connected layers in the head network to predict the probability distribution of keypoints on the horizontal and vertical coordinates provides probabilistic guidance for accurate keypoint localization compared to single coordinate values, helping to handle issues such as partial occlusion and pose changes. Through probabilistic prediction, the model can better handle uncertainties and noise in images, improving robustness in complex environments. Furthermore, this model structure allows the design of the backbone, neck, and head networks to be adjusted according to different application requirements to adapt to different detection tasks, making keypoint detection more flexible and scalable. This model architecture enables the model to accurately handle multi-scale and probabilistic predictions, making it more adaptable to the needs of keypoint detection in complex scenarios such as multi-person pose estimation or dynamic scenes. By using the above keypoint detection model, the accuracy and robustness of keypoint detection can be significantly improved.
[0088] The above scheme, by making minor modifications to the encoding of the network output and training labels, endows the keypoint detection model with the ability to accurately detect keypoints that exceed the bounds. This not only ensures high detection efficiency and accuracy but also simplifies the preprocessing process compared to existing technologies. This method eliminates the need for expanding the detection box, resulting in cleaner model input, which is highly beneficial for keypoint detection in densely populated scenes. Furthermore, the training method described above enables the keypoint detection model to predict keypoints beyond the detection box boundaries as well as those beyond the image boundaries, thus broadening its applicability.
[0089] In one implementation, the length of the first dimension of the first tensor and the second tensor is equal to the total number, the length of the second dimension of the first tensor is equal to M times the third dimension, the length of the second dimension of the second tensor is equal to N times the fourth dimension, the third dimension is equal to the width of the first region, the fourth dimension is equal to the height of the first region, and M and N are both positive integers.
[0090] Step S231.5, based on the first tensor and the second tensor, determine the first coordinates of each keypoint in the first coordinate system, including:
[0091] For the key points represented by the i-th layer of the first dimension
[0092] Find the first element with the largest value among all elements in the first tensor located in the i-th layer, and determine the reference x-coordinate of the key point based on the position index of the first element in the second dimension, where i is a positive integer;
[0093] Find the second element with the largest value among all elements in the second tensor located in the i-th layer, and determine the reference x-coordinate of the key point based on the position index of the second element in the second dimension;
[0094] Based on the reference x-coordinate and reference y-coordinate, determine the first coordinate of the key point in the first coordinate system, where the x-coordinate in the first coordinate system is equal to the reference x-coordinate divided by M, and the y-coordinate in the first coordinate system is equal to the reference y-coordinate divided by N.
[0095] In this embodiment, the third dimension is equal to the width of the first region, and the fourth dimension is equal to the height of the first region. (See reference...) Figure 4The third dimension is 2W, and the fourth dimension is 2H. M and N can be positive integers greater than or equal to 1. In the example where both M and N are equal to 1, the dimensions of the two tensors output by the head network can be [n, 17, 2W] and [n, 17, 2H]. In this example, the first and second tensors output by the head network perform classification at the pixel level, meaning the classification result can be based on the entire pixel, not a sub-pixel. This scheme has relatively low computational cost, is sufficient for most applications, and is relatively simple to implement. In this scheme, the x-coordinate in the first coordinate system can be directly determined to be equal to the reference x-coordinate, and the y-coordinate in the first coordinate system can be equal to the reference y-coordinate.
[0096] In examples where both M and N are greater than 1, the classification results expressed by the first and second tensors output by the head network can achieve finer-grained classification accuracy below the pixel level. For example, sub-pixel, 1 / 3-pixel, and 1 / 4-pixel classifications can be achieved. For instance, when both M and N are equal to 2, the sizes of the two tensors output by the head network can be [n, 17, 4W] and [n, 17, 4H], enabling sub-pixel (0.5-pixel) classification. For example, the head network calculates the maximum values of the horizontal and vertical coordinates of the 17 key points representing the human body in each local image, along with their corresponding indices. The maximum value represents the confidence level, and the corresponding index represents the coordinate position. By calculating the maximum values of the horizontal and vertical coordinates of each key point and their corresponding indices, the sub-pixel-level position of each key point can be accurately determined. In this example, since it is sub-pixel classification, the coordinate positions output by the model need to be divided by 2 to obtain the first coordinates of each key point in the first coordinate system. That is, the horizontal coordinate in the first coordinate system is equal to the reference horizontal coordinate divided by 2, and the vertical coordinate in the first coordinate system is equal to the reference vertical coordinate divided by 2.
[0097] It is understandable that fine-grained classification at the sub-pixel level can provide more refined positional information below the pixel level, thereby improving the accuracy of keypoint localization. Furthermore, in pose estimation, smoother body contours and more natural pose estimates can be generated, enhancing the model's ability to express pose changes, optimizing model performance, and enabling the model to better adapt to different image resolutions and scaling levels. It can also reduce quantization errors and improve the accuracy of coordinate estimation.
[0098] This application also provides a key point detection device. For example... Figure 5 As shown, the key point detection device 500 includes:
[0099] The acquisition module 510 is used to acquire the position of the target detection box in the original image, wherein the original image includes the imaging region of the target object, and the target detection box is the position box of at least part of the imaging region of the target object.
[0100] The cropping module 520 is used to crop the original image based on the target detection box to obtain the cropped local image;
[0101] The detection module 530 is used to input a local image into a key point detection model to obtain a first detection result, wherein the first detection result includes the position information of key points that extend beyond the boundary of the local image;
[0102] The determination module 540 is used to determine the positions of each key point of the target object in the original image based on the first detection result.
[0103] In one embodiment, the detection module 530 includes:
[0104] The detection unit is used to input the local image into the key point detection model and determine the first coordinate of each key point in the first coordinate system, wherein the first coordinate system takes the upper left corner of the first region as the origin, and the first region is obtained by expanding the local image to the corresponding size in multiple expansion directions.
[0105] The determination module 540 includes:
[0106] The coordinate transformation unit is used to perform coordinate transformation on the first coordinate of each key point based on the size corresponding to multiple expansion directions and the position of the target detection box, so as to obtain the second coordinate of each key point in the second coordinate system, wherein the second coordinate system takes the upper left corner of the original image as the origin.
[0107] In one embodiment, the key point detection device 500 further includes:
[0108] The first training acquisition module is used to acquire multiple training images and the first annotation information corresponding to each training image. The first annotation information includes the second annotation coordinates of each key point of each first target object in the training image. The coordinate system of the second annotation coordinates takes the upper left corner of the training image as the origin. The first target object is any target object in the training image.
[0109] The training cropping module is used to crop multiple local image regions of each training large image to obtain multiple training small images, each training small image corresponding to a local image region.
[0110] The training coordinate transformation module is used to transform the second annotation coordinates of each key point of the first target object for each training small image, based at least on the size corresponding to multiple expansion directions and the position of the local image region corresponding to the training small image, to obtain the first annotation coordinates of each key point, which are used as the second annotation information corresponding to the training small image. The coordinate system of the first annotation coordinates is based on the upper left corner of the second region as the origin, and the second region is obtained by expanding the training small image to the corresponding size in multiple expansion directions.
[0111] The training and detection module is used to train the key point detection model using each training small image and the corresponding second annotation information of each training small image until the preset end training condition is reached.
[0112] In one embodiment, the key point detection device 500 further includes:
[0113] The second training acquisition module is used to acquire the position of each labeled detection box in each training large image, wherein the labeled detection box is the bounding box of the complete imaging region of the first target object.
[0114] The shrinking transformation module is used to shrink each labeled detection box according to different shrinking ratios to obtain multiple training detection boxes.
[0115] A local image region determination module is used to identify image regions within multiple training detection boxes as at least a portion of multiple local image regions.
[0116] In one implementation, the keypoint detection model includes a backbone network, a neck network, and a head network. The head network includes convolutional layers and two fully connected layers. The detection unit includes:
[0117] The feature extraction subunit is used to extract features from local images using the backbone network to obtain a first feature map at at least one scale.
[0118] The feature fusion unit is used to fuse the first feature map using the neck network to obtain a fused feature map of at least one scale.
[0119] A convolutional unit is used to convolve the fused feature map using a convolutional layer to obtain a multi-channel feature map, where the number of channels in the multi-channel feature map is equal to the total number of key points for each target object.
[0120] The processing output unit is used to process multiple one-dimensional arrays composed of pixel values of each channel at each pixel position in the multi-channel feature map using two fully connected layers of the head network, and outputs a first tensor and a second tensor. Each element in the first tensor represents the probability of the horizontal coordinate of each key point corresponding to each horizontal coordinate interval covered by the first region in the first coordinate system, and each element in the second tensor represents the probability of the vertical coordinate of each key point corresponding to each vertical coordinate interval covered by the first region in the first coordinate system.
[0121] The first determining unit is used to determine the first coordinate of each key point in the first coordinate system based on the first tensor and the second tensor.
[0122] In one embodiment, the multiple expansion directions include the image width direction and the image height direction, and the key point detection device 500 further includes:
[0123] The first calculation unit is used to calculate the product of the first width and the first multiple to obtain the first size corresponding to the image width direction, wherein the first width is equal to the image width of each training small image and the local image;
[0124] The second calculation unit is used to calculate the product of the first height and the second multiple to obtain the second dimension corresponding to the image height direction, wherein the first height is equal to the image height of each training thumbnail and local image.
[0125] In one implementation, the length of the first dimension of the first tensor and the second tensor is equal to the total number of dimensions; the length of the second dimension of the first tensor is equal to M times the third dimension; the length of the second dimension of the second tensor is equal to N times the fourth dimension; the third dimension is equal to the width of the first region; and the fourth dimension is equal to the height of the first region, where M and N are both positive integers; the first determining unit includes:
[0126] The first search determination sub-unit is used to find the first element with the largest value among the elements of the first tensor located in the i-th layer of the first dimension for the key point represented by the first element, and determine the reference x-coordinate of the key point according to the position index of the first element in the second dimension, where i is a positive integer;
[0127] The second lookup determination subunit is used to find the second element with the largest value among the elements of the second tensor located in the i-th layer of the first dimension for the key point, and determine the reference x-coordinate of the key point based on the position index of the second element in the second dimension.
[0128] The coordinate calculation subunit is used to determine the first coordinate of the key point in the first coordinate system based on the reference abscissa and reference ordinate, wherein the abscissa in the first coordinate is equal to the reference abscissa divided by M, and the ordinate in the first coordinate is equal to the reference ordinate divided by N.
[0129] like Figure 6 As shown, this application embodiment also provides a terminal device 600, including: at least one processor 610 ( Figure 6 The diagram shows only one processor, a memory 620, and a computer program 630 stored in the memory 620 and executable on at least one processor 610. When the processor 610 executes the computer program 630, it implements the steps of the above-described key point detection method.
[0130] Terminal devices may include, but are not limited to, processors and memory. Figure 6 This is merely an example of a terminal device and does not constitute a limitation on the terminal device. It may include more or fewer components than illustrated, or combine certain components, or use different components. The processor may be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0131] It should be noted that the information interaction and execution process between the above-mentioned devices / modules are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0132] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional modules is merely an example. In practical applications, the functions described above can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The functional modules in the embodiments can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules can be implemented in hardware or as software functional modules. Furthermore, the specific names of the functional modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0133] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the steps in the above-described key point detection method.
[0134] This application provides a computer program product that, when run on a terminal device, enables the terminal device to perform the steps in the aforementioned key point detection method.
[0135] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A key point detection method, characterized in that, The method comprises: obtaining a position of a target detection frame in an original image, wherein the original image comprises an imaging region of a target object, and the target detection frame is a position frame of at least part of the imaging region of the target object; cropping the original image according to the target detection frame to obtain a cropped local image; inputting the local image into a key point detection model to obtain a first detection result, wherein the first detection result comprises position information of key points that are outside the boundary of the local image; determining the positions of the key points of the target object in the original image according to the first detection result.
2. The keypoint detection method of claim 1, wherein, The method comprises: inputting the local image into the key point detection model to determine a first coordinate of each key point in a first coordinate system, wherein the first coordinate system takes a top-left corner point of a first region as an origin, and the first region is obtained by expanding the local image in multiple expansion directions by corresponding sizes; The method comprises: performing coordinate transformation on the first coordinate of each key point according to at least the sizes corresponding to the multiple expansion directions and the position of the target detection frame to obtain a second coordinate of each key point in a second coordinate system, wherein the second coordinate system takes a top-left corner point of the original image as an origin.
3. The keypoint detection method of claim 2, wherein, Before the local image is input into the key point detection model, the method further comprises: obtaining a plurality of training large images and first annotation information corresponding to each training large image, wherein the first annotation information comprises second annotation coordinates of the key points of each first target object in the training large image, the coordinate system of the second annotation coordinates takes a top-left corner point of the training large image as an origin, and the first target object is any target object in the training large image; cropping a plurality of local image regions of each training large image to obtain a plurality of training small images, each training small image corresponding to a local image region; for each training small image, performing coordinate transformation on the second annotation coordinates of the key points of the first target object according to at least the sizes corresponding to the multiple expansion directions and the position of the local image region corresponding to the training small image to obtain first annotation coordinates of each key point as second annotation information corresponding to the training small image, wherein the coordinate system of the first annotation coordinates takes a top-left corner point of a second region as an origin, and the second region is obtained by expanding the training small image in the multiple expansion directions by corresponding sizes; training the key point detection model by using each training small image and the second annotation information corresponding to the training small image until a preset end training condition is reached.
4. The keypoint detection method of claim 3, wherein, The method further comprises: for each training large image, obtaining the positions of each annotation detection frame in the training large image, wherein the annotation detection frame is a boundary frame of the complete imaging region of the first target object; performing internal shrinkage transformation on each annotation detection frame according to different internal shrinkage ratios to obtain a plurality of training detection frames; The image regions in the plurality of training bounding boxes are taken as at least part of the plurality of local image regions.
5. The keypoint detection method of claim 2, wherein, The key point detection model comprises a backbone network, a neck network and a head network, the head network comprises a convolution layer and two fully connected layers, and the inputting of the local image into the key point detection model to determine the first coordinates of each key point in the first coordinate system comprises: feature extraction of the local image by using the backbone network to obtain a first feature map of at least one scale; fusion of the first feature map by using the neck network to obtain a fused feature map of at least one scale; convolution of the fused feature map by using the convolution layer to obtain a multi-channel feature map, wherein the number of channels of the multi-channel feature map is equal to the total number of key points of each target object; processing of a plurality of one-dimensional arrays composed of pixel values of each channel of each pixel position in the multi-channel feature map by using the two fully connected layers of the head network, and outputting a first tensor and a second tensor, wherein each element in the first tensor respectively represents a probability of a horizontal coordinate of each key point corresponding to each horizontal coordinate interval covered by the first region in the first coordinate system, and each element in the second tensor respectively represents a probability of a vertical coordinate of each key point corresponding to each vertical coordinate interval covered by the first region in the first coordinate system; determination of the first coordinates of each key point in the first coordinate system according to the first tensor and the second tensor. 6.The key point detection method according to any one of claims 2 to 5, wherein, The plurality of expansion directions comprises an image width direction and an image height direction, and the method further comprises: multiplication of a first width and a first multiple to obtain a first size corresponding to the image width direction, wherein the first width is equal to the image width of each training sub-image and local image; multiplication of a first height and a second multiple to obtain a second size corresponding to the image height direction, wherein the first height is equal to the image height of each training sub-image and local image.
7. The keypoint detection method of claim 6 depending on claim 5, wherein, The length of the first dimension of the first tensor and the second tensor is equal to the total number, the length of the second dimension of the first tensor is equal to M times of the third size, the length of the second dimension of the second tensor is equal to N times of the fourth size, the third size is equal to the width of the first region, the fourth size is equal to the height of the first region, and M and N are positive integers; The determination of the first coordinates of each key point in the first coordinate system according to the first tensor and the second tensor comprises: for the key point represented by the i-th layer of the first dimension, finding a first element with the largest value from each element of the first tensor located at the i-th layer, and determining a reference horizontal coordinate of the key point according to the position index of the first element in the second dimension, wherein i is a positive integer; finding a second element with the largest value from each element of the second tensor located at the i-th layer, and determining a reference horizontal coordinate of the key point according to the position index of the second element in the second dimension; According to the reference abscissa and the reference ordinate, a first coordinate of the key point in the first coordinate system is determined, wherein an abscissa in the first coordinate is equal to the reference abscissa divided by M, and an ordinate in the first coordinate is equal to the reference ordinate divided by N.
8. A key point detection apparatus characterized by, The method comprises the following steps: An acquisition module is configured to acquire a position of a target detection frame in an original image, wherein the original image comprises an imaging region of a target object, and the target detection frame is a position frame of at least part of the imaging region of the target object; A cropping module is configured to crop the original image according to the target detection frame to obtain a cropped local image; A detection module is configured to input the local image into a key point detection model to obtain a first detection result, wherein the first detection result comprises position information of a key point that exceeds a boundary of the local image; A determination module is configured to determine positions of each key point of the target object in the original image according to the first detection result.
9. A terminal device, comprising: The method comprises the following steps: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the key point detection method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the key point detection method according to any one of claims 1 to 7.