Object key point detection method and device, training method and device, and computing device
By adaptively adjusting the width-length ratio and cropping and scaling technology of object boxes, the problems of image distortion and increase background data are solved, and the accuracy and stability of key point detection are improved, especially in lightweight neural networks and video stream detection.
Patent Information
- Application Number
- CN202111294166.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-03
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-11-03
AI Technical Summary
The existing key point detection methods can easily cause image distortion or increase background data before processing the captured images into neural network input, which affects the detection accuracy, especially in lightweight neural networks and video stream detection.
By introducing preset frame expansion coefficients to adaptively adjust the size of the object box, making its width-length ratio close to the preset input width-length ratio of the neural network, and reducing background data, combining cropping and scaling processing, a stable network input image is obtained.
It improves the accuracy of key point detection, reduces the impact of background data in lightweight neural networks, and provides more stable image input in video stream detection, achieving more stable key point detection.
Smart Images

Figure CN114332483B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computers, and more specifically, to a method and apparatus for detecting object key points, a training method and apparatus, and a computing device. Background Art
[0002] Object key points are crucial for describing object postures and predicting object behaviors. Therefore, the detection of object key points is the basis for many applications in the field of computer vision, such as intelligent video surveillance, virtual reality, short videos, fitness applications, and so on. The object can be a human body or an animal. The detection of object key points mainly detects some skeletal key points of the object. For example, for a human body, it can include: left eye, right eye, left ear, right ear, nose, chest, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip joint, right hip joint, left knee, right knee, left ankle, right ankle, etc., so as to describe the skeletal information of the object through the coordinates of the key points. Usually, the key points of the object in the collected image can be detected by collecting the image of the object, so as to perform other operations according to the detected key points, such as determining the limb movements of the object, triggering limb special effects, body search of the human body, driving the actions of virtual avatars, and so on.
[0003] Currently, the detection of key points can be performed based on a neural network. The input data of the neural network has a preset size. Usually, before inputting the collected image into the neural network to detect the key points of the object, it is necessary to first process the collected image into a network input image with the same size as the preset size of the input data of the neural network, and then the neural network will perform key point detection based on the network input image. It can be seen that the accuracy of key point detection by the neural network is closely related to the process from the collected image to the network input image.
[0004] Therefore, a solution that can reasonably process the collected image to improve the accuracy of key point detection based on the neural network is needed. Summary of the Invention
[0005] According to one aspect of the present disclosure, there is provided a method for detecting object key points, including: determining a first object box corresponding to the image to be detected based on the pose of the object in the image to be detected, where the first object box corresponds to the region of interest where the object in the image to be detected is located; adjusting the size of the first object box based on a preset size of the input data of the trained neural network, the box size of the first object box, and a preset box expansion coefficient, to obtain a second object box, where the preset box expansion coefficient makes the aspect ratio of the second object box fall between the aspect ratio corresponding to the preset size and the aspect ratio of the first object box, and the preset size is a fixed size of the input data that can be processed by the trained neural network; obtaining a cropped image of the image to be detected based on the second object box, and adjusting the cropped image based on the preset size to obtain a network input image corresponding to the cropped image, where the size of the network input image is equal to the preset size; and detecting object key points in the network input image based on the trained neural network.
[0006] According to another aspect of the present disclosure, there is also provided a method for training a neural network for object key point detection, including: obtaining a training image set, where each training image includes an object instance and key point annotations for the object instance; for each training image, randomly selecting a plurality of different training box expansion coefficients from a predetermined set; for each training image, processing the training image based on the plurality of different training box expansion coefficients to obtain a plurality of network input images having the same preset size as the input data of the neural network, thereby realizing the augmentation of the network input image set of the training image set, where the preset size is a fixed size of the input data that can be processed by the neural network; and training the neural network using the augmented network input image set to obtain a trained neural network.
[0007] According to another aspect of the present disclosure, there is also provided an apparatus for detecting object key points, including: a determination module, configured to determine a first object box corresponding to the image to be detected based on the pose of the object in the image to be detected, the first object box corresponding to the region of interest where the object is located in the image to be detected; an object box adjustment module, configured to adjust the size of the first object box based on a preset size of the input data of the trained neural network, the box size of the first object box, and a preset box expansion coefficient to obtain a second object box, where the preset box expansion coefficient makes the aspect ratio of the second object box fall between the aspect ratio corresponding to the preset size and the aspect ratio of the first object box, and the preset size is a fixed size of the input data that can be processed by the trained neural network; a cropping module, configured to obtain a cropped image of the image to be detected based on the second object box, and adjust the cropped image based on the preset size to obtain a network input image corresponding to the cropped image, the size of the network input image being equal to the preset size; and a detection module, configured to detect the key points of the object in the network input image based on the trained neural network.
[0008] According to still another aspect of the present disclosure, there is also provided a training apparatus for a neural network for object key point detection, including: an acquisition module, configured to acquire a training image set, each training image including an object instance and key point annotations for the object instance; a selection module, configured to randomly select a plurality of different training box expansion coefficients from a predetermined set for each training image; an image processing module, configured to process each training image based on the plurality of different training box expansion coefficients to obtain a plurality of network input images having the same preset size as the input data of the neural network, thereby realizing the amplification of the network input image set of the training image set, the preset size being a fixed size of the input data that can be processed by the neural network; and a training module, configured to train the neural network using the amplified network input image set to obtain a trained neural network.
[0009] According to another aspect of the present disclosure, there is also provided a computing device, including: a processor; a memory having a computer program stored thereon, and when the computer program is executed by the processor, the one or more processing units are caused to perform various operations described in the steps of the object key point detection method and the training method as described above.
[0010] According to another aspect of the present disclosure, there is also provided a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the processor is caused to perform various operations described in the steps of the object key point detection method and the training method for object key point detection as described above.
[0011] According to another aspect of the present disclosure, there is also provided a computer program product, including a computer program, which when executed by a processor implements various operations described in the steps of the object key point detection method and the training method of the neural network for object key point detection as described above.
[0012] Through the above solution of the present disclosure, by introducing a preset frame expansion coefficient, the first object frame can be appropriately enlarged to be as close as possible to the preset input aspect ratio of the trained neural network (facilitating scaling and avoiding distortion) and including as little background data as possible (improving the accuracy of feature extraction), thereby improving the detection accuracy. On the other hand, in video stream detection, compared with the traditional method, the frame expansion range can be reduced, and thus the degree of amplification of the jitter of the first object frame between adjacent video frames can be reduced. Therefore, in video stream detection, a more stable image input can be provided to the neural network to achieve more stable key point detection. Description of the Drawings
[0013] Figure 1A-1B A schematic diagram showing the process of obtaining the network input image of the neural network from the captured image.
[0014] Figure 2 A schematic diagram showing the overall process of key point detection based on the adaptively adjusted object frame according to an embodiment of the present disclosure.
[0015] Figure 3 A schematic flowchart showing the object key point detection method according to an embodiment of the present disclosure.
[0016] Figure 4 Shows reference Figure 3 More details in step S320 described.
[0017] Figure 5A A schematic diagram showing the process of obtaining the network input image of the neural network from the captured image.
[0018] Figure 5B A schematic diagram showing the process of obtaining the network input image of the neural network from the captured image according to an embodiment of the present disclosure.
[0019] Figure 6A-6B A schematic flowchart showing two determination methods for the preset frame expansion coefficient.
[0020] Figure 7 A schematic flowchart showing the steps of obtaining an optimized neural network by performing frame expansion enhancement on the initially trained neural network according to an embodiment of the present disclosure.
[0021] Figure 8 A schematic flowchart showing the training method of the neural network for object key point detection according to an embodiment of the present disclosure.
[0022] Figure 9-10 The structural block diagram of a detection device for object key points according to an embodiment of the present disclosure is shown.
[0023] Figure 11 The structural block diagram of a training device for a neural network for object key point detection according to an embodiment of the present disclosure is shown.
[0024] Figure 12 The structural block diagram of a computing device according to an embodiment of the present disclosure is shown. Detailed implementation manners
[0025] In order to make the objectives, technical solutions, and advantages of the present disclosure more obvious, exemplary embodiments according to the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure. It should be understood that the present disclosure is not limited by the exemplary embodiments described herein.
[0026] In this specification and the accompanying drawings, steps and elements having substantially the same or similar steps and elements are denoted by the same or similar reference numerals, and repeated descriptions of these steps and elements will be omitted. At the same time, in the description of the present disclosure, terms such as "first", "second", etc. are only used for distinguishing descriptions, and cannot be understood as indicating or implying relative importance or order.
[0027] Machine learning is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how a computing device simulates or implements human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve its own performance. Machine learning is the core of AI and is the fundamental way to endow computing devices with intelligence; the so-called machine learning is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory; it specifically studies how a computing device simulates or implements human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve its own performance. Deep learning, on the other hand, is a technology for machine learning using a deep neural network system; machine learning / deep learning usually includes multiple technologies such as artificial neural networks, reinforcement learning, supervised learning, and unsupervised learning.
[0028] As described above, before inputting the captured image into a trained neural network (keypoint detection neural network) for keypoint detection of an object, it is necessary to first process the captured image (image to be detected) into a network input image with the same size as the preset size of the input data of the trained neural network, and then the trained neural network will perform keypoint detection based on the network input image. Among them, the trained neural network can be a neural network model that has been trained with multiple training images having object instances and keypoint annotations for the object instances. In the description of the present disclosure, the preset size is the fixed size of the input data that can be processed by the neural network.
[0029] First, the object in the image to be detected can be framed with an object box, and the corresponding image part in the object box can be cropped to obtain a cropped image, so as to remove unnecessary information in the captured image and retain all information of the object, which is convenient for reducing the amount of computation and improving the detection accuracy. The size of the object box is related to the pose of the object. The pose of the object can be related to the actions made by the limbs of the object, such as a person standing with limbs outstretched, squatting, lying flat, etc.
[0030] If the aspect ratio of the object box determined based on the pose of the object is the same as the aspect ratio corresponding to the preset size of the input data of the trained neural network, then directly crop the captured image based on the object box, perform proportional scaling on the cropped image, and then input the scaled image into the trained neural network as the network input image. For example, assume that the size of the first object box is 10×10 (unit: pixel, here 10×10 is only an example, and it can actually be more pixels), and the preset size of the input data of the trained neural network is 8×8, then the width and length of the cropped image (10×10) can be scaled from 10 and 10 to 8 respectively.
[0031] However, since the pose of the object can change continuously, in more cases, the aspect ratio of the object box often differs from the aspect ratio corresponding to the preset size of the input data of the trained neural network. In this case, Figure 1A-1B The schematic diagrams of two processing methods are shown.
[0032] In the first method, the cropped image based on the object box can be directly scaled and deformed to the preset size of the input data of the trained neural network by size, that is, the width and length of the cropped image can be scaled correspondingly. As Figure 1A shown, assume that the first object box is 15×15 (aspect ratio is 1:1), then the size of the cropped image is also 15×15 (aspect ratio is 1:1), and the preset size of the input data of the trained neural network is 10×20, then the width can be scaled from 15 to 10, and the length can be scaled from 15 to 20.
[0033] In the second method, the object box can be enlarged first so that the aspect ratio of the object box is the same as the aspect ratio corresponding to the preset size of the trained neural network. Then, a cropped image can be obtained based on the enlarged object box, and then the cropped image can be scaled proportionally to the preset size of the input data of the trained neural network. As Figure 1B shown, assuming that the first object box is 15×15 and the preset size of the input data of the trained neural network is 10×20, the first object box can be enlarged to a second object box of 15×30 (enlarging the length). Then, a cropped image of 15×30 can be obtained based on the second object box, and the width and length of the 15×30 cropped image can be scaled by 1.5 times to 10 and 20 respectively.
[0034] The above two methods can obtain a network input image with the preset size suitable for input to the trained neural network, but there are the following defects.
[0035] In the first method, due to the various postures of the object, directly scaling the size of the cropped image easily causes the object to be distorted, resulting in the trained neural network being unable to detect key points well.
[0036] In the second method, in order to avoid the image distortion in the first method, the object box is enlarged so that the cropped image has the aspect ratio corresponding to the preset size. However, there are also certain defects. Since the key point detection based on the neural network is generally performed on terminal devices such as mobile devices, the neural network is generally lightweight (for example, the neural network can have fewer layers and smaller preset size of input data, or the amount of computation is less than 100 million multiply-accumulate operations). Therefore, the feature extraction ability of the trained neural network for the input image will be limited. In the case of enlarging the object box, some background data (irrelevant data) will be added, resulting in a small proportion of object data but a large proportion of background data in the network input image. Therefore, if the feature extraction ability of the trained neural network is limited, it will affect the key point detection effect. As Figure 1B shown, some background data is added in the long direction, so the proportion of object data will decrease, and the trained neural network will perform feature extraction and processing on more unnecessary background data.
[0037] Therefore, the present disclosure proposes a method for detecting object key points based on an adaptively adjusted object box, so that while ensuring that the object in the network input image will not be greatly distorted relative to the original image, the range that needs to be enlarged can be reduced to reduce the influence of background data on the detection result, so that more object information can be extracted using the limited feature extraction ability even in the case of a lightweight neural network, and the detection accuracy can be improved. However, it should be understood that the various solutions proposed in the embodiments of the present disclosure are also applicable to non-lightweight neural networks.
[0038] Figure 2 A schematic diagram showing the overall process of key point detection based on an adaptively adjusted object box according to an embodiment of the present disclosure.
[0039] First, after obtaining the collected image to be detected (RGB image), an object detector can be called to determine the first object box (the preliminarily determined object box). Optionally, in the case of key point detection for a video stream (i.e., consecutive video frame images), the first object box corresponding to the current video frame image can also be obtained based on the key point detection result of the previous video frame image. For example, the positions of the same part of the human body in two adjacent video frame images often have a certain correlation and are not very different. Therefore, an object box can be determined based on the positions of the human key points detected in the previous video frame image in the previous video frame image and saved as the first object box corresponding to the current video frame image.
[0040] Then, after obtaining the first object box, the first object box can be adaptively adjusted. For example, the length or width can be enlarged to obtain the second object box. The specific adjustment method will be described later.
[0041] Next, the image to be detected is cropped based on the second object box, and the cropped image is scaled according to the preset size of the trained neural network to obtain a network input image for input into the trained neural network. The object in the network input image obtained based on the second object box will not be greatly distorted relative to the original image and does not include excessive background data.
[0042] Finally, key point detection is performed on the network input image based on the trained neural network, and the result is output. In addition, in the case of video stream detection, the first object box of the next video frame image can also be generated and saved according to the positions of the detected key points.
[0043] In the above combination Figure 2 After a schematic description of the process of key point detection based on an adaptively adjusted object box, more details of the key point detection method according to an embodiment of the present disclosure will be described below in combination with Figures 3-5B A flowchart showing the method for detecting object key points according to an embodiment of the present disclosure. This method can be executed at various terminals.
[0044] Figure 3 As
[0045] As Figure 3As shown, in step S310, a first object box corresponding to the image to be detected is determined based on the pose of the object in the image to be detected, and the first object box corresponds to the region of interest (ROI) where the object in the image to be detected is located.
[0046] Optionally, the first object box can be a rectangle or a square, and its specific size is determined according to the region of interest where the object is located. That the first object box corresponds to the region of interest where the object is located can mean that the first object box will enclose the entire object within it. For example, the smallest rectangle enclosing the object can be used as the first object box.
[0047] Optionally, the image to be detected can be a single image or a video frame image in a video stream.
[0048] Optionally, a pre-set object detector can be called to detect the image to be detected to determine the first object box. Or, when the image to be detected is a video frame image, the first video frame image can be used to determine the corresponding first object box by calling the object detector. Starting from the second video frame image, in addition to calling the object detector, the first object box corresponding to the current video frame image can also be generated based on the key point detection results of the previous video frame image.
[0049] In step S320, the size of the first object box is adjusted based on the preset size of the input data of the trained neural network, the box size of the first object box, and a preset box expansion coefficient to obtain a second object box, where the preset box expansion coefficient makes the aspect ratio of the second object box fall between the aspect ratio corresponding to the preset size and the aspect ratio of the first object box. For example, the preset box expansion coefficient is a value between 0 and 1.
[0050] Optionally, the preset size of the input data of the trained neural network, the box size of the first object box, and other sizes mentioned in this article can all be represented by the number of pixels. For example, a size of 128×128 means that both the width and the length are 128 pixels, and 128 is the width value or the length value.
[0051] Optionally, the trained neural network can be a pre-trained key point detection neural network. For example, a model trained using a general training image set (such as the COCO dataset), where reference can be made to Figure 1BThe network input images of each training image are obtained in the second manner shown (the distortion of the first manner is relatively serious, so it is not considered here) for training the model. The trained neural network has the ability to detect key points. To distinguish it from the optimized neural network to be described later, the trained neural network obtained in this way is also referred to as the initially trained neural network in the following (because the frame expansion coefficient is not applied). Therefore, this embodiment of the present disclosure realizes improvements in the inference process (the process of detecting key points of the image to be detected using the trained neural network) to improve the accuracy of key point detection. This embodiment does not involve improvements in the training process for the time being.
[0052] Optionally, the preset frame expansion coefficient is used to determine the degree of expanding the first object frame, for example, how to increase the width or length of the first object frame so that the aspect ratio of the first object frame can be made as close as possible to the aspect ratio of the preset size of the input data of the trained neural network (hereinafter simply referred to as the preset input aspect ratio or the first aspect ratio) (facilitating scaling to avoid distortion) and including as little background data as possible (for example, less background data than that in the method two described in the reference Figure 1B to improve the feature extraction accuracy), so that it can cooperate with the subsequent cropping and scaling processes to enable the trained neural network to more accurately detect key points.
[0053] More details of step S320 will be described in detail later with reference to Figure 4 this.
[0054] In step S330, a cropped image of the image to be detected is obtained based on the second object frame, and the cropped image is adjusted based on the preset size of the trained neural network to obtain a network input image corresponding to the cropped image, and the size of the network input image is the same as the preset size of the trained neural network.
[0055] For example, after obtaining the second object frame, the image to be detected can be cropped based on the second object frame, and the obtained cropped image has the same size as the second object frame. Therefore, the cropped image also needs to be scaled to obtain an image with a size suitable for inputting into the trained neural network as the network input image. This process can also be regarded as a proportional adjustment process.
[0056] For example, if the width and length of the cropped image are 15 and 24 respectively, and the width and length corresponding to the preset size of the trained neural network are 10 and 20, then the width of the cropped image can be scaled down to 10, and the length of the cropped image can be scaled down to 20 to obtain the network input image. It should be noted that although the aspect ratio of the width and length of the network input image is different from that of the cropped image, this can be equivalently regarded as a change in the stretching ratio of a certain amount of the object's body shape, which is equivalent to the changes in height, width, and body shape in vision. Since the trained neural network has a wide distribution of the height, width, and body shape of the objects in the training image set and is highly robust to body shape changes, it can also accurately identify. The stretching here is different from the stretching Figure 1A described above that will cause distortion, because in this case, the second object box is obtained by enlarging the first object box, and the aspect ratio of the width and length of the cropped image obtained based on the second object box is relatively close to the preset size of the trained neural network. Therefore, the stretching amplitude will not be very large and will not cause distortion.
[0057] In step S340, the object key points in the network input image are detected based on the trained neural network.
[0058] Optionally, the network input image includes object data and partial background data. The trained neural network will extract features from the network input image and perform key point detection based on the extracted features.
[0059] For example, the positions of each key point in the network input image can be detected, and based on this, the positions of each key point in the original image to be detected can be obtained. Since the processes from the image to be detected to the cropped image and from the cropped image to the network input image are known, and the corresponding relationship between the pixels of the images obtained in each process is also known, the key point positions of the object in the image to be detected can be determined from the finally determined key point positions.
[0060] By referring to Figure 3 the described object key point detection method, the size of the first object box that frames the object in the image to be detected is adjusted based on the preset size of the input data of the trained neural network, the box size of the first object box, and the preset frame expansion coefficient, so that the first object box can be appropriately enlarged to be as close as possible to the preset input aspect ratio of the trained neural network (facilitating scaling and avoiding distortion) and including as little background data as possible (improving feature extraction accuracy). And through the cropping and scaling process, the network input image is obtained, and the trained neural network performs key point detection on the network input image, which can improve the detection accuracy.
[0061] On the other hand, in video stream detection, the jitter of the first object box corresponding to the video frame image causes the network input image of the trained neural network to be different for each frame. Through the above method, whether it is to call the object detection box or generate the first object box corresponding to the current video frame image based on the key point detection result of the previous video frame image, by appropriately increasing the first object box, relative to the reference Figure 1B described in Method 2, the frame expansion range (the increase amplitude from the first object box to the second object box) can be reduced, and thus the degree of amplification of the jitter of the first object box between adjacent video frames can be reduced. Therefore, in video stream detection, a more stable image input can be provided to the neural network to achieve more stable key point detection.
[0062] The following Figures 4-5B combines Figure 3 to describe more details of step S320 (obtaining the second object box from the first object box) in
[0063] As Figure 4 shown, in step S320-1, the aspect ratio corresponding to the preset size is determined based on the width value and the length value corresponding to the preset size of the input data of the trained neural network, and used as the first aspect ratio.
[0064] For example, the width value and the length value corresponding to the preset size are respectively represented as H net and W net , and the corresponding first aspect ratio is:
[0065]
[0066] In step S320-2, the aspect ratio of the first object box is determined based on the width value and the length value corresponding to the box size of the first object box, and used as the second aspect ratio.
[0067] For example, the width value and the length value corresponding to the box size of the first object box are respectively represented as Hin and Win, and the corresponding second aspect ratio is:
[0068]
[0069] In step S320-3, the stretching coefficient of the first object box is determined based on the first aspect ratio, the second aspect ratio, and the preset frame expansion coefficient.
[0070] Optionally, the product of the difference between the second aspect ratio and the first aspect ratio and the preset frame expansion coefficient can be determined as the aspect ratio adjustment amount, and then the stretching coefficient of the first object box can be determined as the sum of the first aspect ratio and the aspect ratio adjustment amount.
[0071] For example, it can be expressed by the following formula:
[0072] Rr = R net + (R in - R net ) * S(3)
[0073] Wherein, S is a preset frame expansion coefficient, and the range is [0, 1], and (R in - R net ) * S is the aspect ratio adjustment amount. It can be seen from formula (3) that after R in and R net are determined, the larger S is, the larger R r is, and the larger the frame expansion range is; when the preset frame expansion coefficient S = 0, R r = R net , this situation is equivalent to the method in Method 2 described in the previous reference Figure 1B , that is, according to the preset input aspect ratio of the trained neural network, the width or length of the first object box is increased to obtain the second object box (the aspect ratio of the second object box is the same as the stretching coefficient). After obtaining the cropped image based on the second object box, it is scaled while maintaining this aspect ratio to be the same as the preset size as the network input image; when S = 1, R r = R net , this situation is equivalent to the method in Method 1 described in the previous reference Figure 1A , and the first object box is directly used to crop the image to be detected without adjusting the width and length, and then the cropped image is scaled and stretched to the preset size of the trained neural network. Therefore, in the embodiments of the present disclosure, in order to achieve the purpose of improving detection accuracy, etc., the range of S can be, for example, (0, 1).
[0074] In step S320-4, the size of the first object box is adjusted based on the stretching coefficient to obtain a second object box.
[0075] Optionally, when the first aspect ratio is less than or equal to the second aspect ratio (R net ≤ R in ), the width value of the first object box is divided by the stretching coefficient (H' in = W in ÷ R r ) as the length value of the second object box (H' in ), and the width value of the first object box (W in ) is used as the width value of the second object box. And when the first aspect ratio is greater than the second aspect ratio (R net > R in ), the length value of the first object box is multiplied by the stretching coefficient (W' in = H in × R r ) as the width value of the second object box, and the length value of the first object box (Hin ) As the long value of the second object box.
[0076] In this way, the collected image to be detected can be cropped based on the second object box to obtain a cropped image, and then the cropped image can be scaled to obtain a network input image with the same preset size (the same aspect ratio and the same width and length values) as the trained neural network. It can be seen that R r can be regarded as the aspect ratio of the cropped image.
[0077] In this embodiment of the present application, the value of S is between 0 and 1. After the value of the preset expansion box coefficient S is determined, the actual R net and R in can be used to adaptively control the degree of expansion of the first object box to obtain a cropped image (with an aspect ratio of R r ), without having to make the cropped image have the same aspect ratio as the preset input aspect ratio of the trained neural network. This can reduce the proportion of background data in the cropped image to provide detection accuracy, and in video stream detection, it can provide a more stable image input to the neural network to achieve more stable key point detection.
[0078] For example, in order to form a contrast with the method two described in the reference Figure 1B , Figure 5A the expansion box process in Figure 1B is reproduced, and Figure 5B a schematic process of the expansion box process according to an embodiment of the present disclosure is shown. Assuming that Figure 5A the aspect ratio of the first object box in Figure 5B and the aspect ratio of the first object box in Figure 5B are both 1 / 1, and the preset input aspect ratio of the trained neural network is both 1 / 2. In Figure 5A , the aspect ratio of the second object box can be set to 1 / 1.5 by setting S instead of being equal to the preset input aspect ratio 1 / 2 of the trained neural network as in Figure 5B . This can reduce the proportion of background data in the cropped image obtained by cropping based on the second object box, thereby increasing the effective pixel ratio. Then in
[0079] It can be seen that, due to the introduction of the preset frame expansion coefficient, the proportion of background data in the cropped image obtained by cropping based on the second object frame can be reduced to increase the effective pixel ratio, thereby improving the feature extraction ability of the trained neural model, making the accuracy of key point detection higher, and in video stream detection, a more stable image input can be provided to the neural network to achieve more stable key point detection.
[0080] The following will be combined with Figures 6A-7 to exemplarily introduce in detail the method for determining the optimal value of the preset frame expansion coefficient. It should be understood that the two methods described below are only exemplary, and other determination methods can also be used to set the preset frame expansion coefficient S, as long as it can reduce the introduction of background data and can better identify the key points in the image to be detected compared to the method without introducing the preset frame expansion coefficient.
[0081] Figure 6A Shows the steps of the first determination method for the preset frame expansion coefficient S. This method mainly uses the test set and determines the optimal preset frame expansion coefficient S based on the test results of the trained neural network. The trained neural network can still be the trained neural network as described above (that is, the training images are all cropped and scaled as Figure 1B shown for training the neural network).
[0082] As Figure 6A shown, in step S610, multiple candidate values of the preset frame expansion coefficient are obtained.
[0083] For example, each candidate of the preset frame expansion coefficient can be between 0 and 1, such as 0.1, 0.2, 0.3,..., 0.9, etc. Of course, it can also be other values between 0 and 1, and the quantity can be set arbitrarily.
[0084] In step S620, based on each candidate value of the preset frame expansion coefficient, the network input image of each test image in the test set is obtained. Each test image in the test set includes an object instance and key point annotations for the object instance.
[0085] Optionally, the test set can also be selected from the COCO dataset.
[0086] Optionally, the process of obtaining the network input image of each test image in the test set based on each candidate value of the preset bounding box expansion coefficient is similar to the process of obtaining the cropped image of the image to be detected and obtaining the network input image based on the cropped image as described above. For example, for the candidate value 0.5, first, for each test image in the test set, determine the corresponding first object bounding box and its size, then determine the corresponding second object bounding box of the test image based on the first aspect ratio of the first object bounding box, the preset input aspect ratio of the trained neural network, and the candidate value 0.5, and obtain the cropped image based on the second object bounding box. Finally, scale the cropped image to obtain the corresponding network input image, where the network input image has the same preset size and preset input aspect ratio as the trained neural network. That is, for the candidate value 0.5, a set of network input images for the test set is generated. Each candidate value corresponds to a set of network input images.
[0087] In step S630, for each candidate value, use the trained neural network to perform key point detection on the network input image of each test image in the test set to obtain the test result corresponding to the test set.
[0088] Optionally, for any candidate value, input the network input image of each test image in the test set obtained as described above into the trained neural network to obtain the test result of each test image for the candidate value. For example, the detected positions of each key point of the object instance of each test image (the positions converted to the test image).
[0089] As an example, for each test image in the test set, the detected positions of each key point of the object instance of the test image can be compared with the positions of the key point annotations to obtain the detection accuracy measure for the test image (for example, the number of key points whose position error between the detected position and the annotated position is within the threshold range), and based on the detection accuracy measures of all test images in the test set, the test result corresponding to the test set for the current candidate value can be obtained.
[0090] In step S640, based on the test result corresponding to the test set for each candidate value, determine the best candidate value for the test set as the preset bounding box expansion coefficient.
[0091] Optionally, when taking each candidate value, the detection accuracy measures of all test images in the test set can be comprehensively considered. For example, the average detection accuracy measure or the number of test images with detection accuracy measures satisfying the threshold condition, etc., to determine for which candidate value the test result corresponding to the test set is the best (for example, the highest average detection accuracy measure, the largest number of test images with detection accuracy measures satisfying the threshold condition, etc.), so as to determine the best candidate value as the preset bounding box expansion coefficient.
[0092] Each step of another method for determining the preset frame expansion coefficient may be as Figure 6B shown.
[0093] As Figure 6B shown, in step S610’, obtain the mapping relationship between the first aspect ratio, the second aspect ratio, and the preset frame expansion coefficient.
[0094] Optionally, the mapping relationship includes a preset functional relationship:
[0095] S = max(0.5 - |R in - R net | * K, 0.2) (4)
[0096] where S is the preset frame expansion coefficient, Rin is the first aspect ratio, and R net is the second aspect ratio. K can be a positive coefficient, and the value of K can be set as a piecewise value, that is, set according to the absolute difference between the first aspect ratio and the second aspect ratio. For example, the smaller the absolute difference between the first aspect ratio and the second aspect ratio, the larger the corresponding K value.
[0097] For example, for the case of R in ≈R net , it can be known that even if the aspect ratio of the cropped image using the first object frame (without frame expansion, S = 1) is directly adjusted to be the same as the preset input aspect ratio of the trained neural network, it will not cause excessive distortion of the object. Therefore, at this time, a larger coefficient K can be further adopted, so that the maximum value in the max function is 0.2, that is, the preset frame expansion coefficient S can be selected as 0.2, and thus the frame expansion range can be smaller.
[0098] In step S620’, based on the mapping relationship, determine the preset frame expansion coefficient corresponding to the first aspect ratio and the second aspect ratio.
[0099] For example, based on the corresponding relationship between the value of K and the absolute difference between the first aspect ratio and the second aspect ratio, after determining the first aspect ratio and the second aspect ratio, the corresponding preset frame expansion coefficient S can be obtained.
[0100] It should be noted that as described in the present disclosure, other methods can also be used to determine the preset frame expansion coefficient. The present disclosure does not make specific limitations on this. The above two are only presented as examples, and other methods for determining the preset frame expansion coefficient based on the first aspect ratio and the second aspect ratio and used to reduce the frame expansion range are within the protection scope of the present field.
[0101] In summary, refer to Figures 3-6BIn the described key point detection method, a trained neural network is used to detect key points of the network input image of the image to be detected. Among them, a preset frame expansion coefficient is introduced in the process of obtaining the network input image, so that the aspect ratio (the second aspect ratio) of the first object frame corresponding to the image to be detected and the preset input aspect ratio (the first aspect ratio) of the trained neural network can be used together to adaptively adjust the aspect ratio of the first object frame, so as to reduce the frame expansion range, thereby reducing the proportion of background data in the network input image. And if it is to detect the video frame image in the video stream, it can also reduce the range in which the inherent jitter of the first object frame between adjacent video frame images is amplified due to the frame expansion process. Therefore, a more stable input can be provided for the neural network, and more stable key point detection can be achieved.
[0102] In reference Figures 3-6B In the described method, the preset frame expansion coefficient adopted is a fixed value and is determined in the actual application process (inference process). That is, in the inference process, there is an additional step of determining the preset frame expansion coefficient, and the process of determining the preset frame expansion coefficient and adjusting the aspect ratio of the first object frame based on it is implemented on the inference side. Through experimental tests, a +1.5 AP gain can be obtained on the COCO2017 eval test set by this method.
[0103] As described above, the trained neural network used in the above process can be an initially trained neural network trained using a general training image set (for example, the COCO dataset). Among them, the network input image of each training image can be obtained in the manner shown in Figure 1B Way 2 (the distortion situation of Way 1 is more serious, so it is not considered here) for training the neural network to obtain the initially trained neural network. The initially trained neural network has the ability to detect key points. That is, the process of introducing a frame expansion coefficient to adaptively adjust the first object frame is not considered in the training process of obtaining the initially trained neural network.
[0104] In some other embodiments, the trained neural network can be an optimized neural network obtained by performing frame expansion enhancement (optimization) on the initially trained neural network. That is, further optimization of the initially trained neural network can also be considered to obtain an optimized neural network, and the optimized neural network can be used as the trained neural network to detect the key points of the image to be detected.
[0105] Therefore, Figure 7 shows a schematic flowchart of each step of performing frame expansion enhancement on the initially trained neural network according to an embodiment of the present disclosure.
[0106] As Figure 7As shown, in step S710, obtain the preset value range of the training bounding box expansion coefficient. For example, the preset value range can be a subset of the value range (0, 1).
[0107] Optionally, it can be based on the test results corresponding to the test set for each candidate obtained in step S630 above Figure 6A to determine the preset value range of the training bounding box expansion coefficient for the training image set, where for each training bounding box expansion coefficient within this preset value range, the test results of the test set are not worse than the test results corresponding to the test set when no bounding box expansion coefficient (S = 0) is used. For example, the preset value range determined based on multiple test results of the test set for different candidate values can be [0, 0.6].
[0108] In step S720, obtain the training image set, where each training image includes an object instance and key point annotations for this object instance.
[0109] Optionally, the training image set can also come from the COCO dataset like the previous test set, so each training image also correspondingly includes an object instance and key point annotations for this object instance.
[0110] In step S730, for each training image, randomly select multiple different training bounding box expansion coefficients from the preset value range.
[0111] Optionally, the training bounding box expansion coefficients can be selected by random sampling. For example, sampling can be performed from this range with a uniform probability S train = U(0, 0.6) to obtain multiple training bounding box expansion coefficients S train . And for each training image, multiple different training bounding box expansion coefficients can be applied.
[0112] In step S740, for each training image, process this training image based on multiple different training bounding box expansion coefficients to obtain multiple network input images with the same preset size, realizing the amplification of the network input image set of the training image set.
[0113] Optionally, for each training image, processing this training image based on multiple different training bounding box expansion coefficients is also similar to the previous one, and can include: for each training bounding box expansion coefficient, determine the first object box and its size corresponding to this training image for this training image, then determine the second object box corresponding to this test image based on the first aspect ratio of the first object box, the preset input aspect ratio of the trained neural network, and this training bounding box expansion coefficient, and obtain a cropped image based on this second object box, and finally scale the cropped image to obtain the corresponding network input image, where the network input image has the same preset size and preset input aspect ratio as the trained neural network.
[0114] Since each training image can be processed based on multiple training bounding box expansion coefficients, multiple network input images can be obtained for each training image. Moreover, since the positions of the key points of the object instances in each training image and the corresponding positions of the key points on each network input image (position transformation relationship) after image processing are also known, the augmentation of the network input image set corresponding to the training image set can be achieved. That is, assuming that M training bounding box expansion coefficients are randomly sampled, finally, the number of images in the augmented network input image set is M times the number of images in the original corresponding network input image set of the training image set.
[0115] In step S750, the initially trained neural network is further optimized using the augmented network input image set to obtain an optimized neural network. Here, the optimized neural network is used as the trained neural network to detect the object key points of the image to be detected.
[0116] That is, not only is the bounding box expansion coefficient applied to process the image to be detected during the inference process, but also the bounding box expansion coefficients within a preset value range are used during the training process. Such a training process can prompt the neural network (e.g., the above-mentioned optimized neural network) to adapt to different degrees of stretching ratios of the object, and improve the robustness of the neural network to different body types and the aspect ratios of the first object bounding boxes. Experiments have also proven that when the bounding box expansion coefficient is introduced in both the inference and training stages, a total gain of more than +2.5 AP value can be obtained on the COCO2017 eval test set.
[0117] The above-described embodiments separately describe the implementation manners of applying the bounding box expansion coefficient only during the inference process and additionally applying the bounding box expansion coefficients within a preset value range during the training process. In other embodiments, the bounding box expansion coefficients within a preset value range can be applied only during the training process to prompt the neural network (e.g., the above-mentioned optimized neural network) to adapt to different degrees of stretching ratios of the object, and improve the robustness of the neural network to different body types and the aspect ratios of the first object bounding boxes. In this case, the embodiments of the present disclosure also provide a training method for a neural network for object key point detection on the training side.
[0118] Figure 8 A flowchart showing a training method for a neural network for object key point detection according to an embodiment of the present disclosure is shown.
[0119] As Figure 8 shown, in step S810, a training image set is obtained, and each training image includes an object instance and key point annotations for the object instance.
[0120] In step S820, for each training image, a plurality of different training bounding box expansion coefficients are randomly selected from a predetermined set.
[0121] In this case, different training bounding box expansion coefficients can be randomly sampled from the set [0, 1]. For each training image, the different training bounding box expansion coefficients will be used to adjust the first object bounding box corresponding to the training image to a second object bounding box with different aspect ratios to obtain different cropped images.
[0122] In step S830, for each training image, the training image is processed based on a plurality of different training bounding box expansion coefficients to obtain a plurality of network input images having the same preset size as the input data of the neural network, thereby realizing the augmentation of the network input image set of the training image set.
[0123] Optionally, similar to step S730, for each training bounding box expansion coefficient of the training image: the first object bounding box corresponding to the training image can be first determined based on the pose of the object instance in the training image, and the first object bounding box corresponds to the object of interest where the object instance in the training image is located; then, the size of the first object bounding box can be adjusted based on the preset size, the bounding box size of the first object bounding box, and the training bounding box expansion coefficient to obtain a second object bounding box (according to formula 3 above); a cropped image of the training image is obtained based on the second object bounding box, and the cropped image is adjusted based on the preset size to obtain a network input image corresponding to the cropped image that is the same as the preset size (both the preset input aspect ratio and the width and length values are the same).
[0124] In step S840, the neural network is trained using the augmented network input image set to obtain a trained neural network.
[0125] In this way, each training image is augmented into a plurality of network input images based on a plurality of different training bounding box expansion coefficients. Therefore, for the object instance in each training image, by being stretched or shortened, it may have changes in height, width, and body shape, which can be regarded as different object instances. Therefore, the neural network can be trained with more network input images.
[0126] By referring to Figure 8 the described training method, since the training bounding box expansion coefficient is introduced to augment the network input image set of the training image set, it can prompt the trained neural network to adapt to different degrees of object stretching ratios and improve the robustness to object instances of different body shapes and the aspect ratios of the first object bounding boxes.
[0127] In addition, after obtaining a specific trained neural network by referring to Figure 8 the described training method, the trained neural network can be used in the inference process, that is, to determine the first object bounding box corresponding to the image to be detected, based onFigure 1B Adjust the second object box or reference in the second manner Figures 3-6B Introduce the preset frame expansion coefficient determination process described therein and adjust based on it to obtain the second object box, and then obtain the cropped image. Then, adjust the cropped image based on the preset size of the trained neural network to obtain a network input image corresponding to the cropped image with the same preset size (both the preset input aspect ratio and the width and length values are the same). Provide the network input image to the trained neural network to obtain the final key point detection result.
[0128] In addition, in some other embodiments, it is possible to refer to Figure 8 After obtaining the trained neural network using the training method described, use the test set to jointly determine the preset frame expansion coefficient to be used in the inference process. That is, in the previous description, the process of determining the preset frame expansion coefficient was set in the inference process on the inference side. Here, the process of determining the preset frame expansion coefficient can be set in the training-related process on the training side, which can speed up the operation speed in the actual inference process and improve the inference efficiency.
[0129] For example, similarly, similar to the previous reference to Figure 6A After obtaining the trained neural network, multiple candidate values of the preset frame expansion coefficient can be randomly selected from a predetermined set (for example, (0, 1)); based on each candidate value of the preset frame expansion coefficient, obtain the network input image of each test image in the test set, and each test image in the test set includes an object instance and the key point annotation for that object instance; for each candidate value, use the trained neural network to perform key point detection on the network input image of each test image in the test set to obtain the test result corresponding to the test set; and based on the test result corresponding to each candidate value for the test set, determine the best candidate value for the test set as the preset frame expansion coefficient for key point detection of the object in the image to be detected.
[0130] More details about determining the preset frame expansion coefficient have been described in the reference to Figure 6A and will not be repeated here.
[0131] Since the size of the actual image to be detected cannot be known during the training process, the method described in Figure 6B of determining the preset frame expansion coefficient based on the aspect ratio-related mapping relationship of the first object box corresponding to the image to be detected is not adopted.
[0132] In this way, during the actual inference process, the first object box corresponding to the image to be detected is determined, and the second object box is obtained by adjusting it based on the preset box expansion coefficient known in the training stage. Then, the cropped image is obtained, and the cropped image is adjusted based on the preset size of the trained neural network to obtain a network input image corresponding to the cropped image with the same size as the preset size (both the preset input aspect ratio and the width and length values are the same). The network input image is provided to the trained neural network to obtain the final key point detection result. Based on this, since the process of determining the optimal value of the preset box expansion coefficient is placed in the training process, it can be omitted during the actual inference process, thereby improving the inference efficiency.
[0133] According to another aspect of the present disclosure, there is also provided a device for detecting object key points.
[0134] Figure 9 The structural block diagram of the device for detecting object key points according to an embodiment of the present disclosure is shown.
[0135] As Figure 9 shown, the key point detection device 900 may include: a determination module 910, an object box adjustment module 920, a cropping module 930, and a detection module 940.
[0136] The determination module 910 is configured to determine the first object box corresponding to the image to be detected based on the pose of the object in the image to be detected, and the first object box corresponds to the region of interest where the object is located in the image to be detected.
[0137] The object box adjustment module 920 is configured to adjust the size of the first object box based on the preset size of the input data of the trained neural network, the box size of the first object box, and the preset box expansion coefficient to obtain a second object box, where the preset box expansion coefficient makes the aspect ratio of the second object box fall between the aspect ratio of the preset size and the aspect ratio of the first object box, and can be, for example, a value between 0 and 1.
[0138] The cropping module 930 is configured to obtain the cropped image of the image to be detected based on the second object box, and adjust the cropped image based on the preset size to obtain a network input image corresponding to the cropped image, and the size of the network input image is equal to the preset size.
[0139] The detection module 940 is configured to detect the key points of the object in the network input image based on the trained neural network.
[0140] More details of each module have been described in detail in the key point detection method with reference to Figure 3 above, so they will not be repeated here.
[0141] By referring to Figure 9The detection device for the key points of the described object adjusts the size of the first object frame that frames the object in the image to be detected based on the preset size of the input data of the trained neural network, the frame size of the first object frame, and the preset frame expansion coefficient, so that the first object frame can be appropriately enlarged to be as close as possible to the preset input aspect ratio of the trained neural network (facilitating scaling and avoiding distortion) and including as little background data as possible (improving the accuracy of feature extraction), and the network input image is obtained through the cropping and scaling process. The trained neural network performs key point detection on the network input image, which can improve the detection accuracy. On the other hand, in video stream detection, by appropriately enlarging the first object frame, compared with the reference Figure 1B The described method 2 can reduce the frame expansion range (the increase amplitude from the first object frame to the second object frame), and thus can reduce the range in which the jitter of the first object frame between adjacent frames is amplified. Therefore, in video stream detection, a more stable image input can be provided to the neural network to achieve more stable key point detection.
[0142] Furthermore, as Figure 10 shown, the object frame adjustment module 920 may include an aspect ratio determination sub-module 920-1, a stretching coefficient determination sub-module 920-2, and an adjustment sub-module 920-3.
[0143] The aspect ratio determination sub-module 920-1 is used to determine the aspect ratio corresponding to the preset size as the first aspect ratio based on the width value and length value corresponding to the preset size of the input data of the trained neural network, and determine the aspect ratio of the first object frame as the second aspect ratio based on the width value and length value corresponding to the frame size of the first object frame.
[0144] The stretching coefficient determination sub-module 920-2 is used to determine the stretching coefficient of the first object frame based on the first aspect ratio, the second aspect ratio, and the preset frame expansion coefficient.
[0145] For example, the product of the difference between the second aspect ratio and the first aspect ratio and the preset frame expansion coefficient can be determined as the aspect ratio adjustment amount, and then the stretching coefficient (Rr) of the first object frame is determined as the sum of the first aspect ratio and the aspect ratio adjustment amount.
[0146] The adjustment sub-module 920-3 is used to adjust the size of the first object frame based on the stretching coefficient to obtain the second object frame.
[0147] For example, in the case where the first aspect ratio is less than or equal to the second aspect ratio, the width value of the first object frame is divided by the stretching coefficient as the length value of the second object frame, and the width value of the first object frame is used as the width value of the second object frame; and in the case where the first aspect ratio is greater than the second aspect ratio, the length value of the first object frame is multiplied by the stretching coefficient as the width value of the second object frame, and the length value of the first object frame is used as the length value of the second object frame.
[0148] More details of the operations of each sub-module have been described above, so they will not be repeated here.
[0149] In addition, although not shown, the object key point detection device 900 may further include a frame expansion determination module for determining a preset frame expansion coefficient. Additionally, the object key point detection device 900 may further include a model optimization module (not shown) for further optimizing (frame expansion enhancement) the initially trained neural network to obtain an optimized neural network, where the optimized neural network is used as the trained neural network to detect object key points in the image to be detected.
[0150] The specific operations performed by the frame expansion determination module are the same as those in the steps of the method described above with reference to Figure 6A-6B The specific operations performed by the model optimization module are the same as those in the steps of the method described above with reference to Figure 7 So they will not be repeated here.
[0151] Therefore, due to the introduction of the preset frame expansion coefficient, the proportion of background data in the cropped image obtained by cropping based on the second object frame can be reduced to increase the effective pixel ratio, thereby enhancing the feature extraction ability of the trained neural model and making the accuracy of key point detection higher. In addition, not only is the frame expansion coefficient applied in the inference process, but the optimization module also uses the frame expansion coefficient within a preset value range during the training process. Such a training process can prompt the neural network (e.g., the optimized neural network above) to adapt to the stretching ratios of objects of different degrees, improving the robustness of the neural network to different body shapes and the aspect ratios of the first object frames.
[0152] According to another aspect of the present disclosure, there is also provided a training device for a neural network for object key point detection.
[0153] Figure 11 FIG. shows a training device 1000 for a neural network for object key point detection according to an embodiment of the present disclosure.
[0154] The training device 1100 includes an acquisition module 1110, a selection module 1120, a processing module 1130, and a training module 1140.
[0155] The acquisition module 1110 is configured to acquire a training image set, where each training image includes an object instance and key point annotations for the object instance.
[0156] The selection module 1120 is configured to randomly select a plurality of different training frame expansion coefficients from a predetermined set (e.g., (0, 1)) for each training image.
[0157] The processing module 1130 is configured to process each training image based on multiple different training bounding box expansion coefficients, to obtain multiple network input images having the same preset size as the input data of the neural network, thereby achieving the augmentation of the network input image set of the training image set.
[0158] The training module 1140 is configured to train the neural network by using the augmented network input image set, to obtain the trained neural network.
[0159] More details of the operations performed by each module are similar to the operations in the steps of the method described above with reference to Figure 8 and thus will not be repeated here.
[0160] By referring to Figure 11 the described training device, since the training bounding box expansion coefficients are introduced to augment the network input image set of the training image set, it can prompt the trained neural network to adapt to the stretching ratios of objects at different levels, and improve the robustness to object instances of different body types and the aspect ratios of the first object bounding boxes.
[0161] In addition, although not shown, the training device 1100 may include a bounding box determination module for determining the preset bounding box expansion coefficients, that is, implementing the determination process of the preset bounding box expansion coefficients used in the actual inference process in the training device instead of in the detection device as described above, so as to improve the inference efficiency during actual application. The specific process has also been described above with reference to Figure 8 and is not repeated here.
[0162] According to another aspect of the present disclosure, a computing device is also disclosed.
[0163] Figure 12 FIG. shows a schematic block diagram of a computing device 1200 according to an embodiment of the present disclosure.
[0164] As Figure 12 shown, the computing device 1200 includes a processor, a memory, a network interface, an input device, and a display screen connected through a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the terminal stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor can implement various operations described in the steps of the object key point detection method and the training method of the neural network for object key point detection as described above. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor can execute the same various operations described in the steps of the object key point detection method and the training method of the neural network for object key point detection.
[0165] For example, the method for detecting object key points may include: determining a first object bounding box corresponding to the image to be detected based on the pose of the object in the image to be detected, where the first object bounding box corresponds to the region of interest where the object in the image to be detected is located; adjusting the size of the first object bounding box based on a preset size of the input data of the trained neural network, the bounding box size of the first object bounding box, and a preset bounding box expansion coefficient, to obtain a second object bounding box, where the preset bounding box expansion coefficient makes the aspect ratio of the second object bounding box fall between the aspect ratio corresponding to the preset size and the aspect ratio of the first object bounding box, for example, a value between 0 and 1; obtaining a cropped image of the image to be detected based on the second object bounding box, and adjusting the cropped image based on the preset size to obtain a network input image corresponding to the cropped image, where the size of the network input image is equal to the preset size; and detecting object key points in the network input image based on the trained neural network. More details of each step have been described in detail above, so they will not be repeated here.
[0166] The processor may be an integrated circuit chip with the ability to process signals. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc., and may be of the X84 architecture or the ARM architecture.
[0167] The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. It should be noted that the memory for the methods described in the present disclosure is intended to include, but is not limited to, these and any other suitable types of memory.
[0168] The display screen of the computing device may be a liquid crystal display screen or an electronic ink display screen. The input device of the computing device may be a touch layer covering the display screen, or may be a button, a trackball or a touchpad provided on the terminal housing, or may also be an external keyboard, touchpad or mouse, etc.
[0169] The computing device may be a terminal or a server. Among them, the terminal may include but is not limited to: smart phones, tablet computers, laptop computers, desktop computers, smart TVs, etc.; various clients (applications, APPs) can run inside the terminal, such as multimedia playback clients, social clients, browser clients, information flow clients, education clients, and so on. The server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, and so on.
[0170] According to another aspect of the present disclosure, there is also provided a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to perform various operations described in each step of the object key point detection method and the neural network training method for object key point detection as described above.
[0171] According to still another aspect of the present disclosure, there is also provided a computer program product including a computer program, which when executed by a processor implements various operations described in each step of the object key point detection method and the neural network training method for object key point detection as described above.
[0172] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the methods and apparatuses according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains at least one executable instruction for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0173] The exemplary embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art should understand that various modifications and combinations can be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.
Claims
1. A method for detecting object key points, comprising: Determining a first object box corresponding to the image to be detected based on the pose of the object in the image to be detected, where the first object box corresponds to the region of interest where the object in the image to be detected is located; Adjusting the size of the first object box based on the preset size of the input data of the trained neural network, the box size of the first object box, and a preset box expansion coefficient to obtain a second object box, where the preset box expansion coefficient makes the aspect ratio of the second object box fall between the aspect ratio corresponding to the preset size and the aspect ratio of the first object box, and the preset size is a fixed size of the input data that can be processed by the trained neural network; Obtaining a cropped image of the image to be detected based on the second object box, and adjusting the cropped image based on the preset size to obtain a network input image corresponding to the cropped image, where the size of the network input image is equal to the preset size; and Detecting object key points in the network input image based on the trained neural network.
2. The detection method according to claim 1, wherein The adjusting the first object box based on the preset size of the input data of the neural network, the box size of the first object box, and the preset box expansion coefficient to obtain a second object box includes: Determining the aspect ratio corresponding to the preset size based on the width value and length value corresponding to the preset size of the input data of the trained neural network as a first aspect ratio; Determining the aspect ratio of the first object box based on the width value and length value corresponding to the box size of the first object box as a second aspect ratio; Determining the stretching coefficient of the first object box based on the first aspect ratio, the second aspect ratio, and the preset box expansion coefficient; and Adjusting the size of the first object box based on the stretching coefficient to obtain a second object box.
3. The detection method according to claim 2, wherein, The determining the stretching coefficient of the first object box based on the first aspect ratio, the second aspect ratio, and the preset box expansion coefficient includes: Determining the product of the difference between the second aspect ratio and the first aspect ratio and the preset box expansion coefficient as the aspect ratio adjustment amount; Determining the stretching coefficient of the first object box as the sum of the first aspect ratio and the aspect ratio adjustment amount.
4. The detection method according to claim 2, wherein, The adjusting the first object box based on the stretching coefficient to obtain a second object box includes: When the first aspect ratio is less than or equal to the second aspect ratio, dividing the width value of the first object box by the stretching coefficient as the length value of the second object box, and using the width value of the first object box as the width value of the second object box; and When the first aspect ratio is greater than the second aspect ratio, multiplying the length value of the first object box by the stretching coefficient as the width value of the second object box, and using the length value of the first object box as the length value of the second object box.
5. The detection method according to claim 4, wherein, The preset box expansion coefficient is determined by the following method: Obtaining the mapping relationship among the first aspect ratio, the second aspect ratio, and the preset box expansion coefficient; And Based on the mapping relationship, determining the preset box expansion coefficient corresponding to the first aspect ratio and the second aspect ratio.
6. The detection method according to claim 5, wherein, The mapping relationship includes a preset functional relationship: S = max(0.5 - |R in - R net | * K, 0.2), Where S is the preset frame expansion coefficient, Rin is the first aspect ratio, Rnet is the second aspect ratio, K is a preset positive number, and according to R in -R net the absolute difference has piecewise values, and the larger the absolute difference, the larger K is.
7. The detection method according to claim 4, wherein The preset bounding box expansion coefficient is determined by the following method: Obtain multiple candidate values of the preset bounding box expansion coefficient; Based on each candidate value of the preset bounding box expansion coefficient, obtain the network input image of each test image in the test set, where each test image in the test set includes an object instance and key point annotations for the object instance; For each candidate value, use the trained neural network to perform key point detection on the network input image of each test image in the test set, and obtain the test result corresponding to the test set; and Based on the test results corresponding to the test set for each candidate value, determine the best candidate value for the test set as the preset bounding box expansion coefficient.
8. The detection method according to claim 7, wherein, The trained neural network is an optimized neural network obtained by performing bounding box enhancement on the initially trained neural network, wherein, the performing bounding box enhancement on the initially trained neural network includes: Obtain the preset value range of the training bounding box expansion coefficient; Obtain a training image set, where each training image includes an object instance and key point annotations for the object instance; For each training image, randomly select multiple different training bounding box expansion coefficients from the preset value range; For each training image, process the training image based on multiple different training bounding box expansion coefficients to obtain multiple network input images of the same preset size, thereby realizing the amplification of the network input image set of the training image set; and Use the amplified network input image set to further optimize the initially trained neural network to obtain an optimized neural network, where the optimized neural network is used as the trained neural network to detect the object key points in the image to be detected.
9. The detection method according to claim 8, wherein, The obtaining the preset value range of the training bounding box expansion coefficient includes: Based on the test results corresponding to the test set for each candidate, determine the preset value range of the training bounding box expansion coefficient for the training image set, wherein, for each training bounding box expansion coefficient within the preset value range, the test result of the test set is better than the test result corresponding to the test set when no bounding box expansion coefficient is used.
10. The detection method according to claim 1, wherein, The determining the first object box corresponding to the image to be detected based on the pose of the object in the image to be detected includes: Call an object detector to determine the first object box; or In the case of key point detection for consecutive video frame images: Generate the first object box corresponding to the current video frame image based on the key point detection result of the previous video frame image, or call an object detector to determine the first object box of the current video frame.
11. A training method for a neural network for object key point detection, including: Obtain a training image set, where each training image includes an object instance and key point annotations for the object instance; For each training image, randomly select multiple different training bounding box expansion coefficients from a predetermined set; For each training image, the training image is processed based on a plurality of different training bounding box expansion coefficients to obtain a plurality of network input images having the same preset size as the input data of the neural network, thereby realizing the expansion of the network input image set of the training image set. The preset size is a fixed size of the input data that can be processed by the neural network; And The neural network is trained using the expanded network input image set to obtain a trained neural network, wherein, For each training image, the process of processing the training image based on a plurality of different training bounding box expansion coefficients to obtain a plurality of network input images having the same preset size as the input data of the neural network includes: for each training bounding box expansion coefficient of the training image, Determine a first object bounding box corresponding to the training image based on the pose of the object instance in the training image. The first object bounding box corresponds to the object of interest where the object instance in the training image is located; Adjust the size of the first object bounding box based on the preset size, the size of the first object bounding box, and the training bounding box expansion coefficient to obtain a second object bounding box; Obtain a cropped image of the training image based on the second object bounding box, and adjust the cropped image based on the preset size to obtain a network input image corresponding to the cropped image and having the same preset size.
12. The training method according to claim 11, wherein, After obtaining the trained neural network, the training method further includes: Randomly select a plurality of candidate values of the preset bounding box expansion coefficient from the predetermined set; Based on each candidate value of the preset bounding box expansion coefficient, obtain network input images of each test image in the test set. Each test image in the test set includes an object instance and key point annotations for the object instance; For each candidate value, use the trained neural network to perform key point detection on the network input images of each test image in the test set to obtain a test result corresponding to the test set; and Based on the test results corresponding to the test set for each candidate value, determine the best candidate value for the test set as the preset bounding box expansion coefficient for key point detection of the object in the image to be detected.
13. An object key point detection device, comprising: A determination module, configured to determine a first object bounding box corresponding to the image to be detected based on the pose of the object in the image to be detected. The first object bounding box corresponds to the region of interest where the object in the image to be detected is located; An object bounding box adjustment module, configured to adjust the size of the first object bounding box based on the preset size of the input data of the trained neural network, the size of the first object bounding box, and the preset bounding box expansion coefficient to obtain a second object bounding box. The preset bounding box expansion coefficient makes the aspect ratio of the second object bounding box fall between the aspect ratio of the preset size and the aspect ratio of the first object bounding box. The preset size is a fixed size of the input data that can be processed by the trained neural network; A cropping module, configured to obtain a cropped image of the image to be detected based on the second object box, and adjust the cropped image based on the preset size to obtain a network input image corresponding to the cropped image, where the size of the network input image is equal to the preset size; and A detection module, configured to detect key points of an object in the network input image based on a trained neural network.
14. A training device for a neural network for object key point detection, comprising: An acquisition module, configured to acquire a training image set, where each training image includes an object instance and key point annotations for the object instance; A selection module, configured to randomly select a plurality of different training bounding box expansion coefficients from a predetermined set for each training image; An image processing module, configured to process each training image based on a plurality of different training bounding box expansion coefficients to obtain a plurality of network input images having the same preset size as the input data of the neural network, so as to implement the augmentation of the network input image set of the training image set, where the preset size is a fixed size of the input data that can be processed by the neural network; And A training module, configured to train the neural network by using the augmented network input image set to obtain a trained neural network, where the image processing module is configured to: for each training bounding box expansion coefficient of the training image, determine a first object box corresponding to the training image based on the pose of the object instance in the training image, where the first object box corresponds to the object of interest where the object instance in the training image is located; adjust the size of the first object box based on the preset size, the box size of the first object box, and the training bounding box expansion coefficient to obtain a second object box; obtain a cropped image of the training image based on the second object box, and adjust the cropped image based on the preset size to obtain a network input image corresponding to the cropped image and having the same size as the preset size.
15. A computing device, comprising: A processor; A memory, storing a computer program thereon, where when the computer program is executed by the processor, the processor is caused to execute the steps of the object key point detection method according to any one of claims 1-10 and the training method according to any one of claims 11-12.
Citation Information
Patent Citations
Image processing method and device
CN112037143A
Image mask generation using a deep neural network
US20210042928A1