Three-dimensional Pose Detection Method, System, Device and Storage Medium for Dynamic Interactive Objects

By using a single main viewing camera system and a second learning network, the problem of three-dimensional pose detection of multiple dynamic interactive objects is solved, and accurate and real-time detection effects are achieved in complex occlusion and diverse behavior scenarios.

CN118397613BActive Publication Date: 2025-06-27WESTLAKE UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410176441.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-08
Publication Date
2025-06-27
Estimated Expiration
2044-02-08

AI Technical Summary

Technical Problem

The prior art is difficult to accurately detect the three-dimensional poses of multiple dynamic interactive objects, especially when there are complex occlusion relationships and diverse behaviors between objects.

Method used

Using a single main perspective camera system, the three-dimensional position is integrated by training the two-dimensional position of the object limb key points in the two-dimensional image, and a three-dimensional data set is constructed and the second learning network is trained to achieve three-dimensional position detection and dynamic three-dimensional posture determination of the target limb key points.

Benefits of technology

In the process of free interaction between multiple objects, the three-dimensional position of the key points of the target object's limb is accurately and quickly identified and detected, and the dynamic three-dimensional posture is determined in real time, which is suitable for scenes of complex occlusion and diverse behaviors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118397613B_ABST
    Figure CN118397613B_ABST
Patent Text Reader

Abstract

The present application provides a method, system, device and storage medium for detecting the three-dimensional pose of a dynamic interaction object. The method includes: based on the two-dimensional positions of the limb key points of at least one object in the two-dimensional images for training from the master-slave perspectives, processing and integrating to obtain the three-dimensional positions of the limb key points in the first group, and accordingly constructing a three-dimensional data set to train a second learning network; based on the two-dimensional image to be detected with a main perspective including multiple objects, extracting the region of interest containing the target object; based on the region of interest, using the second learning network to detect the three-dimensional positions of the limb key points of the target object; and thereby determining the dynamic three-dimensional pose of the target object. By using a simple shooting system with a single main perspective, during the process of multiple objects freely and dynamically interacting with each other, although the occlusion relationship is complex and changeable, it is still possible to accurately and quickly identify and detect the three-dimensional positions of the limb key points of any target object at each shooting moment, and determine the dynamic three-dimensional pose of the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of object space motion detection in machine vision, and more particularly to a three-dimensional pose detection method, system, device and medium for dynamically interacting objects. Background Art

[0002] People have attempted to introduce machine vision technology into the analysis of the dynamic interaction process of free-state object groups, such as but not limited to animal social activities, robot interaction processes, etc., where each object can spontaneously and freely perform behaviors.

[0003] In order to analyze and study the behavioral characteristics of these dynamically interacting objects, it is first necessary to detect the three-dimensional pose information of each interacting individual. However, the three-dimensional pose detection based on depth cameras can only obtain partial and simple pose information. Although some existing tools can provide dynamic analysis of the limbs of multiple individuals, they require restrictions on individual actions, and thus cannot be applied to application scenarios where each object freely makes interactive behaviors. The methods for detecting free-state individuals are currently limited to single individuals, and there are many technical obstacles in the state analysis of multiple dynamically interacting objects. For example, during the interaction process of multiple objects, complex occlusion relationships will occur between each object. Because the objects have behavioral freedom, this occlusion relationship will also change dynamically and frequently. Further, the interactive behaviors of multiple objects are more diverse than those of a single individual, making it difficult to accurately identify the dynamic three-dimensional positions of the key points of each object. The research on the interaction process of multiple objects not only requires accurately identifying the dynamic three-dimensional poses of each object, but also has strict requirements for the timeliness of identification. Summary of the Invention

[0004] This application is proposed to solve the above technical problems in the prior art.

[0005] This application aims to provide a three-dimensional pose detection method, system, device and computer program product for dynamically interacting objects. Using a simple shooting system with a single main-view camera, during the process where multiple objects freely and dynamically interact with each other, despite the complex and variable occlusion relationships between them, it can still accurately and quickly (including in real time) identify and detect the three-dimensional positions of the limb key points of any target object at each shooting moment, so as to determine the dynamic three-dimensional pose of the target object during this process.

[0006] According to the first aspect of the present application, a method for three-dimensional pose detection of dynamically interacting objects is provided, wherein multiple objects interact dynamically with each other. The method includes the following steps. Based on the two-dimensional positions of the limb key points of at least one object in the training two-dimensional images of the first set of the main view and multiple side views, process and integrate to obtain the three-dimensional positions of the limb key points of the at least one object in the first set. Use the region of interest containing the at least one object and the three-dimensional positions of the limb key points of the at least one object in the training two-dimensional image of the first set of the main view to construct a three-dimensional data set. Use the constructed three-dimensional data set to train a second learning network. Based on the two-dimensional image to be detected with multiple objects in the main view, extract the region of interest containing the target object. Based on the extracted region of interest containing the target object, use the trained second learning network to detect the three-dimensional positions of the limb key points of the target object. Based on the detected three-dimensional positions of the limb key points of the target object at different times, determine the dynamic three-dimensional pose of the target object.

[0007] According to the second aspect of the present application, a system for three-dimensional pose detection of dynamically interacting objects is provided. The system includes a single top view camera and at least one processor. The single top view camera is configured to capture two-dimensional images at different times with multiple objects in the main view as the two-dimensional images to be detected. The at least one processor is configured to execute the method for three-dimensional pose detection of dynamically interacting objects according to various embodiments of the present application. The method includes the following steps. Based on the two-dimensional positions of the limb key points of at least one object in the training two-dimensional images of the first set of the main view and multiple side views, process and integrate to obtain the three-dimensional positions of the limb key points of the at least one object in the first set. Use the region of interest containing the at least one object and the three-dimensional positions of the limb key points of the at least one object in the training two-dimensional image of the first set of the main view to construct a three-dimensional data set. Use the constructed three-dimensional data set to train a second learning network. Based on the two-dimensional image to be detected with multiple objects in the main view, extract the region of interest containing the target object. Based on the extracted region of interest containing the target object, use the trained second learning network to detect the three-dimensional positions of the limb key points of the target object. Based on the detected three-dimensional positions of the limb key points of the target object at different times, determine the dynamic three-dimensional pose of the target object.

[0008] According to a third aspect of the present application, a three-dimensional pose detection device for a dynamic interaction object is provided. The device includes at least one processor configured to execute the three-dimensional pose detection method for a dynamic interaction object according to various embodiments of the present application. The method includes the following steps. Based on the two-dimensional positions of the limb key points of at least one object in the training two-dimensional images of the main view and multiple secondary views of the first group, processing and integration are performed to obtain the three-dimensional positions of the limb key points of the at least one object in the first group. Using the region of interest containing the at least one object and the three-dimensional positions of the limb key points of the at least one object in the training two-dimensional image of the main view of the first group, a three-dimensional data set is constructed. The constructed three-dimensional data set is used to train a second learning network. Based on the two-dimensional image to be detected of the main view containing multiple objects, the region of interest containing the target object is extracted. Based on the extracted region of interest containing the target object, the trained second learning network is used to detect the three-dimensional positions of the limb key points of the target object. Based on the detected three-dimensional positions of the limb key points of the target object at different times, the dynamic three-dimensional pose of the target object is determined.

[0009] According to a fourth aspect of the present application, a computer-readable storage medium is provided, on which an executable program is stored. When the executable program is executed by a processor, the three-dimensional pose detection method for a dynamic interaction object according to various embodiments of the present application is implemented.

[0010] Using the three-dimensional pose detection method, system, device and medium for a dynamic interaction object according to various embodiments of the present application, based on the two-dimensional positions of the limb key points of at least one object in the training two-dimensional images of the main view and multiple secondary views of the first group, processing and integration are performed to obtain the three-dimensional positions of the limb key points of the at least one object in the first group; the region of interest containing the at least one object and the three-dimensional positions of the limb key points of the at least one object in the training two-dimensional image of the main view of the first group are used to construct a three-dimensional data set. The three-dimensional data set constructed in this way reflects the mapping relationship between the region of interest where the object is located in the two-dimensional image of the main view and the three-dimensional positions of the limb key points of the object. By using it to train the second learning network, the second learning network can learn the mapping relationship. When actually detecting the dynamic three-dimensional positions of the limb key points of a target object dynamically interacting with other objects, using the two-dimensional images to be detected of the main view of multiple objects obtained by a simple shooting system (of a camera) with a single main view, during the process of multiple objects freely and dynamically interacting with each other, despite the complex and changeable occlusion relationships between them, it is still possible to accurately and quickly (including in real time) identify and detect the three-dimensional positions of the limb key points of any target object at each shooting moment, so as to determine the dynamic three-dimensional pose of the target object during this process. Description of the Drawings

[0011] The features, advantages, and technical and industrial significance of exemplary embodiments of the present invention will be described below with reference to the accompanying drawings, in which like numerals refer to like elements, and wherein,

[0012] Figure 1 A flowchart showing a method for detecting the three-dimensional pose of a dynamic interactive object according to an embodiment of the present application.

[0013] Figure 2 An example of a training two-dimensional image with Drosophila as the object is shown, in which a two-dimensional position distribution diagram of the limb key points of a single Drosophila is shown.

[0014] Figure 3 A comparison schematic diagram showing two-dimensional images of a main view and multiple side views according to an embodiment of the present application, the two-dimensional positions of the limb key points of the object therein, and the three-dimensional positions of the limb key points of the corresponding object obtained by processing and integrating them.

[0015] FIG. 4(a) shows a block diagram of a first learning network according to an embodiment of the present application, which is used to detect the two-dimensional positions of the limb key points of at least one object based on a first set of training two-dimensional images of a main view and multiple side views.

[0016] FIG. 4(b) shows a block diagram of a second learning network according to an embodiment of the present application, which is used to detect the three-dimensional positions of the limb key points of the target object based on the extracted region of interest containing the target object.

[0017] Figure 5 A flowchart showing processing and integrating the two-dimensional positions of the limb key points of at least one object in a first set of training two-dimensional images of a main view and multiple side views according to an embodiment of the present application to obtain the three-dimensional positions of the first set of limb key points of the at least one object.

[0018] Figure 6 An example of a training two-dimensional image of the top view of multiple Drosophila according to an embodiment of the present application is shown, in which the situation where the Drosophila as the target object is occluded by other Drosophila in front is shown.

[0019] Figure 7 Shown according to Figure 6 An example of a training two-dimensional image of the side view of the Drosophila to be removed in the occluded situation of the Drosophila as the target object shown in

[0020] Figure 8 A diagram showing the result of centralizing and transforming the region of interest containing the target object in the main view according to an embodiment of the present application.

[0021] Figure 9A flowchart showing the construction of a three-dimensional dataset for training the second learning network according to an embodiment of the present application.

[0022] Figure 10 A flowchart showing the detection of the three-dimensional positions of the limb key points of the target object by using the trained second learning network based on the extracted region of interest containing the target object according to an embodiment of the present application.

[0023] Figure 11 A comparison diagram showing the region of interest containing the target object after centralization transformation and the two-dimensional projection of the three-dimensional positions of the limb key points of the target object after superimposing the centralization transformation process according to an embodiment of the present application.

[0024] Figure 12 A schematic block diagram showing a three-dimensional pose detection system for dynamic interaction objects according to an embodiment of the present application. Detailed Description of the Embodiment

[0025] Hereinafter, specific embodiments will be described in detail with reference to the accompanying drawings. In each drawing, the same or corresponding elements are denoted by the same reference numerals, and repeated descriptions are omitted if necessary to make the description clear. Note that the arrows in the flowcharts in the present application are not intended to limit the execution order of each step. As long as the logic of each step does not conflict, they can be executed in parallel, or the order can be swapped, or one step can be split into several steps for execution, or several steps can be combined for execution, etc. The present application does not make any limitation in this regard. Unless otherwise defined, the technical terms or scientific terms used in the present application should have the ordinary meaning understood by those of ordinary skill in the field to which the present application belongs. The "first", "second" and similar terms used in the present application do not indicate any order, quantity or importance, but are only used to distinguish different components. The terms "including" or "comprising" and the like mean that the elements or items appearing before the term cover the elements or items listed after the term and their equivalents, without excluding other elements or items. The terms "connected" or "coupled" and the like are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms "upper", "lower", "left", "right" and the like are only used to represent relative position relationships, and when the absolute position of the object being described changes, the relative position relationship may also change accordingly.

[0026] Figure 1The flowchart shows a three-dimensional pose detection method for dynamic interaction objects according to an embodiment of the present application. The moving object here can be any one of animals (such as insects), movable facilities, and robots. The animals here can include, but are not limited to, insects, small fish, small amphibians, small reptiles, etc. The robots can also include logistics robots, etc. In some embodiments, the object includes any one of small animals, small movable facilities, and small robots. The so-called "small" is relatively small in scale with respect to the shooting field of view. For example, small animals can include, but are not limited to, insects, small fish, small amphibians, small reptiles, etc. For a panoramic field of view, small robots can also include small logistics robots that travel back and forth on the same factory floor. Multiple objects can freely perform dynamic interactions with each other, so that the positions of each object can change dynamically, and the spatial overlap and occlusion relationships between the objects can also change dynamically. According to different dynamic interaction objects, the three-dimensional pose detection method is used for different analysis purposes in different application scenarios. For example, for a group of fruit flies, by allowing them to freely interact and move within a certain space and detecting the dynamic changes in the three-dimensional poses of each fruit fly during the process, the social behavior of the group can be analyzed. Another example is for small logistics robots. Each logistics robot travels back and forth on the factory floor by running an autonomous route dynamic planning function. By detecting the dynamic changes in the three-dimensional poses of each logistics robot during transportation, it can be verified whether they have achieved good mutual avoidance and optimized transportation routes, etc.

[0027] The following takes the example of fruit flies as the object to illustrate the three-dimensional pose detection method for dynamic interaction objects. However, it should be noted that this method can be flexibly applied to other scenarios and other types of small objects. The three-dimensional pose detection method includes the following steps.

[0028] In step 101, based on the two-dimensional positions of the limb key points of at least one object in the training two-dimensional images of the first group of main viewpoints and multiple sub-viewpoints, processing and integration are performed to obtain the three-dimensional positions of the limb key points of the at least one object in the first group. The training two-dimensional images of the main viewpoint and the three-dimensional positions obtained by processing and integration will be used to subsequently construct a three-dimensional data set to train a second learning network. This second learning network is used to detect the three-dimensional positions of the limb key points of the target object based on the extracted region of interest (such as, but not limited to, image patches) containing the target object, which will be described in detail below and will not be elaborated here.

[0029] In the following text, the top view is used as the main perspective and the side view is used as the lateral perspective for illustration. However, it should be noted that the present application is not limited thereto. For the case where multiple objects are widely distributed on a flat surface, such as the fruit flies and logistics robots in the above text, most of the spatial position information of the objects can be reflected in the two-dimensional image obtained by taking the top view as the main perspective. Compared with the spatial position information obtained from the two-dimensional images taken from other perspectives, it is more abundant. Moreover, in the two-dimensional image obtained by taking the top view as the main perspective, the distribution of each object is sparser than that in the two-dimensional images taken from other perspectives, which helps to efficiently extract the region of interest where the target object is located while introducing as little information of other objects as possible. In some embodiments, due to the different spatial characteristics of the object activities, different main perspectives can be adopted accordingly, as long as the two-dimensional image obtained by the perspective can reflect most of the spatial position information of the object and the distribution of each object is sparser.

[0030] To obtain sufficient two-dimensional images for training, a top view camera and multiple side view cameras can be used to simultaneously capture multiple interacting fruit flies. Each fruit fly interacts with other fruit flies on the activity plane. The timing of each camera needs to be synchronized so as to synchronously obtain the captured images from each perspective at each time point. In this way, multiple perspectives of captured images can be obtained each time. By selecting a certain number of captured images and manually annotating the position coordinates of the key points of each limb of each fruit fly, a set of two-dimensional images with annotated limb key points can be obtained.

[0031] In some embodiments, the key points of the limbs in the captured images can be manually annotated to identify their position coordinates. When annotating, only the visible key points of the fruit fly can be annotated, such as Figure 2 shown, the lines connecting the limb key points are only used to show the relationship between these limb key points.

[0032] For the top view captured image and side view captured image of the fruit fly at a certain moment collected synchronously and the position coordinates of each limb key point annotated therein, they can be processed and integrated to obtain the three-dimensional position of the limb key points of the fruit fly at that moment. As Figure 3 shown, the middle image block in the first column of the two columns of two-dimensional images (corresponding to the region of interest) comes from the top view captured image of the fruit fly at a certain moment, and the other image blocks (corresponding to the region of interest) come from the 6 side view captured images of the fruit fly at the same moment. The dot matrix connection line on the right is the three-dimensional position of the limb key points of the fruit fly obtained by processing and integration, and each point represents a limb key point.

[0033] Summarizing the top view captured images of the fruit fly at multiple moments and the three-dimensional positions of the limb key points of the fruit fly can be used for subsequent construction of a three-dimensional dataset. Although Figure 2 and Figure 3A single fruit fly is shown as an example, but it should be noted that the number of objects is not limited to a single one, and two or more are also acceptable.

[0034] In some embodiments, the two-dimensional positions of the limb key points of at least one object (such as, but not limited to, fruit flies) in the training two-dimensional images of the main view and multiple side views of the first group can be obtained by combining manual annotation with the first learning network. For example, a two-dimensional data set can be constructed using a subset of the training two-dimensional images of the main view and multiple side views of the first group and the annotation results (such as manual annotation results) of the limb key points of at least one object in each of the training two-dimensional images. The constructed two-dimensional data set is used to train the first learning network. Then, based on the training two-dimensional images of the main view and multiple side views of the first group, the trained first learning network is used to detect the two-dimensional positions of the limb key points of at least one object in the training two-dimensional images of the main view and multiple side views of the first group. That is, a part is selected from the set of training two-dimensional images as training data to train the first learning network, and then the trained first learning network is used to automatically detect the two-dimensional positions of the limb key points of at least one object in the remaining training two-dimensional images. In this way, the two-dimensional positions of the limb key points of the objects in a large number of training two-dimensional images can be obtained quickly and accurately.

[0035] In some embodiments, the first learning network can adopt a convolutional neural network, and can be implemented using various two-dimensional key point detection networks, such as, but not limited to, HRNet, OpenPose, etc. For example, the first learning network can adopt a deep neural network, including a feature extraction part and a two-dimensional output part. As shown in FIG. 4(a), for example, for the OpenPose detection network, the feature extraction part can include multiple convolutional layers, and the two-dimensional output part can include a PAF (Part Association Field) block and a confidence block, which will not be elaborated here. In some embodiments, the feature extraction part can be composed of convolutional layers, pooling layers, and activation layers, and the two-dimensional output part can be implemented using fully connected layers, upsampling layers, etc., which will not be elaborated here.

[0036] Next, in step 102, a three-dimensional data set is constructed using the region of interest containing the at least one object in the training two-dimensional image of the main view of the first group and the three-dimensional positions of the first group of limb key points of the at least one object for training the second learning network (see step 103).

[0037] In some embodiments, the training two-dimensional images of the main viewpoints of the first group and the three-dimensional positions of the limb key points of at least one object in the first group can also be directly used to construct a three-dimensional data set. However, in step 102, the region of interest containing the at least one object is extracted from the training two-dimensional images of each main viewpoint, such as the region of interest containing a single fruit fly and the three-dimensional positions of the limb key points of the fruit fly, to construct a three-dimensional data set. In this way, in each three-dimensional data, the information of the object in the region of interest is more concentrated, the interfering information is less, and the dimension of the extracted feature information is also lower, which is beneficial to improving the training efficiency of the second learning network.

[0038] In some embodiments, the second learning network may adopt a deep neural network, which can be any convolutional neural network for detecting three-dimensional coordinates from two-dimensional images, such as EpipolarPose. For example, the second learning network deep neural network may include a feature extraction part and a three-dimensional output part. As shown in FIG. 4(b), the three-dimensional output part may be implemented based on a transposed convolution layer and a softargmax activation function, and the feature extraction part may be composed of multiple convolutional layers, and a pooling layer and an activation layer may also be introduced.

[0039] The trained second learning network can then detect the three-dimensional positions of the limb key points of the target object based on the region of interest of the main viewpoint. Compared with detecting the three-dimensional positions of the limb key points of the target object based on the large image of the main viewpoint, the data processing load is significantly reduced, the calculation speed is faster, and rapid (including real-time) detection can be achieved, and it performs better in determining the three-dimensional pose of the target object that may change rapidly and irregularly.

[0040] In some embodiments, the target object whose dynamic three-dimensional position is to be detected may be different from the at least one object. Specifically, the target object to be detected can be different from the at least one object in the training two-dimensional images in terms of quantity and attributes. In some embodiments, the target object to be detected can be different from the at least one object in the training two-dimensional images in terms of attributes, but there is a certain degree of similarity. For example, training two-dimensional images labeled with the two-dimensional positions of the limb key points of multiple fruit flies can be used to construct a three-dimensional data set to train the second learning network, but the trained second learning network is used to detect the dynamic three-dimensional positions of another type of fruit fly from the to-be-detected two-dimensional images containing multiple fruit flies of another strain or with different body sizes (larger or smaller individuals). Preferably, the target object to be detected has the same attributes as the at least one object in the training two-dimensional images, that is, the two-dimensional images of fruit flies are used to construct a three-dimensional data set, and the trained second learning network with this three-dimensional data set is also used to detect the dynamic three-dimensional positions of the same type of fruit flies.

[0041] In some embodiments, the at least one object that annotates the two-dimensional positions of the limb key points is located in the middle of the training two-dimensional image, so that the annotation results of the two-dimensional positions of the limb key points can be as complete as possible, thereby improving the accuracy of processing and integrating the three-dimensional positions of the limb key points, and further improving the quality of the subsequent constructed three-dimensional dataset.

[0042] In step 104, based on the two-dimensional image to be detected with multiple objects in the main view, an interest region containing the target object is extracted. As Figure 8 shown, for example, the interest region containing the target object is extracted to only contain a single complete target object. Alternatively, the single target object can be centered and occupy more than a predetermined ratio of the area of the interest region, such as more than 10%, or more than 15%, or more than 20%, or more than 30%, or more than 50%, etc. The two-dimensional image to be detected with multiple objects in the main view can be obtained by shooting with a single top-view camera. There are various extraction methods. The image of the main-view camera can be segmented, and then an interest region with a fixed size can be cropped, and it is required that the interest region must contain the whole fruit fly, as Figure 8 shown.

[0043] In step 105, based on the extracted interest region containing the target object, the trained second learning network is used to detect the three-dimensional positions of the limb key points of the target object.

[0044] In step 106, based on the detected three-dimensional positions of the limb key points of the target object at different times, the dynamic three-dimensional posture of the target object is determined.

[0045] A three-dimensional pose detection method for dynamic interaction objects according to various embodiments of the present application, which is based on the two-dimensional positions of the limb key points of at least one object in the training two-dimensional images of the main view and multiple side views of the first group, and processes and integrates them to obtain the three-dimensional positions of the limb key points of the at least one object in the first group; uses the region of interest containing the at least one object and the three-dimensional positions of the limb key points of the at least one object in the training two-dimensional image of the main view of the first group to construct a three-dimensional data set. The three-dimensional data set constructed in this way reflects the mapping relationship between the region of interest where the object is located in the two-dimensional image of the main view and the three-dimensional positions of the limb key points of the object, and is used to train the second learning network, so that the second learning network can learn the mapping relationship. When actually detecting the dynamic three-dimensional positions of the limb key points of the target object that dynamically interacts with other objects, using the to-be-detected two-dimensional image of the main view containing multiple objects obtained by a simple shooting system of a single main-view camera, during the process of multiple objects freely and dynamically interacting with each other, despite the complex and changeable occlusion relationships between them, it is still possible to accurately and quickly (including in real time) identify and detect the three-dimensional positions of the limb key points of any target object at each shooting moment, thereby determining the dynamic three-dimensional pose of the target object during this process.

[0046] As Figure 5 shown, the processing and integration for obtaining the three-dimensional positions of the limb key points of the at least one object may further include the following steps. In step 501, based on the training two-dimensional image of the main view at the same time, evaluate the occlusion situation of the at least one object in the training two-dimensional images of each side view at the corresponding time, and remove the two-dimensional positions of the limb key points corresponding to the side views where the occlusion situation of the at least one object reaches a severe condition. As Figure 6 shown, in the top-view shooting image of two fruit flies, it is possible to evaluate the occlusion situation of the target object fruit fly below by other fruit flies above in each lateral view, and remove the shooting images of some lateral views with severe occlusion situations (such as Figure 7As shown (referred to as the occluded captured image), and the two-dimensional positions of the limb key points marked therein. Given that the target object is also detected based on the top-view captured image subsequently, then in the top-view captured image to be detected, the two-dimensional positions of the limb key points of the target object, the fruit fly, usually also have a similar occluded situation in the corresponding side-view. That is to say, the two-dimensional positions of the limb key points that will be occluded in the side-view from the main view in the training two-dimensional images will also be occluded in the corresponding side-view during actual detection. Then, remove the two-dimensional positions of the limb key points corresponding to the perspective where the occlusion situation of the at least one object reaches the severe condition. In step 502, based on the two-dimensional positions of the limb key points of the remaining at least one object, integrate to obtain the three-dimensional positions of the first set of limb key points of the at least one object. The mapping relationship between the three-dimensional positions of the first set of limb key points of the at least one object obtained and the region of interest containing the at least one object in the training two-dimensional image of the main view is more consistent with the mapping relationship between the main-view captured image during actual detection and the actual three-dimensional position of the target object therein in terms of the spatial occlusion relationship. For training the second learning network, it can enable the second learning network to better learn the spatial occlusion relationship, making the detected three-dimensional position of the target object more accurate and reasonable. When constructing the three-dimensional dataset, especially when integrating to obtain the three-dimensional positions of the limb key points for this three-dimensional dataset, use the top-view camera to evaluate the occlusion situation, remove the camera detection results corresponding to the perspectives with severe occlusion, and prevent the detection errors caused by occlusion from affecting the results of the triangulation process (which will be described in detail later), making the three-dimensional key point detection more accurate.

[0047] In some embodiments, the three-dimensional pose detection method may further include: screening according to the confidence levels of the two-dimensional positions of the limb key points of the at least one object in the detected training two-dimensional images, and only the two-dimensional positions of the limb key points with confidence levels higher than the threshold are used to integrate to obtain the three-dimensional positions of the limb key points of the at least one object. In this way, the two-dimensional positions of the limb key points based on which the integration is performed all have a relatively high confidence level, and the three-dimensional positions of the limb key points of the at least one object obtained by integration are more accurate and reasonable.

[0048] Based on the two-dimensional positions of the limb key points of the remaining at least one object, integrate to obtain the three-dimensional positions of the first set of limb key points of the at least one object. This integration process can usually also be referred to as the triangulation process, which refers to estimating the three-dimensional coordinate position for a three-dimensional coordinate point through the two-dimensional positions projected from multiple perspectives. The triangulation process is a mature machine vision algorithm with various implementation methods, which will not be elaborated here. Use the two-dimensional key points provided by the multi-view camera array for triangulation to obtain accurate three-dimensional key points, and thus obtain a high-quality interactive object behavior dataset, whose diversity exceeds that of a single-object behavior dataset.

[0049] In some embodiments, as Figure 9 shown, the construction of the three-dimensional data set specifically includes the following steps. In step 901, a centering transformation is performed on the region of interest containing the at least one object, such that the at least one object faces a predetermined direction and is centered in the region of interest containing the at least one object. Refer to Figure 8 , after the centering transformation of the region of interest with the fruit fly as the object, the head is upward and centered in the region of interest. In step 902, a corresponding centering transformation is also performed on the three-dimensional positions of the limb key points of the at least one object. For example, if the centering transformation process in step 901 is performed via a centering transformation matrix, then the same centering transformation matrix can be applied to perform a corresponding centering transformation on the three-dimensional positions of the limb key points of the at least one object. In step 903, the three-dimensional positions of the limb key points of the at least one object after transformation and the region of interest of the at least one object after transformation are used to construct the three-dimensional data set. In each piece of three-dimensional data constructed in this way, the orientation of the fruit fly in the region of interest is upward with the head and centered. By unifying the orientation and position of the fruit fly in the region of interest, the training computational complexity of the second learning network is significantly reduced and the training is more efficient.

[0050] In some embodiments, as Figure 10 shown, correspondingly, the centering transformation can also be applied to the actual detection process of the three-dimensional positions of the limb key points of the target object. Specifically, in step 1001, a centering transformation is performed on the region of interest containing the target object, such that the target object faces the predetermined direction and is centered in the region of interest containing the target object.

[0051] In step 1002, based on the region of interest containing the target object after transformation, the trained second learning network is used to detect the corresponding three-dimensional positions of the limb key points of the target object after transformation. Figure 11 shows a comparison diagram of the region of interest containing the target object after the centering transformation according to an embodiment of the present application and the three-dimensional positions of the limb key points of the target object after superimposing the centering transformation process. As Figure 11As shown, after the superimposed centering transformation process, the target object is centered in the region of interest. After the corresponding transformed three-dimensional positions of the limb key points of the target object are projected two-dimensionally on the image plane, they are also centered in the region of interest. In step 1003, the inverse transformation of the centering transformation is performed on the transformed three-dimensional positions of the limb key points of the target object to obtain the three-dimensional positions of the limb key points of the target object as the detection result. In some embodiments, through the above processing, the three-dimensional positions of the limb key points of the target object can be automatically marked on the target objects with various orientations and positions in the shooting screen of the main view. By making the target object face the predetermined direction and be centered in the region of interest containing the target object, the detection efficiency of the three-dimensional position by the second learning network can be improved and the detection complexity can be reduced. Further, the detection result of the three-dimensional positions of the limb key points of the target object obtained by the inverse transformation can also be more accurate. By extracting the region of interest in the large shooting screen of the main view and cooperating with the centering transformation-inverse transformation to obtain the detection result, the detection process is faster and more accurate, almost approaching real-time, and the superiority is more significant in determining the rapidly changing or even unpredictable three-dimensional postures of multiple target objects.

[0052] Figure 12 FIG. shows a schematic block diagram of a three-dimensional pose detection system for a dynamic interaction object according to an embodiment of the present application. As Figure 12 shown, the three-dimensional pose detection system includes a single top-view camera 1201 and at least one processor 1202. The single top-view camera 1201 can be configured to capture two-dimensional images of different times of the main view containing multiple objects as the two-dimensional images to be detected. The at least one processor 1202 can be configured to execute the three-dimensional pose detection method for the dynamic interaction object according to various embodiments of the present application. Various examples and details of the above method can be selectively combined here and will not be elaborated. That is, with the simple configuration of the single top-view camera 1201, the two-dimensional images to be detected of the main view containing multiple objects can be used to accurately and quickly (including in real time) identify and detect the three-dimensional positions of the limb key points of any target object at each shooting moment during the process of multiple objects freely and dynamically interacting with each other, so as to determine the dynamic three-dimensional pose of the target object during this process.

[0053] In some embodiments, the at least one processor 1202 may be a processing device, including one or more general processing devices such as, for example, a microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), etc. More specifically, the at least one processor 1202 may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor implementing other instruction sets, or a processor implementing a combination of instruction sets. The at least one processor 1202 may also be one or more dedicated processing devices, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a system on a chip (SOC), etc. As those skilled in the art will appreciate, in some embodiments, the at least one processor 1202 may be a dedicated processor rather than a general purpose processor. The at least one processor 1202 may include one or more known processing devices.

[0054] In some embodiments, there is also provided a three-dimensional pose detection device for a dynamic interaction object, the device including at least one processor configured to execute the three-dimensional pose detection method for a dynamic interaction object according to various embodiments of the present application. Various examples and details of the above method can be selectively combined herein and will not be elaborated herein.

[0055] In some embodiments, there is also provided a computer-readable storage medium having stored thereon an executable program, which when executed by a processor, implements the three-dimensional pose detection method for a dynamic interaction object according to various embodiments of the present application. Various examples and details of the above method can be selectively combined herein and will not be elaborated herein.

[0056] In this document, the term "or" is used to denote a non-exclusive "or", unless otherwise specified, "A or B" includes "A without B", "B without A", and "A and B". In this document, the terms "including" and "in which" are used as plain language equivalents of the respective terms "comprising" and "wherein". Further, in the following claims, the term "including" is open-ended, i.e., a device, system, apparatus, article, composition, concept or process that includes elements other than those listed after such a term in the claims is still considered to fall within the scope of that claim. Further, in the following claims, the terms "first", "second", "third", etc. are used only as labels and are not intended to impose a numerical requirement on their objects.

[0057] The method examples described herein can be at least partially machine or computer-implemented. Some examples can include a computer-readable medium or machine-readable medium encoded with instructions that can be operated to configure an electronic device to perform the methods as described in the above examples. Implementations of these methods can include software code such as microcode, assembly language code, higher-level language code, etc. Various software programming techniques can be used to create various programs or program modules. For example, program segments or program modules can be designed in or through Java, Python, C, C++, assembly language, or any known programming language. One or more such software segments or modules can be integrated into a computer system and / or a computer-readable medium. Such software code can include computer-readable instructions for performing various methods. Such software code can form part of a computer program product or a computer program module. Additionally, in one example, the code can be tangibly stored on one or more volatile, non-transitory, or non-volatile tangible computer-readable media, for example, during runtime or at other times. Examples of such tangible computer-readable media can include, but are not limited to, hard disks, removable disks, removable optical disks (e.g., compact disks and digital video disks), magnetic tapes, memory cards or sticks, random access memory (RAM), read-only memory (ROM), etc.

[0058] Moreover, although exemplary embodiments have been described herein, the scope includes any and all embodiments based on the present disclosure having equivalent elements, modifications, omissions, combinations (e.g., of the aspects of various embodiments), adaptations, or alterations. The elements in the claims will be interpreted broadly based on the language employed in the claims and are not limited to the examples described in this specification or during the prosecution of this application, and the examples will be construed as non-exclusive. Additionally, the steps of the disclosed methods can be modified in any way, including by reordering the steps or inserting or deleting steps. Accordingly, this specification and the examples are intended to be considered only as examples, and the true scope and spirit are indicated by the full scope of the following claims and their equivalents.

[0059] The foregoing description is intended to be illustrative and not restrictive. For example, the above examples (or one or more aspects thereof) may be used in combination with each other. For instance, those of ordinary skill in the art may use other embodiments when reading the above description. Additionally, in the above detailed description, various features may be grouped together to simplify the present disclosure. This should not be construed as an intention that features of the disclosed subject matter not claimed are essential to any one of the claims. Rather, the subject matter of the present invention may be less than all of the features of the specific embodiments disclosed. Accordingly, the following claims are incorporated into the detailed description by way of example or illustration, where each claim stands on its own as a separate embodiment, and it is contemplated that these embodiments may be combined with each other in various combinations or permutations. The scope of the present invention should be determined with reference to the appended claims and the full scope of equivalents to which such claims are entitled.

Claims

1. A method for detecting a three-dimensional posture of a dynamic interactive object, which uses a single main view image to detect the three-dimensional posture, wherein: A plurality of objects dynamically interact with each other, and a dynamically changing occlusion relationship is generated between the plurality of objects during the interaction, and the characteristics include: Based on the two-dimensional positions of the limb key points of at least one object in the first group of training two-dimensional images from a main perspective and a plurality of secondary perspectives, processing and integrating to obtain the three-dimensional positions of the first group of limb key points of the at least one object, wherein the main perspective is a top perspective and the secondary perspective is a side perspective; Constructing a three-dimensional data set using the first group of two-dimensional training images of the main perspective containing an area of ​​interest of the at least one object and the three-dimensional positions of the first group of limb key points of the at least one object; the area of ​​interest in the two-dimensional training images of the main perspective contains an occlusion area of ​​the at least one object and other dynamically interacting objects; Using the constructed three-dimensional dataset to train a second learning network; Extracting a region of interest containing a target object based on a two-dimensional image to be detected from a primary perspective containing multiple objects; Based on the extracted region of interest containing the target object, using the trained second learning network to detect the three-dimensional position of the limb key points of the target object; and Determining the dynamic three-dimensional posture of the target object based on the detected three-dimensional positions of the target object's limb key points at different times; The processing and integration for obtaining the three-dimensional positions of the first group of limb key points of the at least one object specifically includes: Based on the simultaneous training two-dimensional images of the main perspective, evaluating the occlusion condition of the at least one object in the training two-dimensional images of the slave perspectives at the corresponding time, and removing the two-dimensional positions of the limb key points corresponding to the slave perspectives where the occlusion condition of the at least one object reaches a serious condition; and Based on the remaining two-dimensional positions of the limb key points of the at least one object, integrating to obtain the three-dimensional positions of the limb key points of the first group of the at least one object; The processing and integration for obtaining the three-dimensional positions of the first group of limb key points of the at least one object further includes: Screening is performed according to the confidence of the two-dimensional position of the limb key point of the at least one object in the detected training two-dimensional image, and the two-dimensional position of the limb key point with a confidence higher than a threshold is used for integration to obtain the three-dimensional position of the limb key point of the at least one object; The construction of the three-dimensional data set specifically includes: Performing a centering transformation on a region of interest including the at least one object and other objects so that the at least one object faces a predetermined direction and is centered on the region of interest including the at least one object; The three-dimensional position of the limb key point of the at least one object is also subjected to a corresponding centralization transformation; constructing the three-dimensional data set using the transformed three-dimensional positions of the limb key points of the at least one object and the transformed region of interest of the at least one object; Wherein, the occlusion information of the at least one object and the other objects is retained in the region of interest after the centralization transformation.

2. The three-dimensional posture detection method according to claim 1, characterized in that: The two-dimensional position of the limb key point of at least one object in the first group of training two-dimensional images of the main view and multiple slave views is obtained by the following processing: Using a subset of the first group of training two-dimensional images from a main perspective and multiple sub-perspectives and the annotation results of limb key points of at least one object in each of the training two-dimensional images, to construct a two-dimensional data set; wherein the multiple sub-perspectives of the training two-dimensional images contain occlusion areas of the at least one object and other dynamically interacting objects; Using the constructed two-dimensional data set to train a first learning network; Based on the remaining two-dimensional training images of the first group from the main perspective and multiple slave perspectives, the trained first learning network is used to detect the two-dimensional positions of the limb key points of at least one object in the remaining two-dimensional training images of the first group from the main perspective and multiple slave perspectives.

3. The three-dimensional posture detection method according to claim 1, characterized in that: Based on the extracted region of interest containing the target object, using the trained second learning network to detect the three-dimensional position of the limb key points of the target object specifically includes: Performing a centering transformation on the region of interest containing the target object so that the target object faces the predetermined direction and is centered in the region of interest containing the target object; Based on the transformed region of interest containing the target object, using the trained second learning network to detect the corresponding transformed three-dimensional positions of the limb key points of the target object; The inverse transformation of the centering transformation is performed on the transformed three-dimensional positions of the limb key points of the target object to obtain the three-dimensional positions of the limb key points of the target object as the detection results.

4. The three-dimensional posture detection method according to claim 1, characterized in that: The at least one object is located in the middle of the training two-dimensional image, and the target object is different from the at least one object.

5. The three-dimensional posture detection method according to claim 1, characterized in that: The target object-containing region of interest is extracted to contain only a single target object, or to make the single target object be centered and occupy more than a predetermined ratio of the area of ​​the region of interest.

6. The three-dimensional posture detection method according to claim 1, characterized in that: The main-view two-dimensional image to be detected containing multiple objects is obtained by photographing using a single top-view camera.

7. The three-dimensional posture detection method according to claim 2, characterized in that: The first learning network and the second learning network both include convolutional neural networks.

8. The three-dimensional posture detection method according to claim 1, characterized in that: The object includes any one of an animal, a movable facility, and a robot.

9. A three-dimensional posture detection system for a dynamic interactive object, characterized in that: include: A single top-view camera configured to capture two-dimensional images of a main view of a plurality of objects at different times as two-dimensional images to be detected; as well as At least one processor configured to execute the three-dimensional posture detection method of a dynamic interactive object according to any one of claims 1-8.

10. A three-dimensional posture detection device for a dynamic interactive object, characterized in that: The method comprises at least one processor configured to execute the three-dimensional posture detection method of a dynamic interactive object according to any one of claims 1-8.

11. A computer-readable storage medium having an executable program stored thereon, wherein when the executable program is executed by a processor, the three-dimensional posture detection method of a dynamic interactive object according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Insect key point automatic labeling method based on top-down deep learning architecture and application

    CN113516734A

  • Robot grabbing method and device, electronic equipment and storage medium

    CN114387513A