Method, system, and apparatus for detecting three-dimensional posture of dynamic interaction object, and storage medium
By constructing a three-dimensional data set and training learning network, using a single main perspective camera to identify dynamic three-dimensional poses of multiple objects, the difficulty of three-dimensional pose detection caused by the complex occlusion relationship of multiple free interaction objects in the prior art is solved, and fast and accurate three-dimensional pose recognition is achieved.
Patent Information
- Application Number
- PCT/CN2024/103400
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-08
- Filing Date
- 2024-07-03
- Publication Date
- 2025-08-14
AI Technical Summary
The prior art is difficult to accurately identify the dynamic three-dimensional poses of each object during the dynamic process of multiple free interaction objects, especially when the occlusion relationship is complex and changeable, it is impossible to achieve fast and accurate three-dimensional pose detection.
A single main viewing camera is used to combine multiple viewing cameras to build a three-dimensional data set and train a second learning network by training the location of limb key points in the two-dimensional image. This network is used to extract the area of interest from the main viewing image, identify the three-dimensional position of the limb key points of the target object, and determine its dynamic three-dimensional posture.
During the free interaction of multiple objects, the three-dimensional position and dynamic three-dimensional pose of the key points of the target object's limbs can be accurately and quickly identified, real-time three-dimensional pose detection is achieved.
Smart Images

Figure CN2024103400_14082025_PF_FP_ABST
Abstract
Description
Three-dimensional posture detection method, system, device and storage medium for dynamic interactive objects Technical Field
[0001] The present application relates to the technical field of object spatial motion detection in machine vision, and more specifically to a method, system, device and medium for detecting three-dimensional posture of a dynamic interactive object. Background Art
[0002] People have tried to introduce machine vision technology to analyze the dynamic interaction process of groups of objects in a free state, such as but not limited to animal social activities, robot interaction processes, etc., in which each object can behave spontaneously and freely.
[0003] To analyze the behavioral characteristics of these dynamically interacting objects, the first step is to detect the 3D pose information of each interacting entity. However, 3D pose detection based on depth cameras can only obtain partial and simple pose information. While some existing tools can provide dynamic body movement analysis of multiple entities, they require constraints on individual movements, making them inapplicable to scenarios where entities interact freely. Currently, methods for detecting freely interacting entities are limited to single entities, and state analysis of multiple dynamically interacting entities faces numerous technical obstacles. For example, during the interaction of multiple objects, complex occlusion relationships arise between them. Due to the objects' freedom of movement, these occlusion relationships frequently and dynamically change. Furthermore, the interactive behavior of multiple objects is more diverse than that of a single entity, making it difficult to accurately identify the dynamic 3D positions of key points of each object. Research on the interaction of multiple objects requires not only accurate identification of the dynamic 3D pose of each object, but also stringent requirements for timely recognition.
[0004] Summary of the Invention
[0005] This application is proposed to solve the above technical problems in the prior art.
[0006] The present application aims to provide a method, system, device and computer program product for detecting the three-dimensional posture of dynamic interactive objects. The method utilizes a simple shooting system of a camera with a single main perspective. In the process of multiple objects freely and dynamically interacting with each other, despite the complex and changeable occlusion relationships between each other, it can still accurately and quickly (including in real time) identify and detect the three-dimensional position of the key points of the limbs of any target object at each shooting moment, thereby determining the dynamic three-dimensional posture of the target object during this process.
[0007] According to the first aspect of the present application, a method for detecting the three-dimensional posture of a dynamic interactive object is provided, wherein a plurality of objects dynamically interact with each other. The method comprises the following steps. Based on the two-dimensional positions of the limb key points of at least one object in the first group of main-view and multiple secondary-view training two-dimensional images, the three-dimensional positions of the limb key points of the first group of the at least one object are processed and integrated to obtain the three-dimensional positions of the limb key points of the first group of the at least one object. A three-dimensional data set is constructed using the region of interest containing the at least one object and the three-dimensional positions of the limb key points of the first group of the at least one object in the first group of main-view training two-dimensional images. A second learning network is trained using the constructed three-dimensional data set. Based on the two-dimensional image to be detected from the main-view containing multiple objects, a region of interest containing the target object is extracted. Based on the extracted region of interest containing the target object, the trained second learning network is used to detect the three-dimensional positions of the limb key points of the target object. The dynamic three-dimensional posture of the target object is determined based on the detected three-dimensional positions of the limb key points of the target object at different times.
[0008] According to a second aspect of the present application, a system for detecting the three-dimensional posture of a dynamic interactive object is provided. The system includes a single top-view camera and at least one processor. The single top-view camera is configured to capture two-dimensional images from a primary perspective of multiple objects at different times as the two-dimensional images to be detected. The at least one processor is configured to execute a method for detecting the three-dimensional posture of a dynamic interactive object according to various embodiments of the present application. The method includes the following steps: Based on the two-dimensional positions of the limb key points of at least one object in a first set of primary-view and multiple secondary-view training two-dimensional images, processing and integrating the two-dimensional positions of the limb key points of the at least one object is performed to obtain the three-dimensional positions of the first set of limb key points of the at least one object. A three-dimensional dataset is constructed using the regions of interest containing the at least one object and the three-dimensional positions of the first set of limb key points of the at least one object in the first set of primary-view training two-dimensional images. A second learning network is trained using the constructed three-dimensional dataset. Based on the two-dimensional images to be detected from the primary-view of multiple objects, a region of interest containing a target object is extracted. Based on the extracted region of interest containing the target object, the trained second learning network is used to detect the three-dimensional positions of the limb key points of the target object. The dynamic three-dimensional posture of the target object is determined based on the detected three-dimensional positions of the limb key points of the target object at different times.
[0009] According to a third aspect of the present application, a device for detecting the three-dimensional posture of a dynamic interactive object is provided. The device includes at least one processor configured to execute a method for detecting the three-dimensional posture of a dynamic interactive object according to various embodiments of the present application. The method includes the following steps: Based on the two-dimensional positions of the limb key points of at least one object in a first set of main-view and multiple secondary-view training two-dimensional images, processing and integrating are performed to obtain the three-dimensional positions of the first set of limb key points of the at least one object. A three-dimensional dataset is constructed using the regions of interest of the at least one object and the three-dimensional positions of the first set of limb key points of the at least one object in the first set of main-view training two-dimensional images. A second learning network is trained using the constructed three-dimensional dataset. Based on the two-dimensional images to be detected from the main-view containing multiple objects, a region of interest containing a target object is extracted. Based on the extracted region of interest containing the target object, the trained second learning network is used to detect the three-dimensional positions of the limb key points of the target object. The dynamic three-dimensional posture of the target object is determined based on the detected three-dimensional positions of the limb key points of the target object at different times.
[0010] According to a fourth aspect of the present application, a computer-readable storage medium is provided, on which an executable program is stored. When the executable program is executed by a processor, the three-dimensional posture detection method of the dynamic interactive object according to each embodiment of the present application is implemented.
[0011] Utilizing the methods, systems, devices, and media for detecting the three-dimensional posture of dynamically interactive objects according to various embodiments of the present application, the methods, systems, devices, and media are configured to process and integrate the two-dimensional positions of the limb key points of at least one object in a first set of primary-view and multiple secondary-view training two-dimensional images to obtain the three-dimensional positions of the first set of limb key points of the at least one object; and to construct a three-dimensional dataset using the regions of interest of the at least one object and the three-dimensional positions of the first set of limb key points of the at least one object in the first set of primary-view training two-dimensional images. The thus constructed three-dimensional dataset reflects the mapping relationship between the regions of interest of the objects in the primary-view two-dimensional images and the three-dimensional positions of the limb key points of the objects, and is used to train a second learning network, thereby enabling the second learning network to learn the mapping relationship. When actually detecting the dynamic three-dimensional positions of the limb key points of a target object dynamically interacting with other objects, using a simple capture system (camera) with a single primary-view to capture multiple primary-view two-dimensional images of the objects to be detected, during the process of multiple objects freely and dynamically interacting with each other, despite complex and changing occlusion relationships between them, the three-dimensional positions of the limb key points of any target object at each capture moment can be accurately and rapidly (including in real time) detected, thereby determining the dynamic three-dimensional posture of the target object during this process. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Features, advantages, and technical and industrial significance of exemplary embodiments of the present invention will be described below with reference to the accompanying drawings, wherein like numerals refer to like elements, and wherein:
[0013] FIG1 shows a flow chart of a method for detecting a three-dimensional posture of a dynamic interactive object according to an embodiment of the present application.
[0014] FIG2 shows an example of a two-dimensional image for training with fruit flies as the object, wherein a two-dimensional position distribution map of key points of the limbs of a single fruit fly is shown.
[0015] Figure 3 shows a comparison diagram of two-dimensional images of the main perspective and multiple sub-perspectives according to an embodiment of the present application and the two-dimensional positions of the limb key points of the objects therein, as well as the three-dimensional positions of the limb key points of the corresponding objects obtained by processing and integrating them.
[0016] Figure 4(a) shows a block diagram of a first learning network according to an embodiment of the present application, which is used to detect the two-dimensional positions of limb key points of at least one object based on a first set of training two-dimensional images from a main perspective and multiple secondary perspectives.
[0017] FIG4( b ) shows a block diagram of a second learning network according to an embodiment of the present application, which is used to detect the three-dimensional positions of the limb key points of the target object based on the extracted region of interest containing the target object.
[0018] Figure 5 shows a flowchart of processing and integrating the two-dimensional positions of the limb key points of at least one object in the training two-dimensional images of the first group of main perspectives and multiple slave perspectives to obtain the three-dimensional positions of the limb key points of the first group of at least one object according to an embodiment of the present application.
[0019] FIG6 shows an example of a top-view two-dimensional image for training of multiple fruit flies according to an embodiment of the present application, in which a fruit fly as a target object is blocked by other fruit flies in front.
[0020] FIG. 7 shows an example of a two-dimensional training image of a side view of a fruit fly to be removed according to the occlusion of the fruit fly as the target object shown in FIG. 6 .
[0021] FIG8 shows a diagram of a region of interest containing a target object in a main viewing angle after a centralization transformation according to an embodiment of the present application.
[0022] FIG9 shows a flowchart of constructing a three-dimensional data set for training the second learning network according to an embodiment of the present application.
[0023] FIG10 shows a flowchart of detecting the three-dimensional positions of the limb key points of the target object based on the extracted region of interest containing the target object and using the trained second learning network according to an embodiment of the present application.
[0024] FIG11 shows a comparison diagram of a region of interest including a target object after centralization transformation and a two-dimensional projection of the three-dimensional position of key points of the limbs of the target object after superimposition of centralization transformation processing according to an embodiment of the present application.
[0025] FIG12 shows a schematic block diagram of a three-dimensional posture detection system for a dynamic interactive object according to an embodiment of the present application. DETAILED DESCRIPTION
[0026] Specific embodiments are described in detail below with reference to the accompanying drawings. In the various figures, identical or corresponding elements are denoted by the same reference numerals, and repeated descriptions are omitted where necessary to clarify the description. Please note that the arrows in the flowcharts of this application are not intended to limit the order in which the steps are executed. As long as there is no logical conflict, the steps can be executed in parallel, in a different order, or split into multiple steps, or combined, etc., without limitation in this application. Unless otherwise defined, technical or scientific terms used in this application should have the same meaning as those commonly understood by persons of ordinary skill in the art to which this application belongs. The terms "first," "second," and similar terms used in this application do not denote any order, quantity, or importance; they are simply used to distinguish different components. Terms such as "include" or "comprising" mean that the element or object preceding the term includes the elements or objects listed after the term, and their equivalents, without excluding other elements or objects. Terms such as "connected" or "connected" are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect. “Up,” “down,” “left,” “right,” etc. are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0027] Figure 1 shows a flow chart of a method for detecting the three-dimensional posture of dynamically interactive objects according to an embodiment of the present application. The moving objects herein can be any of animals (such as insects), movable facilities, and robots. Animals herein may include, but are not limited to, insects, small fish, small amphibians, small reptiles, etc., and robots may also include logistics robots, etc. In some embodiments, the objects include any of small animals, small movable facilities, and small robots. "Small" means relatively small relative to the field of view. For example, small animals may include, but are not limited to, insects, small fish, small amphibians, small reptiles, etc. For a panoramic field of view, small robots may also include small logistics robots that travel on the same factory floor. Multiple objects can interact dynamically with each other, allowing the positions of each object to change dynamically, and the spatial overlap and occlusion relationships between objects can also change dynamically. Depending on the type of dynamically interactive object, the three-dimensional posture detection method can be used for different analysis purposes in different application scenarios. For example, for a colony of fruit flies, the social behavior of the colony can be analyzed by allowing them to interact freely within a certain space and detecting the dynamic changes in the three-dimensional posture of each fruit fly during the process. For example, for small logistics robots, each logistics robot transports back and forth on the factory floor by running the autonomous route dynamic planning function. By detecting the dynamic changes in the three-dimensional posture of each logistics robot during transportation, it can be verified whether they have achieved good avoidance of each other and optimized transportation routes.
[0028] The following uses a fruit fly as an example to illustrate a method for detecting the three-dimensional posture of a dynamic interactive object. However, it should be noted that the method can be flexibly applied to other scenes and other types of small objects. The three-dimensional posture detection method includes the following steps.
[0029] In step 101, the two-dimensional positions of at least one subject's limb key points in a first set of two-dimensional training images from a primary perspective and multiple secondary perspectives are processed and integrated to obtain the three-dimensional positions of the first set of limb key points of the at least one subject. The two-dimensional training images from the primary perspective and the three-dimensional positions obtained from the processing and integration are subsequently used to construct a three-dimensional dataset for training a second learning network. This second learning network is used to detect the three-dimensional positions of the target subject's limb key points based on an extracted region of interest (e.g., but not limited to an image block) containing the target subject. This will be described in detail below and is not repeated here.
[0030] In the following text, the main perspective is taken as the top perspective, and the secondary perspective is taken as the side perspective for illustration, but it should be noted that the present application is not limited to this. For situations where multiple objects are widely distributed on a flat surface, such as the fruit flies and logistics robots mentioned above, the two-dimensional image obtained by taking the top perspective as the main perspective can reflect most of the spatial position information of the object, and the spatial position information obtained by the two-dimensional images taken from other perspectives is richer, and the distribution of each object in the two-dimensional image taken from the top perspective as the main perspective is sparser than that of the two-dimensional images taken from other perspectives, which helps to efficiently extract the area of interest where the target object is located while introducing as little information as possible about other objects. In some embodiments, the spatial characteristics of the object activities are different, and different main perspectives can be used accordingly, as long as the two-dimensional image taken from the perspective can reflect most of the spatial position information of the object and the distribution of each object is sparser.
[0031] To obtain sufficient 2D images for training, a top-view camera and multiple side-view cameras can be used to simultaneously capture multiple interacting fruit flies, each fly interacting with the others on their own surface. The timing of each camera must be synchronized to simultaneously capture images from various viewpoints at each time point. This allows for multiple viewpoints to be captured during each capture. A certain number of these images are then manually annotated with the coordinates of key points on each fly's limbs, resulting in a set of 2D images with these key points annotated.
[0032] In some embodiments, the limb key points in the captured image can be manually annotated to identify their position coordinates. When annotating, only the visible key points of the fruit fly can be annotated, as shown in Figure 2. The connecting lines between the limb key points are only used to show the relationship between these limb key points.
[0033] The top and side views of a fruit fly captured simultaneously at a specific moment, along with the coordinates of the key points of each limb annotated therein, can be processed and integrated to obtain the three-dimensional positions of the key points of the fruit fly's limbs at that moment. As shown in Figure 3, the center image block (corresponding to the region of interest) in the first column of the two columns of two-dimensional images is from the top view of the fruit fly at that moment, while the other image blocks (corresponding to the region of interest) are from the six side views of the fruit fly captured at the same moment. The lines connecting the dots on the right are the three-dimensional positions of the key points of the fruit fly's limbs at that moment, obtained through processing and integration, with each dot representing a key point of the limb.
[0034] The top-view images of the fruit flies at multiple times and the 3D positions of the key points of the fruit flies' limbs can be aggregated to subsequently construct a 3D dataset. Although Figures 2 and 3 illustrate a single fruit fly as an example, the number of subjects is not limited to one; two or more can also be used.
[0035] In some embodiments, manual annotation can be used in conjunction with a first learning network to obtain the two-dimensional positions of key points on the limbs of at least one object (e.g., but not limited to, a fruit fly) in a first set of training two-dimensional images from a primary perspective and multiple secondary perspectives. For example, a subset of the first set of training two-dimensional images from a primary perspective and multiple secondary perspectives, along with the annotation results (e.g., manual annotation results) of key points on the limbs of at least one object in each of these training two-dimensional images, can be used to construct a two-dimensional dataset. The constructed two-dimensional dataset is then used to train the first learning network. Based on the first set of training two-dimensional images from a primary perspective and multiple secondary perspectives, the trained first learning network is then used to detect the two-dimensional positions of key points on the limbs of at least one object in the first set of training two-dimensional images from a primary perspective and multiple secondary perspectives. In other words, a portion of the set of training two-dimensional images is selected as training data to train the first learning network, and the trained first learning network is then used to automatically detect the two-dimensional positions of key points on the limbs of at least one object in the remaining training two-dimensional images. In this way, the two-dimensional positions of key points on the limbs of objects in a large number of training two-dimensional images can be quickly and accurately obtained.
[0036] In some embodiments, the first learning network may adopt a convolutional neural network, for example, it may be implemented using various two-dimensional key point detection networks, such as but not limited to HRNet, OpenPose, etc. For example, the first learning network may adopt a deep neural network, including a feature extraction unit and a two-dimensional output unit. As shown in Figure 4(a), for example, for the OpenPose detection network, the feature extraction unit may include multiple convolutional layers, and the two-dimensional output unit may include a PAF (part association vector field) block and a confidence block, which are not described in detail here. In some embodiments, the feature extraction unit may be composed of a convolutional layer, a pooling layer, and an activation layer, and the two-dimensional output unit may be implemented using a fully connected layer, an upsampling layer, etc., which are not described in detail here.
[0037] Next, in step 102, a three-dimensional dataset is constructed using the first group of training two-dimensional images of the main perspective containing the region of interest of the at least one object and the three-dimensional positions of the first group of limb key points of the at least one object for training the second learning network (see step 103).
[0038] In some embodiments, the first set of primary-view training 2D images and the 3D positions of the first set of limb key points of at least one object can also be directly used to construct a 3D dataset. However, in step 102, a region of interest containing the at least one object, such as a region of interest containing a single fruit fly and the 3D positions of the limb key points of the fruit fly, is extracted from each primary-view training 2D image to construct the 3D dataset. This allows for more concentrated information about the object in the region of interest within each piece of 3D data, with less interference, and correspondingly, lower dimensionality in the extracted feature information, which improves the training efficiency of the second learning network.
[0039] In some embodiments, the second learning network can adopt a deep neural network, which can be any convolutional neural network that detects three-dimensional coordinates based on a two-dimensional image, such as EpipolarPose. For example, the second learning network deep neural network can include a feature extraction unit and a three-dimensional output unit. As shown in Figure 4(b), the three-dimensional output unit can be implemented based on a deconvolution layer and a softargmax activation function, and the feature extraction unit can be composed of multiple convolutional layers, and a pooling layer and an activation layer can also be introduced.
[0040] The trained second learning network can detect the three-dimensional position of the target object's limb key points based on the area of interest of the main perspective. Compared with detecting the three-dimensional position of the target object's limb key points based on a large image from the main perspective, the data processing load is significantly reduced, the calculation speed is faster, and rapid (including real-time) detection can be achieved. It performs better in determining the three-dimensional posture of the target object that may change rapidly and irregularly.
[0041] In some embodiments, the target object whose dynamic three-dimensional position is to be detected may be different from the at least one object. Specifically, the target object to be detected may be different in both quantity and attributes from at least one object in the training two-dimensional image. In some embodiments, the target object to be detected may be different in attributes from at least one object in the training two-dimensional image, but may share a certain degree of similarity with the at least one object in the training two-dimensional image. For example, a training two-dimensional image containing annotations of the two-dimensional positions of key points of the limbs of multiple fruit flies may be used to construct a three-dimensional dataset to train a second learning network. However, the trained second learning network is then used to detect the dynamic three-dimensional position of another fruit fly from a two-dimensional image containing multiple fruit flies of a different strain or with different body shapes (larger or smaller individuals). Preferably, the target object to be detected has the same attributes as at least one object in the training two-dimensional image, that is, a two-dimensional image of a fruit fly is used to construct a three-dimensional dataset, and the second learning network trained with this three-dimensional dataset is also used to detect the dynamic three-dimensional position of the same type of fruit fly.
[0042] In some embodiments, the at least one object that marks the two-dimensional positions of the limb key points is located in the middle of the training two-dimensional image, so that the marking results of the two-dimensional positions of the limb key points can be as complete as possible, thereby improving the accuracy of the processed and integrated three-dimensional positions of the limb key points, and thereby improving the quality of the subsequently constructed three-dimensional data set.
[0043] In step 104, based on the two-dimensional image to be detected from the main perspective containing multiple objects, a region of interest containing the target object is extracted, as shown in FIG8 . For example, the region of interest containing the target object is extracted to contain only a single complete target object. Alternatively, the single target object can be centered and occupy more than a predetermined ratio of the area of the region of interest, such as more than 10%, or more than 15%, or more than 20%, or more than 30%, or more than 50%, and so on. The two-dimensional image to be detected from the main perspective containing multiple objects can be obtained by taking pictures using a single top-view camera. There are many ways to extract the image. The camera shooting picture from the main perspective can be segmented to crop out a region of interest of a fixed size, and the region of interest must contain the entire fruit fly, as shown in FIG8 .
[0044] In step 105 , based on the extracted region of interest containing the target object, the trained second learning network is used to detect the three-dimensional positions of the key points of the limbs of the target object.
[0045] In step 106 , the dynamic three-dimensional posture of the target object is determined based on the detected three-dimensional positions of the limb key points of the target object at different moments.
[0046] Using a method for detecting the three-dimensional posture of a dynamically interactive object according to various embodiments of the present application, the method processes and integrates the two-dimensional positions of the limb key points of at least one object in a first set of primary-view and multiple secondary-view training two-dimensional images to obtain the three-dimensional positions of the first set of limb key points of the at least one object; and constructs a three-dimensional dataset using the regions of interest of the at least one object and the three-dimensional positions of the first set of limb key points of the at least one object in the first set of primary-view training two-dimensional images. The thus constructed three-dimensional dataset reflects the mapping relationship between the regions of interest of the objects in the primary-view two-dimensional images and the three-dimensional positions of the limb key points of the objects, and is used to train a second learning network, thereby allowing the second learning network to learn the mapping relationship. When actually detecting the dynamic three-dimensional positions of the limb key points of a target object dynamically interacting with other objects, a simple capture system using a single primary-view camera, including a primary-view two-dimensional image of multiple objects to be detected, can accurately and quickly (including in real time) detect the three-dimensional positions of the limb key points of any target object at each capture moment during the process of multiple objects freely and dynamically interacting with each other, despite complex and changing occlusion relationships between them, thereby determining the dynamic three-dimensional posture of the target object during this process.
[0047] As shown in Figure 5, the processing integration for obtaining the three-dimensional positions of the first group of limb key points of the at least one object can further include the following steps. In step 501, based on the simultaneous training two-dimensional images of the main perspective, the occlusion of the at least one object in the training two-dimensional images of each slave perspective at the corresponding time is evaluated, and the two-dimensional positions of the limb key points corresponding to the slave perspectives where the occlusion of the at least one object reaches a serious condition are removed. As shown in Figure 6, in the top-view shot of the two fruit flies, the situation in which the target object fruit fly below is occluded by other fruit flies above in each side perspective can be evaluated, and some side perspective shots with serious occlusion (as shown in Figure 7, referred to as occluded shots) and the two-dimensional positions of the limb key points marked therein are removed. Since the target object is subsequently detected based on the top-view footage, the two-dimensional positions of the limb key points of the target object fruit fly in the top-view footage to be detected usually also show similar occlusion in the corresponding side-view. That is, the two-dimensional positions of the limb key points in the training two-dimensional image that are occluded in the side-view from the main perspective will also be occluded in the corresponding side-view during actual detection. Therefore, the two-dimensional positions of the limb key points corresponding to the perspectives where the occlusion of the at least one object reaches a serious condition are removed. In step 502, based on the remaining two-dimensional positions of the limb key points of the at least one object, the three-dimensional positions of the limb key points of the first group of the at least one object are integrated to obtain the three-dimensional positions of the limb key points of the first group of the at least one object. The obtained mapping relationship between the three-dimensional positions of the limb key points of the first group of the at least one object and the region of interest containing the at least one object in the first group of the main-view training two-dimensional images is more consistent with the mapping relationship between the main-view footage during actual detection and the actual three-dimensional position of the target object therein in terms of spatial occlusion relationship. When used to train the second learning network, the second learning network can better learn the spatial occlusion relationship, making the detected three-dimensional position of the target object more accurate and reasonable. When constructing a three-dimensional dataset, especially when integrating the three-dimensional positions of the limb key points used in the three-dimensional dataset, a top-view camera is used to evaluate the occlusion situation, and the camera detection results corresponding to the perspective with severe occlusion are removed to prevent the detection errors caused by occlusion from affecting the results of the triangulation processing (described in detail below), so as to make the three-dimensional key point detection more accurate.
[0048] In some embodiments, the 3D pose detection method may further include screening the 2D positions of the at least one object's limb key points in the detected 2D training image based on their confidence levels, whereby only the 2D positions of the limb key points with confidence levels above a threshold are used to integrate the 3D positions of the at least one object's limb key points. In this way, the 2D positions of the limb key points used for integration all have a high confidence level, resulting in a more accurate and reasonable 3D position of the at least one object's limb key points.
[0049] Based on the remaining two-dimensional positions of the limb key points of the at least one object, the three-dimensional positions of the limb key points of the first group of the at least one object are integrated. This integration process can also be generally referred to as triangulation processing, which means that for a three-dimensional coordinate point, the three-dimensional coordinate position is estimated through the two-dimensional positions projected from multiple perspectives. Triangulation processing is a mature machine vision algorithm with multiple implementation methods, which will not be described here. Triangulation processing is performed using the two-dimensional key points provided by the multi-view camera array to obtain accurate three-dimensional key points, thereby obtaining a high-quality interactive object behavior dataset whose diversity exceeds that of a single-object behavior dataset.
[0050] In some embodiments, as shown in FIG9 , constructing the three-dimensional dataset specifically includes the following steps. In step 901, a centering transformation is performed on the region of interest containing the at least one object, such that the at least one object faces a predetermined direction and is centered within the region of interest containing the at least one object. Referring to FIG8 , after the centering transformation of the region of interest, the fruit fly, as an object, is oriented head-up and centered within the region of interest. In step 902, a corresponding centering transformation is also performed on the three-dimensional positions of the key points of the limbs of the at least one object. For example, the centering transformation in step 901 can be performed using a centering transformation matrix. Then, the same centering transformation matrix can be applied to the three-dimensional positions of the key points of the limbs of the at least one object. In step 903, the transformed three-dimensional positions of the key points of the limbs of the at least one object and the transformed region of interest of the at least one object are used to construct the three-dimensional dataset. In each piece of three-dimensional data constructed in this manner, the fruit fly in the region of interest is oriented head-up and centered. By unifying the orientation and position of the fruit fly in the region of interest, the computational complexity of training the second learning network is significantly reduced, making training more efficient.
[0051] In some embodiments, as shown in FIG10 , a centering transformation can also be applied to the actual detection of the three-dimensional positions of the target object's limb key points. Specifically, in step 1001, a centering transformation is performed on the region of interest containing the target object so that the target object faces the predetermined direction and is centered within the region of interest containing the target object.
[0052] In step 1002, based on the transformed region of interest containing the target object, the trained second learning network is used to detect the corresponding transformed three-dimensional positions of the limb key points of the target object. Figure 11 shows a comparative diagram of the region of interest containing the target object after the centering transformation and the three-dimensional positions of the limb key points of the target object after the superimposed centering transformation according to an embodiment of the present application. As shown in Figure 11, after the superimposed centering transformation, the target object is centered in the region of interest, and the corresponding transformed three-dimensional positions of the limb key points of the target object are also centered in the region of interest after two-dimensional projection on the image plane. In step 1003, the inverse transformation of the centering transformation is performed on the transformed three-dimensional positions of the limb key points of the target object to obtain the three-dimensional positions of the limb key points of the target object as the detection result. In some embodiments, through the above processing, the three-dimensional positions of the limb key points of the target object can be automatically identified on target objects of various orientations and positions on the shooting screen of the main perspective. By aligning the target object with the predetermined direction and centered within the region of interest (ROI) containing the target object, the second learning network improves its 3D position detection efficiency and reduces detection complexity. Furthermore, the inverse transformation results in more accurate 3D position detection of key points on the target object's limbs. By extracting the ROI from a large primary perspective and performing a coordinated centralization-inverse transformation to generate detection results, the detection process is faster and more accurate, approaching real-time. This is particularly advantageous for determining the 3D poses of multiple target objects that change rapidly or even unpredictably.
[0053] FIG12 shows a schematic block diagram of a system for detecting a three-dimensional posture of a dynamic interactive object according to an embodiment of the present application. As shown in FIG12 , the three-dimensional posture detection system includes a single top-view camera 1201 and at least one processor 1202. The single top-view camera 1201 can be configured to capture two-dimensional images of a main perspective of multiple objects at different times as two-dimensional images to be detected. The at least one processor 1202 can be configured to execute the three-dimensional posture detection method for a dynamic interactive object according to various embodiments of the present application. Various examples and details of the above methods can be selectively incorporated herein and are not described in detail here. That is, using the simple configuration of a single top-view camera 1201, it is possible to utilize the two-dimensional images to be detected from the main perspective of multiple objects. In the process of multiple objects freely and dynamically interacting with each other, despite the complex and changeable occlusion relationships between each other, it is still possible to accurately and quickly (including in real time) identify and detect the three-dimensional positions of the key points of the limbs of any target object at each shooting moment, thereby determining the dynamic three-dimensional posture of the target object during this process.
[0054] In some embodiments, the at least one processor 1202 may be a processing device, including one or more general-purpose processing devices such as a microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), etc. More specifically, the at least one processor 1202 may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor that implements other instruction sets, or a processor that implements a combination of instruction sets. The at least one processor 1202 may also be one or more special-purpose processing devices, such as an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a system on a chip (SOC), etc. As will be appreciated by those skilled in the art, in some embodiments, the at least one processor 1202 may be a special-purpose processor rather than a general-purpose processor. The at least one processor 1202 may include one or more known processing devices.
[0055] In some embodiments, a device for detecting a three-dimensional posture of a dynamic interactive object is further provided, comprising at least one processor configured to execute the method for detecting a three-dimensional posture of a dynamic interactive object according to various embodiments of the present application. Various examples and details of the above method may be selectively incorporated herein and are not further elaborated upon.
[0056] In some embodiments, a computer-readable storage medium is further provided, storing an executable program thereon. When executed by a processor, the executable program implements the three-dimensional posture detection method for a dynamic interactive object according to various embodiments of the present application. Various examples and details of the above methods may be selectively incorporated herein and are not further described here.
[0057] In this document, the term "or" is used to mean a non-exclusive "or", and unless otherwise stated, "A or B" includes "A without B", "B without A", and "A and B". In this document, the terms "including" and "in which" are used as plain-language equivalents of the respective terms "comprising" and "wherein". In addition, in the following claims, the term "comprising" is open-ended, that is, devices, systems, apparatuses, articles, compositions, concepts or processes that include elements in addition to those elements listed after such term in the claim are still deemed to fall within the scope of the claim. In addition, in the following claims, the terms "first", "second", and "third", etc. are used merely as labels and are not intended to impose numerical requirements on their objects.
[0058] The method examples described herein may be machine or computer-implemented, at least in part. Some examples may include a computer-readable medium or machine-readable medium encoded with instructions that can be operated to configure an electronic device to perform the methods described in the above examples. The implementation of these methods may include software codes such as microcode, assembly language code, higher-level language code, etc. Various software programming techniques may be used to create various programs or program modules. For example, a program segment or program module may be designed using or through Java, Python, C, C++, assembly language, or any known programming language. One or more such software segments or modules may be integrated into a computer system and / or computer-readable medium. Such software code may include computer-readable instructions for performing various methods. Such software code may form part of a computer program product or computer program module. In addition, in one example, for example, during runtime or at other times, the code may be tangibly stored on one or more volatile, non-transitory, or non-volatile tangible computer-readable media. Examples of such tangible computer-readable media may include, but are not limited to, hard disks, removable disks, removable optical disks (e.g., compact discs and digital video disks), magnetic tapes, memory cards or sticks, random access memory (RAM), read-only memory (ROM), etc.
[0059] In addition, although exemplary embodiments have been described herein, the scope includes any and all embodiments based on the present disclosure with equivalent elements, modifications, omissions, combinations (e.g., of the various embodiments), adaptations, or changes. The elements in the claims are to be interpreted broadly based on the language employed in the claims and are not limited to the examples described in this specification or during the prosecution of this application, which examples are to be interpreted as non-exclusive. In addition, the steps of the disclosed methods may be modified in any way, including by reordering the steps or inserting or deleting steps. Therefore, this specification and examples are intended to be considered as examples only, with the true scope and spirit being indicated by the following claims and the full scope of their equivalents.
[0060] The above description is intended to be illustrative and not restrictive. For example, the above examples (or one or more of their solutions) may be used in combination with each other. For example, a person of ordinary skill in the art may use other embodiments when reading the above description. In addition, in the above-mentioned specific embodiments, various features may be grouped together to simplify the present disclosure. This should not be interpreted as an intention to make the disclosed features that are not required to be protected necessary for any claim. Instead, the subject matter of the present invention may be less than all the features of the disclosed specific embodiments. Therefore, the following claims are incorporated into the specific embodiments as examples or embodiments, wherein each claim itself is independently a separate embodiment, and it is believed that these embodiments can be combined with each other in various combinations or arrangements. The scope of the present invention should be determined with reference to the appended claims and the full scope of equivalents to which these claims are entitled.
Claims
1. A three-dimensional posture detection method for a dynamic interactive object, wherein: Multiple objects dynamically interact with each other, which is characterized by: Based on the two-dimensional positions of the key points of the limbs of at least one object in the first set of training two-dimensional images from the primary perspective and the plurality of secondary perspectives, processing and integrating them to obtain the three-dimensional positions of the key points of the limbs of the first set of the at least one object; constructing a three-dimensional dataset using the first set of two-dimensional training images of the primary view including the region of interest of the at least one object and the three-dimensional positions of the first set of limb key points of the at least one object; Use the constructed three-dimensional dataset to train the second learning network; Extracting a region of interest containing a target object based on a two-dimensional image to be detected from a primary perspective containing multiple objects; Based on the extracted region of interest containing the target object, using the trained second learning network to detect the three-dimensional positions of the limb key points of the target object; and The dynamic three-dimensional posture of the target object is determined based on the detected three-dimensional positions of the target object's limb key points at different times.
2. The three-dimensional posture detection method according to claim 1, characterized in that: The two-dimensional positions of the limb key points of at least one object in the first set of two-dimensional training images from the main view and multiple secondary view are obtained by the following processing: Constructing a two-dimensional dataset using a subset of the first group of training two-dimensional images from a primary perspective and multiple secondary perspectives and the annotation results of limb key points of at least one object in each of the training two-dimensional images; Using the constructed two-dimensional dataset to train a first learning network; Based on the first group of training two-dimensional images from the main perspective and multiple training two-dimensional images from the secondary perspective, the trained first learning network is used to detect the two-dimensional positions of the limb key points of at least one object in the first group of training two-dimensional images from the main perspective and multiple training two-dimensional images from the secondary perspective.
3. The three-dimensional posture detection method according to claim 1, characterized in that: The processing integration for obtaining the three-dimensional positions of the first group of limb key points of the at least one object specifically includes: Based on the simultaneous training two-dimensional images from the main perspective, evaluating the occlusion of at least one object in the training two-dimensional images from the slave perspectives at the corresponding time, and removing the two-dimensional positions of the limb key points corresponding to the slave perspectives where the occlusion of the at least one object reaches a severe condition; and Based on the remaining two-dimensional positions of the limb key points of the at least one object, the three-dimensional positions of the first group of limb key points of the at least one object are integrated.
4. The three-dimensional posture detection method according to claim 2 or 3, characterized in that: Also includes: The two-dimensional positions of the limb key points of the at least one object in the detected two-dimensional training image are screened according to their confidence levels, and only the two-dimensional positions of the limb key points with confidence levels higher than a threshold are used to integrate and obtain the three-dimensional positions of the limb key points of the at least one object.
5. The three-dimensional posture detection method according to claim 1, characterized in that: The construction of the three-dimensional data set specifically includes: performing a centering transformation on a region of interest containing the at least one object so that the at least one object faces a predetermined direction and is centered on the region of interest containing the at least one object; Performing corresponding centralization transformation on the three-dimensional position of the limb key point of the at least one object; The three-dimensional dataset is constructed using the transformed three-dimensional positions of the limb key points of the at least one object and the transformed region of interest of the at least one object.
6. The three-dimensional posture detection method according to claim 1, characterized in that: Based on the extracted region of interest containing the target object, the trained second learning network is used to detect the three-dimensional positions of the limb key points of the target object, specifically including: Performing a centering transformation on the region of interest containing the target object so that the target object faces the predetermined direction and is centered in the region of interest containing the target object; Based on the transformed region of interest containing the target object, using the trained second learning network to detect the corresponding transformed three-dimensional positions of the limb key points of the target object; The inverse transformation of the centering transformation is performed on the transformed three-dimensional positions of the limb key points of the target object to obtain the three-dimensional positions of the limb key points of the target object as the detection results.
7. The three-dimensional posture detection method according to claim 1, characterized in that: The main viewing angle is a top viewing angle, and the secondary viewing angle is a side viewing angle.
8. The three-dimensional posture detection method according to claim 1, characterized in that: The at least one object is located in the middle of the training two-dimensional image, and the target object is different from the at least one object.
9. The three-dimensional posture detection method according to claim 1, characterized in that: The target object-containing region of interest is extracted to contain only a single target object, or to make the single target object centered and occupy a predetermined ratio or more of the area of the region of interest.
10. The three-dimensional posture detection method according to claim 1, characterized in that: The main-view two-dimensional image to be detected containing multiple objects is obtained by photographing using a single top-view camera.
11. The three-dimensional posture detection method according to claim 1, characterized in that: The first learning network and the second learning network both include convolutional neural networks.
12. The three-dimensional posture detection method according to claim 1, characterized in that: The object includes any one of an animal, a movable facility, and a robot.
13. A three-dimensional posture detection system for dynamic interactive objects, characterized in that: include: a single top-view camera configured to capture two-dimensional images of a main viewpoint containing multiple objects at different times as two-dimensional images to be detected; as well as At least one processor is configured to execute the three-dimensional posture detection method of a dynamic interactive object according to any one of claims 1-12.
14. A three-dimensional posture detection device for a dynamic interactive object, characterized in that: The method comprises at least one processor configured to execute the three-dimensional posture detection method of a dynamic interactive object according to any one of claims 1-12.
15. A computer-readable storage medium having an executable program stored thereon, wherein when the executable program is executed by a processor, the method for detecting a three-dimensional posture of a dynamic interactive object according to any one of claims 1 to 12 is implemented.
Citation Information
Patent Citations
Target three-dimensional key point extraction model construction and posture recognition method in two-dimensional graph
CN110634160A
Motion capture method and related equipment
CN115311472A
Method and device for constructing three-dimensional attitude estimation data set based on multiple view angles
CN115841602A
Automated data capture
US10417781B1