Method, apparatus and storage medium for three-dimensional skeleton reconstruction
Patent Information
- Application Number
- CN202610914015.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2046-06-24
AI Technical Summary
[0003]本申请的主要目的在于提供一种三维骨架重建方法、设备和存储介质,旨在解决对实验场地适配性差的技术问题
[0012] This application provides a 3D skeleton reconstruction method. It involves inputting multiple animal images acquired from various observation perspectives into a lightweight detection model in parallel. The method obtains 2D keypoints and their confidence levels for each observation perspective in real time. Based on the confidence levels, a validity indicator for each 2D keypoint is determined. When the validity indicator determines that a certain 2D keypoint has valid 2D observations in at least two observation perspectives, a corresponding 3D keypoint is generated based on that 2D keypoint. These 3D keypoints are then connected according to a pre-defined skeleton topology to generate a real-time 3D skeleton. By utilizing the validity indicator to filter 2D observations, this method does not rely on fixed camera setups and viewpoint combinations. Even when the size, layout, or type of the experimental site changes, the method can identify the actual valid observation perspectives in the current location using only the validity indicator, thus enabling 3D reconstruction. This method demonstrates high adaptability to different experimental sites.
Smart Images

Figure CN122435168B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of animal behavior analysis technology, and in particular to a three-dimensional skeleton reconstruction method, device and storage medium. Background Technology
[0002] Currently, analyzing animal behavior can reveal their underlying neural activity mechanisms, cognitive functions, disease states, and drug efficacy. Existing 3D reconstruction systems prioritize offline high precision; therefore, camera deployment, calibration procedures, and algorithms are designed for fixed geometric relationships in specific experimental environments. In complex experimental scenarios or with non-standard setups, certain camera perspectives are easily obstructed, or the detection model may experience brief false positives. This indicates poor adaptability to the size, layout, and type of experimental sites, thus the solutions described in related technologies suffer from poor adaptability. Summary of the Invention
[0003] The main purpose of this application is to provide a three-dimensional skeleton reconstruction method, device and storage medium, which aims to solve the technical problem of poor adaptability to experimental sites.
[0004] To achieve the above objectives, this application provides a three-dimensional skeleton reconstruction method, which includes: Multiple animal images acquired from multiple observation perspectives are input into a lightweight detection model in parallel to obtain the two-dimensional key points and the confidence level of the two-dimensional key points for each observation perspective in real time. Determine the confidence threshold for each two-dimensional key point, where the confidence threshold differs for different locations; If the confidence level of a two-dimensional keypoint is greater than or equal to the confidence threshold, the two-dimensional keypoint is associated with the effective indicator. Otherwise, associate the two-dimensional keypoints with invalid indicators, where valid indicators include both valid and invalid indicators; Determine the two-dimensional displacement between the two-dimensional key point and the corresponding target two-dimensional key point in the previous frame; Identify abnormal two-dimensional key points whose two-dimensional displacement is greater than or equal to a preset two-dimensional displacement threshold; When the validity indicator determines that there are valid 2D keypoints in at least two observation views for the actual 3D keypoints, the corresponding target 3D keypoints are generated based on the valid 2D keypoints. The steps for generating the corresponding target 3D keypoints based on the valid 2D keypoints include any of the following: generating target 3D keypoints based on other 2D keypoints besides the abnormal 2D keypoints; determining the smoothed 2D keypoints corresponding to the abnormal 2D keypoints and the target 2D keypoints, and generating target 3D keypoints based on the smoothed 2D keypoints and other 2D keypoints; reducing the construction weights corresponding to the abnormal 2D keypoints based on a preset strategy, and constructing target 3D keypoints by weighting based on the updated construction weights. Connect the target's 3D key points according to the preset skeleton topology to generate a real-time 3D skeleton.
[0005] In one embodiment, multiple animal images acquired from multiple observation perspectives are input in parallel into a lightweight detection model to obtain the two-dimensional keypoints and their confidence levels for each observation perspective in real time, including: Each observation perspective is assigned an independent processing thread, and a model instance of the same lightweight detection model is loaded in each processing thread. Each processing thread reads the animal image from the corresponding observation viewpoint, performs model inference based on the model instance, and outputs the two-dimensional keypoints and their confidence scores for each observation viewpoint in parallel.
[0006] In one embodiment, generating corresponding target 3D key points based on valid 2D key points includes: Randomly select at least two valid two-dimensional key points corresponding to observation viewpoints, and calculate candidate three-dimensional key points; Determine the reprojection error of candidate 3D keypoints at the observation viewpoint; When the reprojection error is less than the error threshold, it is determined that the observation viewpoint is consistent with the candidate 3D keypoint. The candidate 3D keypoint with the most consistent observation perspectives is selected as the target 3D keypoint.
[0007] In one embodiment, after selecting the candidate 3D keypoint with the most consistent observation viewpoints as the target 3D keypoint, the method further includes: The weights of each observation perspective are determined based on the confidence level of the effective two-dimensional key points; The target's 3D key points are then refined using a weighted approach to obtain the refined 3D key points.
[0008] In one embodiment, generating a real-time 3D skeleton by connecting target 3D key points according to a preset skeleton topology includes: Based on the target's 3D key points and the target's 3D key points in the previous frame, calculate the smoothed target's 3D key points. Calculate the 3D displacement between the target 3D keypoint and the target 3D keypoint in the previous frame. When the 3D displacement is less than or equal to the 3D displacement threshold, determine the target 3D keypoint as the corrected target 3D keypoint.
[0009] In one embodiment, after generating a real-time 3D skeleton by connecting the target 3D key points according to a preset skeleton topology, the method further includes: Based on a real-time 3D skeleton, pose and motion features are extracted to form feature vectors; The feature vector is used to determine or decode its state to obtain the current behavior state or motion intention; Control signals are generated based on the current behavior state or motion intention, and the control signals are output to external devices for control.
[0010] In addition, to achieve the above objectives, this application also provides a three-dimensional skeleton reconstruction device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the above-described three-dimensional skeleton reconstruction method.
[0011] In addition, to achieve the above objectives, this application also provides a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the above-described three-dimensional skeleton reconstruction method.
[0012] This application provides a 3D skeleton reconstruction method. It involves inputting multiple animal images acquired from various observation perspectives into a lightweight detection model in parallel. The method obtains 2D keypoints and their confidence levels for each observation perspective in real time. Based on the confidence levels, a validity indicator for each 2D keypoint is determined. When the validity indicator determines that a certain 2D keypoint has valid 2D observations in at least two observation perspectives, a corresponding 3D keypoint is generated based on that 2D keypoint. These 3D keypoints are then connected according to a pre-defined skeleton topology to generate a real-time 3D skeleton. By utilizing the validity indicator to filter 2D observations, this method does not rely on fixed camera setups and viewpoint combinations. Even when the size, layout, or type of the experimental site changes, the method can identify the actual valid observation perspectives in the current location using only the validity indicator, thus enabling 3D reconstruction. This method demonstrates high adaptability to different experimental sites. Attached Figure Description
[0013] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0014] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is a flowchart illustrating an embodiment of the three-dimensional skeleton reconstruction method of this application. Figure 2 A schematic diagram of the data acquisition provided for the three-dimensional skeleton reconstruction method of this application; Figure 3This is a schematic diagram of the overall process of the three-dimensional skeleton reconstruction method provided in the embodiments of this application; Figure 4 A simplified flowchart illustrating the three-dimensional skeleton reconstruction method provided in this application embodiment; Figure 5 This is a schematic diagram of the structure of the three-dimensional skeleton reconstruction device in the embodiments of this application.
[0016] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0017] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application. To better understand the technical solutions of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0018] Currently, analyzing animal behavior can reveal its underlying neural mechanisms, cognitive functions, disease states, and drug efficacy. Existing 3D reconstruction systems prioritize offline high accuracy; therefore, camera deployment, calibration procedures, and algorithms are designed for fixed geometric relationships in specific experimental environments. In complex experimental scenarios or non-standard setups, certain camera views are easily obstructed, or the detection model may experience brief false positives. This indicates poor adaptability to the size, layout, and type of experimental sites, resulting in poor adaptability of the solutions described in related technologies. Furthermore, while existing multi-camera-based 3D reconstruction schemes can obtain relatively complete pose information from multi-view images, they are mostly computationally intensive, have heavy models, and high processing latency. They are more suitable for offline analysis and struggle to meet the requirements of high frame rate, low latency, and stability. They typically only focus on 3D skeleton reconstruction or behavior recording, lacking effective integration with real-time closed-loop applications such as drug administration, stimulus detection, or motion intent decoding, thus failing to meet the application needs of real-time monitoring, intervention, and decoding in animal experiments.
[0019] This application uses a lightweight detection model to input multiple animal images acquired from various observation perspectives in parallel. It obtains the two-dimensional keypoints and their confidence levels for each observation perspective in real time. Based on the confidence levels, it determines the validity indicator of the two-dimensional keypoints. When the validity indicator determines that a certain two-dimensional keypoint has valid two-dimensional observations in at least two observation perspectives, it generates a corresponding three-dimensional keypoint. The three-dimensional keypoints are then connected according to a pre-defined skeleton topology to generate a real-time three-dimensional skeleton. By using the validity indicator to filter two-dimensional observations, this method does not rely on fixed camera setups and viewpoint combinations. Even when the size, layout, or type of the experimental site changes, the validity indicator alone can identify the actual valid observation perspectives in the current location, enabling three-dimensional reconstruction. This demonstrates a high degree of adaptability to different experimental sites.
[0020] Based on this, Embodiment 1 of this application proposes a three-dimensional skeleton reconstruction method, please refer to... Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the three-dimensional skeleton reconstruction method of this application. The three-dimensional skeleton reconstruction method includes steps S10 to S40: Step S10 involves inputting multiple animal images collected from various observation perspectives into a lightweight detection model in parallel to obtain the two-dimensional key points and their confidence levels for each observation perspective in real time.
[0021] Image coordinates refer to the position of a 2D keypoint in its corresponding camera image coordinate system, represented by a tuple (u, v), where u is the x-coordinate and v is the y-coordinate. A 2D keypoint is a pre-defined location in an image representing a specific part of an animal's body, such as the tip of the nose, ears, limb joints, rump, or tail root. Confidence refers to the degree of certainty the lightweight detection model has regarding the predicted position of each 2D keypoint; it is a scalar value between [0, 1]. A higher confidence level indicates a more reliable prediction.
[0022] In this embodiment, multiple animal images acquired from various observation perspectives are input in parallel into a lightweight detection model for keypoint prediction. Specifically, two-dimensional keypoint detection is performed on each image acquired from multiple perspectives, and multiple two-dimensional keypoints in the current image are output. Simultaneously, the model can be adapted or fine-tuned based on the imaging differences between different perspectives. The image coordinates and confidence score of the j-th keypoint are represented as follows:
[0023] in This represents the information of the j-th two-dimensional keypoint detected from the c-th image. and These represent the horizontal and vertical coordinates in the image plane, respectively. This indicates the confidence level of the key point. , This represents the total number of keypoints. Through a lightweight, integrated model, accurate and complete 2D keypoint observation data is provided for each viewpoint, offering a reliable data foundation for subsequent 3D reconstruction.
[0024] During model inference, a graphics processing unit (GPU) or other parallel computing hardware is used to perform keypoint prediction for each image stream using a lightweight detection model. Independent computing resources are allocated to each image captured by the industrial camera; for example, multiple parallel computing threads are launched on the GPU. Each image stream and its corresponding detection model are simultaneously loaded onto the hardware for inference computation. If the... The 2D inference time for the road image under hardware acceleration is: Then the total time consumed during multi-path parallel inference is approximately:
[0025] The total time for parallel 2D keypoint detection of multiple industrial camera images is determined solely by the inference time of the slowest image stream. As the number of industrial cameras increases, the latency increase in the 2D prediction stage remains manageable. Hardware acceleration and parallel computing significantly reduce the total time for 2D keypoint detection.
[0026] As an alternative implementation, the lightweight detection model is a lightweight convolutional neural network focused on keypoint detection.
[0027] Specifically, lightweight convolutional neural networks (CNNs) offer advantages such as smaller model size, faster inference speed, and readily available pre-trained weights, making them suitable for deployment in multi-view real-time scenarios. An independent processing thread is created for each observation view. After acquiring multiple animal images, the image data from each camera and its corresponding processing thread are bound together. Each thread independently loads and runs an instance of the same lightweight CNN, responsible for performing forward inference on that image, predicting the image coordinates and confidence scores of two-dimensional keypoints in the animal's body. By using lightweight CNNs, the requirements for high frame rates and low latency in real-time processing can be met, while the mature pre-trained weights provide more reliable prediction results.
[0028] As another alternative implementation, the lightweight detection model is a lightweight detection-keypoint integrated network, such as YOLO11n-pose (You Only Look Once Version 11 Nano Pose, YOLO version 11 nano-scale pose estimation model).
[0029] Specifically, the lightweight detection-keypoint integrated network design enables it to simultaneously output the target region, 2D coordinates of body keypoints, and confidence scores of the animal in a single forward pass. This results in lower single-frame processing time. Furthermore, pre-trained weights trained on its own dataset guide the model's computation, allowing it to focus on learning visual features highly relevant to the experiment. An independent processing thread is allocated to each observation view, with each thread running a YOLO11n-pose model instance. Within each thread, a single forward pass of the input single-path image simultaneously outputs the image coordinates and confidence scores of the target region and 2D keypoints. The target region, i.e., the bounding box or region of interest identifying the animal's position in the image, assists the model in locating the animal and improves the stability and accuracy of keypoint detection. By using the YOLO11n-pose model for 2D keypoint prediction, the intra-thread serial latency and additional intermediate data transmission overhead associated with first detecting the target and then predicting keypoints based on the detection box are avoided, further improving the overall efficiency of the multi-view real-time processing link.
[0030] In addition, please refer to Figure 2 , Figure 2 This diagram illustrates the acquisition process for the 3D skeleton reconstruction method described in this application. It is primarily used for 3D skeleton reconstruction of rodents, such as mice and rats, but can be extended to other animals. Before the experiment, the animal is placed on a platform, and a three-dimensional frame equipped with industrial cameras is constructed, forming a surround-image array from the top and four sides to capture the animal from multiple angles. Furthermore, the height of the frame is adjustable, for example, from 30cm to 2000cm, allowing for flexible adjustment of the spatial scale according to different experimental scenarios and animal sizes. During animal image acquisition, multiple industrial cameras are used to acquire images of the target animal from different perspectives, with each image corresponding to a different viewpoint. A multiple-channel industrial camera refers to a dedicated industrial imaging device with at least two channels that acquire images of the target animal from different spatial perspectives. The platform on which the animal is placed can be flat, inclined, stepped, or a platform with a specific structure. The surround-image array can be deployed on the top and four sides, or at the bottom. The height of the frame can be less than 30cm or more than 2000cm; this embodiment does not impose specific limitations. Assume the number of industrial cameras is C, where C≥2, indicating that images of the target animal are acquired from at least two different perspectives. For the c-th industrial camera, c is used to identify and distinguish different industrial cameras, and is an integer starting from 1 up to the total number of cameras C.
[0031] Specifically, the self-integrated Python algorithm acquires image data by calling the API interface provided by the industrial camera SDK (Software Development Kit), which internally interfaces with the camera driver interface at the operating system level. The image acquired by the c-th industrial camera at time t is denoted as... After images are acquired from each industrial camera, each image is permanently bound to that camera and its corresponding timestamp is recorded. All industrial cameras are bound to the calibrated cameras in a pre-defined, fixed order to ensure consistent data flow during subsequent processing and prevent geometric errors caused by out-of-order sequences. Simultaneously, an independent acquisition thread is created for each industrial camera to enable parallel acquisition of multiple images. In parallel acquisition, the total acquisition time does not increase linearly with the number of industrial cameras, significantly reducing the overall acquisition time and meeting the real-time requirements of high frame rates. Furthermore, a limited buffer queue is set up for each camera, retaining only the latest frame or a few recent frames to prevent data backlog when backend processing speed lags temporarily, keeping overall latency within a limited range and ensuring that real-time output closely reflects the animal's current state.
[0032] Step S20: Determine the validity indicator of the two-dimensional key points based on the confidence level.
[0033] The validity indicator is a variable used to mark whether the observation data of each two-dimensional keypoint is valid under each camera view.
[0034] In this embodiment, a validity indicator is constructed to determine whether a two-dimensional key point j is valid in the c-th view.
[0035] in Whether a point is valid is determined by comparing its confidence level with a preset confidence threshold. A point is considered valid when its confidence level is not less than the preset threshold. When the coordinates are valid, the validity indicator is 1; otherwise, it is 0. By filtering out unreliable detection results with low confidence, low-quality data can be prevented from interfering with subsequent 3D calculations.
[0036] As an optional implementation, a uniform confidence threshold is used to determine the validity of all two-dimensional key points.
[0037] Specifically, a global confidence threshold is preset. For the detection result of the c-th viewpoint and the j-th keypoint, its confidence level is compared with this global confidence threshold. If the confidence level of the two-dimensional keypoint is not less than the global confidence threshold, the two-dimensional keypoint is determined to be valid, and the corresponding validity indicator is set to 1; if the confidence level of the two-dimensional keypoint is less than the global confidence threshold, the two-dimensional keypoint is determined to be invalid, and the corresponding validity indicator is set to 0. By using a unified confidence threshold, the validity indicator can be quickly determined.
[0038] As another alternative implementation, a differentiated confidence threshold is used to determine the validity indicator for each two-dimensional key point.
[0039] Specifically, different thresholds are preset for different types of 2D keypoints. A pre-defined table showing the relationship between 2D keypoint types and confidence thresholds is read. For example, a relatively high confidence threshold is used for core keypoints such as the head, while a relatively low confidence threshold is used for keypoints that are easily occluded, such as the tail. Based on the type of 2D keypoint j, the corresponding threshold is selected from the table and then compared with the confidence score of the 2D keypoint. If the confidence score of the 2D keypoint is not less than the corresponding confidence threshold, the 2D keypoint is considered valid, and a validity indicator is generated. By setting differentiated confidence thresholds, the integrity and accuracy of the 3D skeleton can be better balanced.
[0040] Step S30: When it is determined based on the validity indicator that there are valid two-dimensional key points in at least two observation views, the corresponding target three-dimensional key points are generated based on the valid two-dimensional key points.
[0041] Valid two-dimensional observation points are two-dimensional key points from a specific camera viewpoint that have been screened for confidence levels.
[0042] In this embodiment, for each 2D keypoint, all camera viewpoints are traversed, and the validity indicator value under each viewpoint is checked one by one. The camera viewpoint numbers with a validity indicator value of 1 are collected to form a special set, which is the valid 2D observation set of that keypoint, denoted as .
[0043] Let be the set of indices of all valid industrial cameras that can observe a 2D keypoint j at time t. This set will serve as the direct basis for determining whether the conditions for 3D reconstruction are met and which viewpoint data to use for triangulation calculation. Based on the set of valid 2D observations, it is determined whether the 2D keypoint has valid 2D observations from two or more viewpoints. If two or more viewpoints exist, 3D reconstruction is performed based on the 2D keypoint. Check the number of elements in the set, i.e., the number of valid viewpoints. The number of valid viewpoints is determined if and only if there are at least two valid viewpoints. When a 2D keypoint is determined to satisfy the geometric constraints of 3D reconstruction, a corresponding 3D keypoint is generated based on that 2D keypoint. The 3D keypoints obtained from solving all the 2D keypoints are then integrated to obtain all the 3D keypoints.
[0044] in Let J be the set of 3D coordinates of all J keypoints at time t, which is a Jx3 matrix. Through random sampling and consistency verification mechanisms, abnormal observation data caused by misdetection from individual viewpoints, occlusion, or noise are effectively eliminated, thus stably calculating accurate 3D keypoint coordinates even in complex experimental scenarios. By determining the number of 2D keypoint viewpoints, the system focuses only on which cameras provided valid data at the current time and experimental site, adapting to any camera deployment method.
[0045] As an alternative implementation, based on the projection matrix of each industrial camera, the two-dimensional key points are triangulated using the RANSAC (Random Sample Consensus) method to obtain the three-dimensional key points.
[0046] Specifically, for each 2D keypoint, a minimum subset of observations containing two different viewpoints is randomly selected multiple times from its corresponding valid 2D observations. Multiple candidate points are calculated using the projection matrix corresponding to the subset. For each candidate point, its reprojection error across all valid viewpoints is calculated and compared with a preset threshold to construct a consensus set supporting that point. Through iterative iteration, the candidate point with the most consensus viewpoints is selected as the 3D keypoint. By employing the RANSAC method for triangulation, abnormal observation data caused by temporary occlusion or false detections from a single viewpoint can be eliminated through random sampling and consensus verification mechanisms, thus enabling accurate calculation of 3D keypoint coordinates even when some viewpoint data is unreliable.
[0047] As another alternative implementation, based on the projection matrix of each industrial camera, a weighted multi-view reprojection error minimization method is used to perform triangulation to obtain the three-dimensional key points.
[0048] Specifically, the effective two-dimensional observations and their confidence levels of the two-dimensional keypoint are obtained. Different weights are assigned based on the confidence levels of observations from different viewpoints; that is, observations with higher confidence levels receive greater weights. An optimization algorithm is used to find a three-dimensional spatial point that minimizes the sum of weighted errors between its theoretical and actual observed positions projected onto the effective camera images. The three-dimensional spatial point with the smallest sum of weighted errors is the three-dimensional keypoint. By assigning greater optimization weights to high-confidence two-dimensional observations, the three-dimensional reconstruction results rely more heavily on reliable observation data, thus improving the accuracy of the three-dimensional keypoint coordinate calculation.
[0049] Step S40: Connect the target 3D key points according to the preset skeleton topology relationship to generate a real-time 3D skeleton.
[0050] Skeletal topology refers to the set of rules describing how key points of an animal's body should be connected to form a complete skeleton. For example, it stipulates that the tip of the nose connects to the neck, and the neck connects to the back. The three-dimensional skeleton refers to the spatial structure representing the complete body posture of an animal, formed by connecting three-dimensional key points with line segments according to skeletal topology.
[0051] In this embodiment, based on the predefined skeleton topology Connect the key points in 3D to generate a real-time 3D animal skeleton.
[0052] in This refers to the set of connections between key points. For example, it could include connections such as nose tip to neck, neck to back, back to rump, rump to tail, neck to forelimbs, and rump to hindlimbs. For all 3D keypoints, This results in a three-dimensional skeleton. By connecting discrete three-dimensional key points according to the anatomical structure of an organism, an intuitive three-dimensional skeleton model can be formed.
[0053] As an optional implementation, a pre-defined unified skeleton topology is used for connection to obtain a real-time three-dimensional skeleton.
[0054] Specifically, a fixed skeletal topology is built-in and applied to the target animal. This topology defines which key points need to be connected by line segments, such as nose tip to neck, neck to left anterior shoulder, etc. Each connection definition in the skeletal topology is traversed, and line segments are drawn between the coordinates of the corresponding two 3D key points. This connects these discrete 3D spatial points into a wireframe model that conforms to the anatomical structure of the target animal—a real-time 3D skeleton. Through preset fixed skeletal connection rules, discrete 3D key points are quickly combined into a coherent wireframe model that conforms to the anatomical structure of the target animal, efficiently outputting a standardized 3D pose representation that can be used for subsequent real-time analysis and control.
[0055] As another alternative implementation, different keypoint definitions and topological structures are used for different animals to obtain differentiated real-time 3D skeletons.
[0056] Specifically, a configurable template library is pre-defined, with exclusive keypoint sets and skeletal topological relationships defined for different animal species, such as mice, rats, zebrafish, and fruit flies. For example, for mice or rats, keypoints can be defined as the nose tip, ears, neck, back, rump, limb joints, tail root, and tail, with topological relationships connecting the mammalian trunk and limb structures. For zebrafish, keypoints can be defined as the head, multiple points along the trunk's central axis, and the caudal peduncle, with topological relationships represented by continuous body axis lines. For fruit flies, keypoints can be defined as the head, thorax-abdomen connection points, and coxae of each leg. In actual operation, the corresponding keypoints and topological templates are selected or imported according to the experimental animal species. The resulting 3D keypoints are connected strictly according to the predefined topological connections of the template, thereby generating a 3D skeleton that conforms to the animal's actual anatomical structure in real time. By using different keypoint definitions and topological structures for different animals, the applicability is expanded.
[0057] This embodiment provides a three-dimensional skeleton reconstruction method. First, multiple animal images acquired from various observation perspectives are input in parallel into a lightweight detection model. Two-dimensional keypoints and their confidence levels for each observation perspective are obtained in real time. Based on the confidence levels, a validity indicator for each two-dimensional keypoint is determined. When the validity indicator determines that a certain two-dimensional keypoint has valid two-dimensional observations in at least two observation perspectives, a corresponding three-dimensional keypoint is generated based on that keypoint. These three-dimensional keypoints are then connected according to a preset skeleton topology to generate a real-time three-dimensional skeleton. By utilizing the validity indicator to filter two-dimensional observations, this method does not rely on fixed camera setups and viewpoint combinations. Even when the size, layout, or type of the experimental site changes, the validity indicator alone can identify the actual valid observation perspectives in the current location, enabling three-dimensional reconstruction. This method demonstrates high adaptability to different experimental sites.
[0058] Based on Embodiment 1, in Embodiment 2 of this application, the content that is the same as or similar to that in Embodiment 1 can be referred to the above description, and will not be repeated hereafter. Based on this, the determination of the validity indicator of two-dimensional key points based on confidence level includes: Step S21: Determine the confidence threshold corresponding to each two-dimensional key point, wherein the confidence threshold corresponding to different parts is different.
[0059] As one implementation method, different confidence thresholds are set for 2D keypoints of different body parts. Based on a predefined animal anatomy model, each preset body keypoint is assigned a fixed type label, such as head, nose tip, tail tip, etc. Different types of 2D keypoints correspond to different confidence thresholds. For example, relatively high thresholds are used for core trunk points such as the head, neck, back, and buttocks; relatively low thresholds are used for peripheral points such as ears, paws, and tail tip. The confidence threshold for each 2D keypoint is determined based on the preset confidence threshold corresponding to its type. By using differentiated confidence thresholds based on the 2D keypoint type, the reliability of trunk point data participating in subsequent 3D reconstruction is ensured.
[0060] As one implementation method, a uniform confidence threshold is applied to all 2D keypoints. A global confidence threshold is preset, and for the detection result of each 2D keypoint, its confidence score is compared with the uniform global confidence threshold. By using a uniform confidence threshold, computational overhead is low and data processing speed is fast.
[0061] Step S22: If the confidence level of the two-dimensional key point is greater than or equal to the confidence threshold, associate the two-dimensional key point with the effective indication.
[0062] Step S23: Otherwise, associate the two-dimensional keypoints with invalid indicators, where valid indicators include both valid and invalid indicators.
[0063] As one implementation, the confidence level of a 2D keypoint is compared with a determined confidence threshold. If the confidence level of the 2D keypoint is greater than or equal to the confidence threshold, the 2D keypoint is associated with a valid indicator; for example, if the valid indicator is 1, the 2D keypoint is valid. If the confidence level of the 2D keypoint is less than the confidence threshold, the 2D keypoint is associated with an invalid indicator; for example, if the invalid indicator is 0, the 2D keypoint is invalid. Let the confidence threshold for the j-th keypoint be... The filtering rules are as follows:
[0064] in This represents the result after applying the confidence filtering rules to the j-th key point. This is the preset confidence threshold for the j-th keypoint. The validity indicator is used to determine whether a 2D keypoint j is valid in the c-th view.
[0065] If the confidence level is greater than or equal to the confidence threshold, then Equivalent to the original image coordinates ( If the confidence level is less than the threshold, then the two-dimensional keypoint is valid, meaning it is associated with a valid indicator; if the confidence level is less than the threshold, then... Marked as (NaN, NaN), this indicates that the 2D keypoint is invalid and associated with invalid indicators; it will be considered missing in subsequent steps. Filtering out unreliable detection results with excessively low confidence prevents low-quality data from interfering with subsequent 3D calculations.
[0066] In this embodiment, by setting differentiated confidence thresholds for key points in different parts and generating validity indicators accordingly, observation data can be intelligently filtered, thereby improving the accuracy and robustness of the 3D skeleton output as a whole.
[0067] Based on any of the above embodiments of this application, Embodiment 3 of this application proposes a three-dimensional skeleton reconstruction method, which can be referred to the above description and will not be repeated hereafter. In addition, after determining the validity indicator of the two-dimensional key points based on the confidence level, the method further includes: Step S201: Determine the two-dimensional displacement between the two-dimensional key point and the corresponding target two-dimensional key point in the previous frame.
[0068] Step S202: Identify abnormal two-dimensional key points whose two-dimensional displacement is greater than or equal to a preset two-dimensional displacement threshold.
[0069] As one implementation method, the two-dimensional displacement between the two-dimensional keypoint and the corresponding target two-dimensional keypoint in the previous frame is calculated, and abnormal two-dimensional keypoints are identified. For the same camera and the same keypoint, the two-dimensional displacement between adjacent frames is defined as:
[0070] in, For the same keypoint j, within the same camera c (i.e., the same viewpoint), the two-dimensional pixel displacement between two adjacent frames. Exceeding the preset two-dimensional displacement threshold When an anomaly occurs, the current point is considered an outlier. By identifying anomalous 2D key points, the data quality input into subsequent 3D reconstruction processes is improved.
[0071] The steps for generating corresponding target 3D key points based on valid 2D key points include any of the following methods: Step S301: Generate target 3D key points based on other 2D key points besides the abnormal 2D key points.
[0072] As one implementation method, abnormal keypoints in the two-dimensional keypoints are removed, and only the other two-dimensional keypoints are used for three-dimensional reconstruction to generate the target three-dimensional keypoints. By removing the two-dimensional keypoints that are judged to be abnormal, the contamination of the single three-dimensional coordinate calculation results by momentary false detections or jitter can be avoided, thereby improving the accuracy of the reconstruction results.
[0073] Step S302: Determine the smoothed two-dimensional key points corresponding to the abnormal two-dimensional key points and the target two-dimensional key points, and generate the target three-dimensional key points based on the smoothed two-dimensional key points and other two-dimensional key points.
[0074] As one implementation method, abnormal two-dimensional key points are smoothed to obtain smoothed two-dimensional key points. An exponential smoothing method is used.
[0075]
[0076] in For smoothing coefficients, The larger the value, the higher the weight of the current frame, and the weaker the smoothing effect. The x-coordinate after smoothing the current frame. The x-coordinate is the smoothed x-coordinate of the previous frame. The ordinate is the smoothed ordinate of the current frame. The ordinate of the previous frame after smoothing is used to obtain the smoothed 2D keypoints. By smoothing abnormal 2D keypoints, local jitter caused by detection noise can be reduced, making the 3D reconstruction results more continuous.
[0077] Step S303: Reduce the construction weights corresponding to abnormal two-dimensional key points based on a preset strategy, and construct the target three-dimensional key points based on the updated construction weights.
[0078] As one implementation method, based on a preset strategy, such as anomaly detection based on temporal displacement, potential anomalous observations are identified, and their weights in the weighted optimization are reduced accordingly. For example, when calculating the weights, the original confidence level is multiplied by a penalty factor calculated based on the degree of anomalousness to obtain the reduced weights, and the target 3D keypoints are constructed based on the updated construction weights. By weakening the influence of anomalous observations, their interference is suppressed while preserving the geometric constraint information of all data, making the 3D keypoint solution results smoother and more stable.
[0079] Furthermore, the two-dimensional keypoint results from a single viewpoint are denoted as:
[0080] Let be the set of all 2D keypoints detected by the c-th camera at time t, which is a 2D array. The 2D keypoint results from each viewpoint are integrated to obtain the 2D keypoint results for all views, denoted as:
[0081] It represents the set of detection results of all C cameras and all J key points at time t, which is a three-dimensional array.
[0082] In this embodiment, by performing any of the following processing on abnormal two-dimensional key points—not participating in the three-dimensional reconstruction of this frame, replacing them with short-term smoothing results, or reducing their weight in subsequent triangulation—it is possible to effectively suppress the interference of detection noise and instantaneous false detections, and improve the temporal smoothness, stability, and overall robustness of the three-dimensional skeleton output.
[0083] Based on any of the above embodiments of this application, Embodiment 4 of this application proposes a three-dimensional skeleton reconstruction method, which can be referred to the above description and will not be repeated hereafter. Based on this, corresponding target three-dimensional key points are generated based on effective two-dimensional key points, including: Step S31: Randomly select at least two valid two-dimensional key points corresponding to observation viewpoints, and calculate the candidate three-dimensional key points.
[0084] As one implementation method, the RANSAC method is used to reconstruct each two-dimensional keypoint from its corresponding valid two-dimensional observation set. In this process, two different viewpoints are randomly selected, and the corresponding 2D keypoint coordinate data are obtained. Based on the pre-obtained intrinsic and extrinsic parameters and necessary distortion parameters of each industrial camera, candidate 3D keypoints are obtained by solving the 2D keypoints corresponding to the two different viewpoints according to the relationship between the projection matrix and the viewpoint projection. For the first... The projection matrix of the road industrial camera is denoted as .
[0085] in For the camera intrinsic parameter matrix, For rotation matrix, This is the translation vector. For candidate 3D keypoints... In its first The projection relationship in the road camera is satisfied.
[0086] in As a scale factor, These are the homogeneous coordinates of the two-dimensional observation point on the image. For the same key point... When valid 2D observations exist in two or more viewpoints, candidate 3D keypoints are obtained based on the projection relationships of these viewpoints. By randomly selecting the smallest subset of observations, multiple different initial 3D point assumptions are generated for subsequent triangulation, avoiding the algorithm results being dominated by false detections from individual viewpoints.
[0087] As one implementation method, three or more valid two-dimensional keypoints corresponding to observation perspectives are randomly selected to obtain candidate three-dimensional keypoints. When computational resources permit, multiple valid two-dimensional keypoints corresponding to observation perspectives can be randomly selected for calculation, which can improve the quality of the initial hypothesis. The number of randomly selected observation perspectives can be flexibly adjusted according to computing power and accuracy requirements, allowing for an optimized configuration between computational efficiency and result reliability.
[0088] Specifically, the geometric parameters of each industrial camera are calibrated using either the Charuco (Chessboard ArUco board) calibration method or the moving ball calibration method, yielding the intrinsic, extrinsic, and necessary distortion parameters for each camera. When the relative viewing angles between industrial cameras are small or they cannot simultaneously observe the same moving target, the Charuco calibration method is used. The Charuco calibration board is fabricated on a float glass substrate, and its size can be customized according to the specific experimental site and shooting range. By acquiring images of the calibration board at different spatial positions and orientations, the intrinsic, extrinsic, and spatial geometric relationships of each camera can be stably determined. When there are relative viewing angles between cameras and at least two cameras can simultaneously observe the same moving ball or the same rigid multi-sphere structure, the moving ball calibration method is used. The same rigid multi-sphere structure refers to a multi-sphere structure rigidly connected by multiple balls according to a fixed geometric relationship. By moving a ball with distinct characteristics within the experimental space, its multi-view trajectory at different times and positions is acquired, and the geometric parameters of the multiple cameras are estimated by combining the corresponding relationships. By flexibly selecting different calibration methods based on the structure of the experimental setup, the degree of overlap in the field of view, and the convenience of actual operation, we can obtain geometric parameters that are more in line with reality.
[0089] Preferably, if the geometric parameters include lens distortion parameters, then the two-dimensional points are first subjected to distortion correction processing before participating in the solution of candidate three-dimensional key points.
[0090] Step S32: Determine the reprojection error of the candidate 3D key points at the observation viewpoint.
[0091] In this embodiment, reprojection error refers to the Euclidean distance between the theoretical projection point coordinates obtained by reprojecting a candidate point onto the image plane of a certain camera viewpoint and the actual two-dimensional key point coordinates observed from that viewpoint. It is used to measure the degree of agreement between the candidate point and the observation data of a single viewpoint.
[0092] As one implementation method, for candidate points From the perspective The reprojection error on is,
[0093] in Representing candidate 3D key points The projection in the c-th camera, Let c be the coordinates of the 2D keypoints. For each effective viewpoint c, the candidate 3D keypoints are projected using the projection function of the camera's projection matrix. Projecting this onto the camera's image plane yields the coordinates of a theoretical projection point. Calculate the theoretical projection point Two-dimensional key points actually observed from this perspective The Euclidean distance between them is the reprojection error of the candidate point at that viewpoint. By quantifying the degree of agreement between candidate 3D key points and actual 2D observations, a crucial criterion for subsequent screening is provided, ensuring the robustness of 3D reconstruction.
[0094] Step S33: When the reprojection error is less than the error threshold, it is determined that the observation viewpoint is consistent with the candidate 3D keypoint.
[0095] As one implementation method, when In other words, if the reprojection error is less than a preset reprojection error, the viewpoint is considered to be consistent with the candidate point. By calculating the reprojection error of the candidate point under all viewpoints and comparing it with the threshold, reliable observation data that matches the candidate point are selected, thereby quantifying the reliability of the candidate 3D keypoint.
[0096] Step S34: Select the candidate 3D key point with the most consistent observation perspectives as the target 3D key point.
[0097] As one implementation method, after generating multiple candidate 3D points through multiple random samplings, the number of consistent observation views corresponding to each candidate point is traversed and counted, and the candidate 3D keypoint with the most consistent observation views is selected as the target 3D keypoint. This ensures that the selected target 3D keypoint receives the cross-validation of the most reliable observation data, effectively eliminating outliers caused by false detections of individual views or temporary occlusion.
[0098] In this embodiment, by generating hypotheses through random sampling, verifying consistency, and selecting the optimal solution, the coordinates of the target's 3D key points can be accurately calculated even when there are false detections, occlusions, or noise in some camera perspectives, thereby improving the robustness and reliability of 3D skeleton reconstruction in complex real-world scenes.
[0099] Based on any of the above embodiments of this application, Embodiment 5 of this application proposes a three-dimensional skeleton reconstruction method, which can be referred to the above description and will not be repeated hereafter. Based on this, after selecting the candidate three-dimensional keypoint with the largest number of consistent observation viewpoints as the target three-dimensional keypoint, the method includes: Step S35: Determine the weight of each observation viewpoint based on the confidence level of the effective two-dimensional keypoints.
[0100] As one implementation method, the weights of each observation perspective are defined based on the confidence level of the effective two-dimensional keypoints.
[0101] in The confidence level for detecting the effective 2D keypoint by the c-th camera. To prevent small constants with a denominator of zero, Let c be the weight of the observation viewpoint. The higher the confidence level of the viewpoint, the greater its weight in the triangulation. By calculating the circles of each observation viewpoint based on the confidence level, the reconstruction results become more dependent on high-quality observations.
[0102] Step S36: Perform weighted refinement of the target 3D key points based on weights to obtain the refined target 3D key points.
[0103] As one implementation method, a weighted multi-view reprojection error minimization solution is performed on the target's 3D key points. Based on these target 3D key points, the refined target 3D key points are then solved to minimize their reprojection errors across multiple viewpoints.
[0104] in Representing a three-dimensional point In the Projection in industrial cameras, This indicates the weight of the viewpoint. The Euclidean distance between the theoretical projection point and the actual observation point is given. By adjusting the position of the 3D point, the overall error between its theoretical projection position and the actual observation position in each effective camera is minimized, resulting in the refined 3D keypoints of the target. By refining the positions of the candidate 3D keypoints to the geometrically optimal solution, higher reconstruction accuracy is achieved while ensuring robustness.
[0105] In this embodiment, by using the selected candidate 3D key points as the initial solution and minimizing the reprojection error from multiple perspectives, the position of the 3D key points is optimized, thereby improving the reconstruction accuracy and noise resistance.
[0106] Based on any of the above embodiments of this application, Embodiment Six of this application proposes a three-dimensional skeleton reconstruction method, which can be referred to the above description and will not be repeated hereafter. On this basis, a real-time three-dimensional skeleton is generated by connecting target three-dimensional key points according to a preset skeleton topology relationship, including: Step S41: Calculate the smoothed target 3D key points based on the target 3D key points and the target 3D key points in the previous frame.
[0107] As one implementation method, exponential smoothing is applied to the target's three-dimensional key points.
[0108] in For smoothing coefficients, Let j be the 3D coordinates of the target's 3D key point in the current frame. The 3D coordinates of the target's 3D keypoint j, after smoothing in the previous frame. Smoothing suppresses the jitter of the 3D coordinates over time, making the output skeleton motion smoother and more natural.
[0109] Step S42: Calculate the three-dimensional displacement between the target three-dimensional key point and the target three-dimensional key point in the previous frame. When the three-dimensional displacement is less than or equal to the three-dimensional displacement threshold, determine the target three-dimensional key point as the corrected target three-dimensional key point.
[0110] As one implementation method, the 3D displacement between the target 3D keypoints in the current frame and the previous frame is calculated and compared with a preset 3D displacement threshold. If the 3D displacement between the current frame and the previous frame satisfies the threshold,
[0111] This point is then considered an abnormal jump point. To preset a 3D displacement threshold, abnormal jump points are ignored. When the 3D displacement is less than or equal to the 3D displacement threshold, the target 3D keypoint is determined as the corrected target 3D keypoint. By correcting the 3D keypoint, the impact of peripheral point jitter and local false detections on real-time skeleton stability can be reduced.
[0112] In this embodiment, by detecting and suppressing abnormal temporal jumps in the target 3D key points, temporal smoothing of the 3D skeleton is achieved, thereby improving the visual fluency, motion continuity and overall stability of the skeleton output while maintaining high real-time performance.
[0113] Based on any of the above embodiments of this application, Embodiment Seven of this application proposes a three-dimensional skeleton reconstruction method, which can be referred to the above description and will not be repeated hereafter. Based on this, after generating a real-time three-dimensional skeleton by connecting the target three-dimensional key points according to the preset skeleton topology, the method further includes: Step S401: Based on the real-time 3D skeleton, extract the posture and motion features to form a feature vector.
[0114] In this embodiment, posture and motion features refer to a series of numerical indicators used to quantitatively describe the animal's body shape and dynamic changes, such as the body's main axis direction, head direction, angles between key bone segments, distances between key points, velocity, acceleration, head height, and trunk curvature. The feature vector refers to a numerical array formed by sequentially arranging and combining all the calculated posture and motion features mentioned above.
[0115] As one implementation method, based on the obtained real-time 3D skeleton, pose and motion features are extracted to form a feature vector.
[0116] in For attitude and motion features, by transforming the spatiotemporal geometric information of the 3D skeleton, i.e. attitude and motion features, into structured numerical feature vectors, a computable and discriminable high-dimensional data representation is provided for subsequent real-time analysis.
[0117] Step S402: Perform state discrimination or decoding on the feature vector to obtain the current behavior state or motion intention.
[0118] In this embodiment, the current behavioral state refers to the clearly categorized behavior that the animal is currently performing. Examples include raising its head, turning around, running, and remaining still. Motion intention refers to the continuously predictable or inferred motion parameters or goals that the animal is about to perform. Examples include the angular velocity of an imminent left turn and the speed at which it will move forward.
[0119] As one implementation method, the obtained feature vector is input into an online state discriminator or decoder to obtain the current behavioral state or motion intention output.
[0120] Where F represents the online state discrimination model, online decoding model, or rule-based judgment function. By classifying or regressing the feature vectors, continuous posture changes of animals are identified online as discrete behavioral categories or decoded into continuous motion parameters, enabling real-time interpretation of the intentions associated with animal behavior or neural activity.
[0121] Step S403: Generate a control signal based on the current behavior state or motion intention, and output the control signal to an external device for control.
[0122] In this embodiment, the control signal refers to the signal used to directly drive or trigger external devices to perform corresponding operations. External devices refer to experimental instruments or devices connected through hardware I / O interfaces, such as stimulation devices, like optogenetic stimulators, electrical stimulators, etc.; and drug delivery devices, such as microinfusion pumps, olfactory stimulators, etc.
[0123] As one implementation method, a control signal is generated based on the current behavioral state, the state of the motion intention, or the decoding result.
[0124] in To control the mapping function, This refers to the current behavioral state or movement intention. This represents the feature vector. Control signals can be output to external devices via hardware I / O interfaces, such as stimulation devices, drug delivery devices, neural signal acquisition systems, or other execution devices. By connecting to external devices, online stimulation, real-time drug delivery, real-time behavioral state triggering, or real-time decoding control based on the animal's real-time three-dimensional skeleton can be achieved.
[0125] In this embodiment, by converting the real-time three-dimensional skeleton of an animal into a computable feature vector and performing online analysis, the output control signal can precisely regulate the animal's behavior or neural activity. At the same time, the subsequent behavioral changes of the animal are perceived and analyzed in real time, thereby forming a new round of control decisions, realizing closed-loop regulation and interactive experimental control of animal behavior or neural activity.
[0126] For example, to help understand the technical concept or principle of the three-dimensional skeleton reconstruction method of this application, please refer to... Figure 3 , Figure 3 The overall flowchart of the three-dimensional skeleton reconstruction method provided in the embodiments of this application is shown below: Please also refer to Figure 4 , Figure 4This is a simplified flowchart illustrating the three-dimensional skeleton reconstruction method provided in this application embodiment. First, animal behavior and posture data are acquired in parallel using multi-view industrial cameras. Two-dimensional keypoints are extracted from multiple animal images using the same lightweight detection model. A validity indicator for each two-dimensional keypoint is determined based on confidence level. When a two-dimensional keypoint is determined to have valid two-dimensional observations in at least two viewing angles based on the validity indicator, a three-dimensional keypoint is obtained through 2D-3D reconstruction triangulation and RANSAC algorithm processing. Finally, a high-precision real-time state discrimination and motion decoding result is output through the constructed real-time three-dimensional skeleton. Simultaneously, it supports the input of external neural signals or task events as auxiliary judgment criteria. Based on the real-time decoded state information, the closed-loop decision-maker dynamically sets the intervention target, threshold, and time parameters, thereby driving the I / O control command output to achieve precise control of external intervention devices, such as optogenetic stimulators and drug delivery pumps. Relevant data is saved in real-time to an image / 2D / 3D event library. The saved data includes multiple original image or video streams, timestamps for each stream, 2D keypoints and confidence levels from each viewpoint, 3D keypoint coordinates, 3D skeleton feature behavior, online state discrimination results, control output events and their times, and timestamps and stimulation signals from other access devices acquired synchronously. Based on this data saving mechanism, the saved data can not only be used for online closed-loop processing, but also, after the experiment, be input into a larger-scale, more complex offline model for further analysis. This gives the application a dual-layer working capability combining real-time and offline modes. The real-time mode is responsible for stable and rapid output of 3D pose, while the offline mode can utilize the saved data for more complex and higher-precision subsequent calculations.
[0127] This application provides a three-dimensional skeleton reconstruction device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the three-dimensional skeleton reconstruction method in the above embodiment 1.
[0128] The following is for reference. Figure 5 The diagram illustrates a structural schematic suitable for implementing the three-dimensional skeleton reconstruction device of the embodiments of this application. The three-dimensional skeleton reconstruction device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), tablet computers, and in-vehicle terminals, as well as fixed terminals such as digital TVs and desktop computers. Figure 5 The illustrated 3D skeleton reconstruction device is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0129] like Figure 5As shown, the 3D skeleton reconstruction device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the 3D skeleton reconstruction device. The processing unit 1001, the read-only memory 1002, and the RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the 3D skeleton reconstruction device to communicate wirelessly or wiredly with other devices to exchange data. Although a 3D skeleton reconstruction device with various systems is shown in the figure, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.
[0130] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0131] The three-dimensional skeleton reconstruction device provided in this application, employing the three-dimensional skeleton reconstruction method described in the above embodiments, can solve the technical problem of poor adaptability to experimental sites. Compared with the prior art, the beneficial effects of the three-dimensional skeleton reconstruction device provided in this application are the same as those of the three-dimensional skeleton reconstruction device provided in the above embodiments, and other technical features of this three-dimensional skeleton reconstruction device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.
[0132] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0133] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0134] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the three-dimensional skeleton reconstruction method in the above embodiments.
[0135] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, radio frequency (RF), or any suitable combination thereof.
[0136] The aforementioned computer-readable storage medium may be included in the three-dimensional skeleton reconstruction device; or it may exist independently and not be assembled into the three-dimensional skeleton reconstruction device.
[0137] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by the 3D skeleton reconstruction device, the 3D skeleton reconstruction device: inputs multiple animal images acquired from multiple observation perspectives into a lightweight detection model in parallel, and obtains in real time the two-dimensional keypoints and their confidence levels for each observation perspective; determines the validity indicator of the two-dimensional keypoints based on the confidence levels; when the validity indicator determines that there are valid two-dimensional keypoints in at least two observation perspectives, it generates the corresponding target 3D keypoints based on the valid two-dimensional keypoints; and connects the target 3D keypoints according to a preset skeleton topology to generate a real-time 3D skeleton.
[0138] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0139] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0140] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0141] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described three-dimensional skeleton reconstruction method, thereby solving the technical problem of poor adaptability to experimental sites. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the three-dimensional skeleton reconstruction method provided in the above embodiments, and will not be repeated here.
[0142] This application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the three-dimensional skeleton reconstruction method described above.
[0143] The computer program product provided in this application can solve the technical problem of poor adaptability to experimental sites. Compared with the prior art, the beneficial effects of the computer program product provided in the embodiments of this application are the same as the beneficial effects of the three-dimensional skeleton reconstruction method provided in the above embodiments, and will not be repeated here.
[0144] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent scope of this application.
Claims
1. A method for reconstructing a three-dimensional skeleton, characterized in that, The three-dimensional skeleton reconstruction method includes: Multiple animal images acquired from multiple observation perspectives are input into a lightweight detection model in parallel to obtain the two-dimensional key points corresponding to each observation perspective and the confidence level of the two-dimensional key points in real time. Determine the confidence threshold corresponding to each of the two-dimensional key points, wherein the confidence threshold is different for different parts; If the confidence level of the two-dimensional key point is greater than or equal to the confidence threshold, the two-dimensional key point is associated with a valid indicator. Otherwise, the two-dimensional keypoints are associated with invalid indicators, wherein the valid indicators include both valid and invalid indicators, and the valid indicators are identifier variables used to mark whether the observation data of each two-dimensional keypoint is valid in each observation view. Determine the two-dimensional displacement between the two-dimensional key point and the corresponding target two-dimensional key point in the previous frame; Identify abnormal two-dimensional key points whose two-dimensional displacement is greater than or equal to a preset two-dimensional displacement threshold. For each of the two-dimensional key points, all the observation views are traversed, and the validity indicator of the two-dimensional key points under the observation view is checked one by one. All view numbers with a validity indicator of 1 are collected to obtain the valid two-dimensional observation set of the actual three-dimensional key points. When it is determined, based on the validity indicator in the effective two-dimensional observation set, that the actual three-dimensional keypoint has valid two-dimensional keypoints in at least two observation perspectives, a corresponding target three-dimensional keypoint is generated based on the valid two-dimensional keypoints. The step of generating the corresponding target three-dimensional keypoint based on the valid two-dimensional keypoints includes any of the following methods: generating the target three-dimensional keypoint based on other two-dimensional keypoints besides the abnormal two-dimensional keypoints; determining the smoothed two-dimensional keypoints corresponding to the abnormal two-dimensional keypoints and the target two-dimensional keypoints, and generating the target three-dimensional keypoint based on the smoothed two-dimensional keypoints and the other two-dimensional keypoints; reducing the construction weight corresponding to the abnormal two-dimensional keypoints based on a preset strategy, and constructing the target three-dimensional keypoint based on the updated construction weights. The target 3D key points are connected according to the preset skeleton topology to generate a real-time 3D skeleton.
2. The three-dimensional skeleton reconstruction method as described in claim 1, characterized in that, The process involves inputting multiple animal images acquired from various observation perspectives into a lightweight detection model in parallel, and obtaining in real time the two-dimensional keypoints corresponding to each observation perspective and the confidence level of the two-dimensional keypoints, including: Each observation perspective is assigned an independent processing thread, and a model instance of the same lightweight detection model is loaded in each processing thread. Each processing thread reads the animal image from the corresponding observation viewpoint, performs model inference based on the model instance, and outputs the two-dimensional keypoints and the confidence level of the two-dimensional keypoints corresponding to each observation viewpoint in parallel.
3. The three-dimensional skeleton reconstruction method as described in claim 1, characterized in that, The process of generating corresponding target 3D key points based on the effective 2D key points includes: Randomly select at least two effective two-dimensional key points corresponding to the observation viewpoints, and calculate candidate three-dimensional key points; Determine the reprojection error of the candidate 3D key points on the observation viewpoint; When the reprojection error is less than the error threshold, it is determined that the observation viewpoint is consistent with the candidate 3D keypoint. The candidate 3D key point with the most consistent observation perspectives is selected as the target 3D key point.
4. The three-dimensional skeleton reconstruction method as described in claim 3, characterized in that, After selecting the candidate 3D keypoint with the most consistent observation viewpoints as the target 3D keypoint, the process further includes: The weights of each observation viewpoint are determined based on the confidence level of the effective two-dimensional keypoints. The target 3D key points are refined based on the weights to obtain the refined target 3D key points.
5. The three-dimensional skeleton reconstruction method as described in claim 1, characterized in that, The step of connecting the target 3D key points according to the preset skeleton topology to generate a real-time 3D skeleton includes: Based on the target 3D key points and the target 3D key points in the previous frame, calculate the smoothed target 3D key points; Calculate the three-dimensional displacement between the target three-dimensional key point and the target three-dimensional key point in the previous frame. When the three-dimensional displacement is less than or equal to the three-dimensional displacement threshold, determine the target three-dimensional key point as the corrected target three-dimensional key point.
6. The three-dimensional skeleton reconstruction method as described in claim 1, characterized in that, After connecting the target 3D key points according to the preset skeleton topology to generate a real-time 3D skeleton, the process further includes: Based on the real-time 3D skeleton, posture and motion features are extracted to form feature vectors; The feature vector is subjected to state discrimination or decoding to obtain the current behavior state or motion intention; Based on the current behavior state or the motion intention, a control signal is generated, and the control signal is output to an external device for control.
7. A three-dimensional skeleton reconstruction device, characterized in that, The three-dimensional skeleton reconstruction device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the three-dimensional skeleton reconstruction method as described in any one of claims 1 to 6.
8. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the three-dimensional skeleton reconstruction method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
High-precision multi-view motion capture method and system
CN118506458A
Method and system of multi-view image processing with accurate skeleton reconstruction
WO2023087164A1