Long-short soldier training track analysis method and system based on three-dimensional reconstruction

By decomposing the training scene into a static background and a dynamic foreground, and optimizing weapon motion using neural radiation fields and B-spline functions, the problems of high equipment cost and insufficient accuracy in existing technologies are solved, and high-precision analysis of weapon trajectories is achieved.

CN121169971BActive Publication Date: 2026-03-27GUANGDONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing optical motion capture systems and traditional computer vision technologies suffer from high equipment costs, cumbersome and easily detached markers, difficulty in handling occlusion and rapid movement, and inability to accurately recover the weapon's three-dimensional attitude and high-order motion parameters in weapon trajectory analysis.

Method used

The training scene is decomposed into a static background and a dynamic foreground, and the neural radiation field is used to represent the two respectively. By combining a parameterized six-DOF B-spline function and a deformable human body model, the rigid body motion of the weapon and the non-rigid deformation of the human body are optimized through volume rendering technology, and the three-dimensional motion trajectory and high-order motion parameters of the weapon are extracted.

Benefits of technology

It can accurately reconstruct the three-dimensional motion trajectory and attitude of weapons without the need for marker points, and directly calculate linear velocity, angular velocity and acceleration, providing refined data support for training and evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121169971B_ABST
    Figure CN121169971B_ABST
Patent Text Reader

Abstract

The application provides a long-short soldier training trajectory analysis method and system based on three-dimensional reconstruction, which decouples the training scene into a static background and a dynamic foreground containing a trainer and a weapon; a parameterized six-degree-of-freedom B-spline function is introduced for accurate modeling of the rigid body motion of the weapon; a deformable human body model is used to represent the non-rigid deformation of the human body; the spatial points in the dynamic foreground at time t are transformed from the observation space to a unified canonical space, and the canonical space coordinates are input into a second neural radiation field; using volume rendering technology, the camera light rays are sampled, the color and density are queried in combination with the two neural radiation fields, and the rendering image is integrated to generate; after joint optimization and convergence, the optimal weapon B-spline function is extracted, the zero-order, first-order and second-order derivatives are calculated, and the three-dimensional motion trajectory, attitude, linear velocity, angular velocity and acceleration of the weapon in the entire training process are output.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of trajectory analysis, and particularly relates to a long and short weapon training trajectory analysis method and system based on three-dimensional reconstruction. BACKGROUND

[0002] In the fields of martial arts, sports competition and military training, by obtaining key motion parameters such as the attitude, speed and acceleration of a weapon in three-dimensional space, objective and quantitative data support can be provided for the motion optimization, technical evaluation and tactics formulation of athletes or soldiers. An optical motion capture system reconstructs a high-precision three-dimensional motion trajectory by pasting reflective marker points on the weapon and the body of the trainer, and tracking by using multiple high-speed cameras. However, such a marker point-based method has significant defects: the equipment cost is high, it needs to be deployed in a professional site, and the pasting process of the marker points is cumbersome, which is prone to displacement or falling off during intense exercise, and causes interference to the training process itself. Traditional computer vision technology performs markerless motion analysis by using a single view or multiple views of ordinary video, but these methods are difficult to meet the high-precision requirements of weapon trajectory analysis in dealing with complex occlusion problems, fast motion and accurate recovery of three-dimensional attitude. NeRF implicitly learns the geometry and appearance information of a scene by using a multi-layer perceptron, and can synthesize highly realistic new view images from sparse multi-view images. In order to process dynamic scenes, traditional NeRF is extended, for example, time is introduced as input, or a deformation field from the observation space to the canonical space is constructed to model the motion of the object. These dynamic NeRF methods have achieved remarkable results in the reconstruction and new view synthesis of dynamic scenes such as people and vehicles, and usually decompose the scene into a static background and a dynamic foreground, and uniformly model the deformation of the dynamic foreground. The existing dynamic NeRF methods mainly aim to achieve high-quality visual rendering, and usually model the dynamic foreground such as the trainer and the weapon in his hand as a whole for fuzzy uniform non-rigid deformation modeling, and cannot effectively decouple and explicitly express the rigid motion of the weapon and the non-rigid motion of the human body, so it is difficult to directly and accurately extract the continuous six-degree-of-freedom motion trajectory, angular velocity and other high-order motion parameters of the weapon as an independent rigid body, which limits its application in the field of professional motion analysis. SUMMARY

[0003] In view of the above problems, in a first aspect of the application, a long and short weapon training trajectory analysis method based on three-dimensional reconstruction is provided, comprising the following steps:

[0004] obtaining a multi-view synchronous video sequence and camera internal and external parameters for a weapon training process; decomposing the training scene into a static background field and a dynamic foreground field, the dynamic foreground field containing a trainer and a weapon, respectively using a first neural radiance field to represent the static background field, and using a second neural radiance field to represent the dynamic foreground field;

[0005] a parameterized six-degree-of-freedom B-spline function for describing the rigid body motion of the weapon, and a deformable human model for characterizing the non-rigid deformation of the human body; for any spatial point in the dynamic foreground field at time t, a mapping parameterized by the B-spline function and the deformable human model is used to transform the spatial point from the observation space to a unified canonical space, and coordinates of the spatial point in the canonical space are input as features into the second neural radiance field;

[0006] a rendering image is obtained by integrating along the light ray; a photometric rendering loss between the rendering image and the real observation image is constructed, and the first neural radiance field and the second neural radiance field, the six-degree-of-freedom B-spline function of the weapon, and the deformable human model are jointly optimized;

[0007] After the joint optimization converges, the optimal six-degree-of-freedom B-spline function of the weapon is extracted, and zero-order, first-order and second-order derivatives are calculated to obtain three-dimensional motion trajectories, postures, linear velocities, angular velocities and accelerations of the weapon during the entire training process.

[0008] In a second aspect of the present application, a long-short weapon training trajectory analysis system based on three-dimensional reconstruction is provided, comprising the following modules:

[0009] a scene decomposition module for obtaining a multi-view synchronous video sequence and camera internal and external parameters for a weapon training process; the training scene is decomposed into a static background field and a dynamic foreground field, the dynamic foreground field containing a trainer and a weapon, the first neural radiance field being used to represent the static background field and the second neural radiance field being used to represent the dynamic foreground field;

[0010] a mapping module for constructing a parameterized six-degree-of-freedom B-spline function for describing the rigid body motion of the weapon, and a deformable human model for characterizing the non-rigid deformation of the human body; for any spatial point in the dynamic foreground field at time t, a mapping parameterized by the B-spline function and the deformable human model is used to transform the spatial point from the observation space to a unified canonical space, and coordinates of the spatial point in the canonical space are input as features into the second neural radiance field;

[0011] an optimization module for sampling along the camera light ray by volume rendering technology, for each sampling point, querying and synthesizing color and density in combination with the first neural radiance field and the second neural radiance field, and integrating along the light ray to obtain a rendering image; a photometric rendering loss between the rendering image and the real observation image is constructed, and the first neural radiance field and the second neural radiance field, the six-degree-of-freedom B-spline function of the weapon, and the deformable human model are jointly optimized;

[0012] An analysis module is configured to extract optimal weapon six-degree-of-freedom B-spline functions after convergence of the joint optimization, and to calculate zero-order, first-order and second-order derivatives to obtain three-dimensional motion trajectories, postures, linear velocities, angular velocities and accelerations of the weapon during the entire training process.

[0013] In an optional embodiment, the acquisition of the multi-view synchronous video sequence for the weapon training process and the camera internal and external parameters comprises:

[0014] Eight to twelve high-speed industrial cameras with a frame rate of not less than 100 fps are arranged in a ring around the training site, and a checkerboard calibration plate and a joint beam adjustment algorithm are used to solve the internal and external parameters of all the cameras.

[0015] A hardware synchronization controller is used to ensure that the starting time stamp error of exposure of all the cameras is less than 1 millisecond.

[0016] In an optional embodiment, the decomposition of the training scene into a static background field and a dynamic foreground field comprises:

[0017] An instance segmentation network is used to extract a pixel-level mask of the trainer and the weapon in each frame of image.

[0018] In the joint optimization process, a photometric loss for optimizing the first neural radiance field is calculated based on a pixel region outside the mask, and a photometric loss for optimizing the second neural radiance field is calculated based on a pixel region covered by the mask.

[0019] In an optional embodiment, the transformation from the observation space to the unified canonical space by using the mapping parameterized by the B-spline function and the deformable human body model comprises:

[0020] For any sampling point in the observation space at time t, the sampling point is transformed to the unified canonical space, wherein an inverse transformation of the deformable human body model is applied to a point belonging to the trainer, and an inverse matrix of a transformation corresponding to the six-degree-of-freedom B-spline function is applied to a point belonging to the weapon.

[0021] In the unified canonical space, the trainer and the weapon jointly constitute a static canonical scene.

[0022] In an optional embodiment, the joint optimization comprises:

[0023] A total loss function is constructed wherein is a photometric error of a rendered image and a real image is a binary cross-entropy loss between a cumulative opacity of a rendered light and a foreground mask, is a parameter regularization term of the deformable human body model, A motion smoothness regularization term is imposed on the B-spline function to penalize excessive velocity or acceleration.

[0024] In an optional embodiment, the parameterized six-degree-of-freedom B-spline function is constructed to describe the rigid body motion of the weapon, comprising:

[0025] A cubic B-spline function with control points as a series of SE(3) group elements is constructed to fit the continuous motion trajectory of the weapon in the time dimension by optimizing the control points, wherein the SE(3) group elements encode both the three-dimensional rotation matrix and the translation vector of the weapon.

[0026] In an optional embodiment, the three-dimensional motion trajectory, pose, linear velocity, angular velocity and acceleration of the weapon during the entire training process are obtained, comprising:

[0027] From the optimized B-spline function, the SE(3) matrix at any time t is obtained by interpolation as the three-dimensional pose and position of the weapon at that time;

[0028] The first and second time derivatives of the B-spline function are calculated and mapped from the Lie algebra SE(3) space to the Euclidean space to analytically obtain the linear velocity, angular velocity, linear acceleration and angular acceleration of the weapon.

[0029] The present application does not require any marker points to be attached to the weapon or the trainer, and can be analyzed using only ordinary multi-view video; the rigid body motion of the weapon and the non-rigid deformation of the human body are decoupled, the weapon motion is explicitly and continuously mathematically described by a parameterized six-degree-of-freedom B-spline function, and the deformable human body model is jointly optimized in the neural radiance field framework, which not only accurately reconstructs the three-dimensional motion trajectory and pose of the weapon during the entire training process, but also directly calculates the high-order kinematic parameters such as linear velocity, angular velocity and acceleration of the weapon by taking the derivative of the B-spline function, providing fine and quantitative data support for training evaluation that traditional markerless methods cannot achieve. BRIEF DESCRIPTION OF DRAWINGS

[0030] Figure 1 Flowchart for specific embodiments;

[0031] Figure 2 High-speed camera distribution diagram;

[0032] Figure 3 Diagram for separating static background and dynamic foreground;

[0033] Figure 4 Diagram for mapping to a unified canonical space;

[0034] Figure 5 Diagram for joint optimization. DETAILED DESCRIPTION

[0035] For the purpose of promoting the technical solutions of the present application, the present application will be further described below in conjunction with the drawings.

[0036] The terms "first" and "second" and the like in the description, claims, and drawings of the present application merely mean to distinguish different objects, and are not intended to describe a particular order. In addition, the terms "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device, etc. comprising a series of steps or units is not limited to the listed steps or units, but can optionally further comprise steps or units not listed, etc., or can optionally further comprise other steps or units inherent to the process, method, product, or device, etc.

[0037] In this document, "embodiment" means that the specific features, structures, or characteristics described in connection with the embodiment can be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily mean the same embodiment is referred to, nor does it mean that the embodiments are mutually exclusive or alternative to each other. Those skilled in the art can explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0038] In this application, "at least one" means one or more, "multiple" means two or more, "at least two" means two or three and more, and "and / or" is used to describe the association relationship of the associated objects, which means that there can be three relationships, for example, "A and / or B" can mean: only A, only B, and A and B exist at the same time, where A and B can be singular or plural. "Or" means that there can be two relationships, such as only A, only B; when A and B are not mutually exclusive, it can also mean that there are three relationships, such as only A, only B, and A and B exist at the same time. The character " / " generally represents an "or" relationship between the associated objects before and after it. "At least one of the following" or similar expressions means any combination of these items. For example, at least one of a, b, or c, can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c".

[0039] In one embodiment, the present application proposes a long-short soldier training trajectory analysis method based on three-dimensional reconstruction, as shown in Figure 1 The method comprises the following steps:

[0040] S1, acquire multi-view synchronous video sequences and camera intrinsic and extrinsic parameters for weapon training process; decompose the training scene into static background field and dynamic foreground field, the dynamic foreground field contains the trainer and the weapon, represent the static background field by first neural radiance field, represent the dynamic foreground field by second neural radiance field;

[0041] Deploy multiple, for example, 4 to 8 high-speed cameras around the training site, as shown in Figure 2 As shown, ensure that all video frames are strictly aligned in time by means of hardware synchronization or post-vision alignment; use open-source Structure-from-Motion toolkits such as COLMAP to process the initial few frames of static images, and automatically calibrate the position, orientation, and other extrinsic parameters of each camera, as well as the focal length, distortion, and other intrinsic parameters.

[0042] Use instance segmentation networks such as Mask R-CNN to process each frame of image, generate foreground masks of the trainer and the weapon, and thus divide the image pixels into static background and dynamic foreground; target detection such as YOLO series model can also be used to decompose the training scene, and the decomposition result is as shown in Figure 3 The first neural radiance field is a multi-layer perceptron model that inputs three-dimensional space coordinates and outputs the color and volume density of the point; the second neural radiance field is also a multi-layer perceptron model that inputs three-dimensional coordinates in the canonical space and outputs the color and volume density of the point in the canonical state.

[0043] S2, construct a parameterized six-degree-of-freedom B-spline function for describing the rigid body motion of the weapon, and a deformable human body model for representing the non-rigid deformation of the human body; for any spatial point in the dynamic foreground field at time t, use the mapping parameterized by the B-spline function and the deformable human body model to transform from the observation space to the unified canonical space, and input the coordinates of the spatial point in the canonical space as features into the second neural radiance field;

[0044] The six-degree-of-freedom motion of the weapon, i.e., three-dimensional translation and three-dimensional rotation, is modeled by a time-varying quartic B-spline function, and the control points of the function are the parameters to be optimized; the non-rigid deformation of the human body is represented by the commonly used SMPL model in the industry, and its pose and body shape parameters are also used as optimization variables; for any spatial point in the foreground region at time t, determine whether it belongs to the weapon or the human body, for example, by target recognition, target detection, etc., if it belongs to the weapon, apply the inverse transformation of the rigid body transformation defined by the B-spline function at time t to map it to the canonical model space of the weapon; if it belongs to the human body, apply the inverse linear blend skinning transformation of the non-rigid deformation defined by the SMPL model at time t to map it to the standard A-pose canonical space of the human body, thereby obtaining the coordinates of the point in the unified canonical space, as shown in Figure 4 .

[0045] S3, sampling along camera rays by volume rendering technique, for each sampling point, querying and synthesizing color and density from the first and second neural radiance fields, integrating along the ray to get the rendered image; constructing the photometric rendering loss between the rendered image and the real observed image, and jointly optimizing the first and second neural radiance fields, the weapon six-degree-of-freedom B-spline function, and the deformable human body model;

[0046] For any pixel in the image, hierarchical random sampling is performed along its corresponding camera ray to obtain a series of three-dimensional sampling points; for each sampling point, if it is located in the background region, the color and density are directly queried from the first neural radiance field; if it is located in the foreground region, the aforementioned dynamic mapping transformation to the canonical space is first performed, the color and density are queried from the second neural radiance field; then, the standard volume rendering integral formula is used to accumulate the color and density of all sampling points along the ray to calculate the final rendering color of the pixel; the L2 norm, i.e., the mean square error, between the rendering color of all pixels and the real video image color is calculated as the photometric loss, and the Adam optimizer is used to simultaneously update the weights of the two neural radiance field networks, the control points of the B-spline function, and all parameters of the SMPL model by the back propagation algorithm, as shown in Figure 5 .

[0047] S4, after the joint optimization converges, the optimal weapon six-degree-of-freedom B-spline function is extracted, and the zero-order, first-order, and second-order derivatives are calculated to obtain the three-dimensional motion trajectory, attitude, linear velocity, angular velocity, and acceleration of the weapon during the entire training process.

[0048] Optimization convergence means that the control points of the B-spline function have reached the optimal solution; the optimal B-spline function is densely sampled on the time axis to evaluate the zero-order derivative value, i.e., the three-dimensional spatial position and rotational attitude of the weapon at each time, constituting the motion trajectory; using the analytical derivability of the B-spline function, the first-order derivative is calculated to obtain the linear velocity vector of the weapon's center of mass and the angular velocity vector of the attitude change; similarly, the second-order derivative is calculated to obtain the linear acceleration of the weapon's center of mass and the angular acceleration of the attitude change.

[0049] In an optional embodiment, the acquisition of the multi-view synchronous video sequence and the camera internal and external parameters for the weapon training process comprises:

[0050] Eight to twelve high-speed industrial cameras with a frame rate of not less than 100 fps are deployed in a ring around the training site, and a checkerboard calibration plate and a joint beam adjustment algorithm are used to solve the internal and external parameters of all cameras;

[0051] A hardware synchronization controller is used to ensure that the starting timestamp error of all camera exposures is less than 1 millisecond.

[0052] To capture the details of a trainer's high-speed weapon swing, such as a sword slash or a spear thrust, the embodiment evenly deploys 10 high-speed industrial cameras along the circumference of a 5-meter-radius circular field. The frame rate of these cameras is set to 200 fps, much higher than the 30 fps of regular video, which can clearly record the instantaneous state of each action and avoid motion blur. Such a ring deployment ensures that no matter which direction the trainer faces, his actions can be unobstructed by at least three cameras, providing sufficient visual data for subsequent three-dimensional reconstruction.

[0053] Before data collection, the spatial positions and optical characteristics of all cameras need to be accurately calibrated. The staff moves a 9x6 checkerboard calibration plate around the field, allowing all cameras to take pictures of the calibration plate from different angles. Through these images, the intrinsic parameters such as focal length and distortion coefficient of each camera, and the extrinsic parameters, i.e., the position and orientation of the camera relative to the world coordinate system, are initially calculated. Subsequently, a joint bundle adjustment algorithm is used to globally optimize the intrinsic and extrinsic parameters of all cameras, reducing the average three-dimensional point reprojection error in all camera views to less than 0.5 pixels, establishing a high-precision unified three-dimensional space. To achieve accurate time alignment of multi-angle pictures, a hardware synchronization controller sends a trigger signal to all cameras, ensuring that the start time difference of each camera's shutter exposure is less than 0.5 milliseconds. Sub-millisecond synchronization accuracy is crucial for analyzing actions with weapon tip speeds exceeding 20 meters per second, as it can avoid tearing or distortion of the three-dimensional reconstruction model caused by time asynchronization.

[0054] In an optional embodiment, the decomposition of the training scene into a static background field and a dynamic foreground field includes:

[0055] Using an instance segmentation network to extract a pixel-level mask of the trainer and the weapon in each frame of image;

[0056] In the joint optimization process, the photometric loss for optimizing the first neural radiance field is calculated based on the pixel region outside the mask, and the photometric loss for optimizing the second neural radiance field is calculated based on the pixel region covered by the mask.

[0057] The complex training scene is decoupled into two parts for modeling to improve reconstruction efficiency and quality. The first part is the static background field, i.e., the training field itself, such as the ground, walls, and fixed training equipment. The second part is the dynamic foreground field, which includes the trainer and the weapon he uses. To achieve this separation, a pre-trained instance segmentation network, such as YOLACT or Mask R-CNN, is used to process each frame of input image. The instance segmentation network can output a mask accurate to the pixel level, clearly defining which pixels belong to the trainer, which belong to the weapon, and which belong to the background.

[0058] In the subsequent neural radiance field training phase, when optimizing the first neural radiance field representing the static background, its photometric loss function only calculates the pixel area outside the mask, that is, only the rendered background color is compared with the background color in the real image. Conversely, when optimizing the second neural radiance field representing the dynamic foreground, the photometric loss function only focuses on the pixel area covered by the mask. In this way, the geometry and texture information of the static background will not be disturbed by the moving person and weapon, and the modeling of the dynamic foreground also does not need to consider the complex background changes, so that both neural radiance fields can learn a more pure and accurate scene representation, and the finally synthesized video is also more clear and real.

[0059] In an optional embodiment, the transforming from the observation space to the unified canonical space by using the mapping parameterized by the B-spline function and the deformable human body model comprises:

[0060] For any sampling point in the observation space at time t, the sampling point is transformed to the unified canonical space, wherein an inverse transformation of the deformable human body model is applied to the point belonging to the trainer, and an inverse matrix of the transformation corresponding to the six-degree-of-freedom B-spline function is applied to the point belonging to the weapon;

[0061] In the unified canonical space, the trainer and the weapon jointly constitute a static canonical scene.

[0062] The canonical space is a static three-dimensional reference system, in which the trainer always maintains a standard static posture, such as T-pose, and the weapon is also in a fixed initial position and posture. The complex time-varying four-dimensional problem, i.e. three-dimensional space plus time, is decomposed into a simple static three-dimensional geometry learning problem and a dynamic motion learning problem. At any time t, when the color and density of a three-dimensional point P in the observation space need to be queried, the point P is first mapped back to the canonical space. If it is judged according to the segmentation information that the point P belongs to the trainer, the inverse transformation corresponding to the posture parameter at time t of the deformable human body model is applied. For example, if the right arm is lifted by 30 degrees forward at time t, the inverse transformation will rotate the point on the right arm in the observation space by 30 degrees backward, so as to return to the position of the T-pose in the canonical space. If the point P belongs to the weapon, the inverse matrix of the posture transformation matrix calculated by the six-degree-of-freedom B-spline function at time t is applied. For example, if the weapon translates by 2 meters and rotates by 45 degrees at time t, the inverse matrix transformation will move it back to the origin position in the canonical space. Through such mapping, all dynamic foregrounds at different times are unified to the same static canonical scene for learning, which greatly simplifies the modeling difficulty of the neural radiance field.

[0063] In an optional embodiment, the joint optimization comprises:

[0064] Constructing a total loss function wherein is a photometric error between the rendered image and the real image is a binary cross-entropy loss between the accumulated opacity of the rendered ray and the foreground mask, is a parameter regularization term for the deformable human model, is a motion smoothness regularization term imposed on the B-spline function to penalize excessive velocity or acceleration.

[0065] To make the final reconstruction result optimal in both visual and geometric and physical laws, a joint loss function containing multiple sub-terms is adopted. The total loss L is a weighted sum of four key components, each of which is aimed at a specific optimization goal. The weight coefficients a, b, g are preset hyperparameters, which can be set to 0.1, 0.01 and 0.01 respectively, for example, to balance the relative importance of different loss terms.

[0066] The first term is a photometric loss, which ensures the visual fidelity of the reconstructed scene by calculating the color difference between the pixel color of the rendered image and the color of the real camera-captured image, for example, using the L2 norm. The second term is a mask loss, which compares the accumulated opacity along each ray rendered with the foreground mask using the binary cross-entropy, ensuring that the rendered foreground contour is accurately aligned with the contour extracted by the instance segmentation network, avoiding the problem of model edge blur. The third term is a human pose regularization term, which imposes constraints on the pose parameters of the deformable human model, penalizing poses that do not conform to human kinematics, such as excessive joint bending or limb penetration, ensuring the naturalness of human motion. The fourth term is a weapon motion smoothness regularization term, which penalizes the high-order derivative of the B-spline function, i.e., excessive velocity and acceleration, to ensure that the motion trajectory of the weapon is physically smooth and continuous, eliminating unrealistic situations such as instantaneous displacement or violent shaking.

[0067] More specifically, in one embodiment, the photometric loss is calculated as where R is a set of rays randomly sampled from all camera perspectives in one optimization iteration, r is a ray in the set R, is the real pixel color of the ray r, is the rendered pixel color, if the pixel is outside the mask, is rendered by the first neural radiance field, otherwise, it is rendered by the first and second neural radiance fields. The mask loss is calculated as wherein is a set of rays that only pass through the foreground region, is the real mask value corresponding to the ray r, if the pixel belongs to the foreground, If it belongs to the background, ; The rendered opacity of the light, i.e. the soft mask or the probability of the light hitting the object, is accumulated for each light. The regularization term of the morphable human model is calculated as where is the pose prior loss, which penalizes unnatural joint angles, preferably implemented by a pre-trained Variational Autoencoder (VAE) or Gaussian Mixture Model (GMM); is the body shape prior loss, which penalizes unrealistic body proportions, preferably the principal component analysis (PCA) coefficients of the body shape parameters of the human model. The weapon motion smoothness regularization term is calculated as where S(t) is the B-spline function of SE(3) of the weapon's six degrees of freedom motion, S(t) outputs a SE(3) matrix at any time t, representing the pose (rotation matrix) and position (translation vector) of the weapon, and are the first and second order derivatives of the B-spline function with respect to time t. In the above formula, denotes the square of the L2 norm, for example and are RGB three-channel color vectors, for example (R, G, B), the difference between them is a difference vector (AR, AG, AB), the norm is , then .

[0068] In an optional embodiment, the parameterized six degrees of freedom B-spline function is constructed to describe the rigid body motion of the weapon, comprising:

[0069] A cubic B-spline function with control points as a series of SE(3) group elements is constructed, which fits the continuous motion trajectory of the weapon in the time dimension by optimizing the control points, wherein the SE(3) group elements encode the three-dimensional rotation matrix and translation vector of the weapon.

[0070] In order to accurately represent the continuous motion of the rigid body of the weapon during the training process, a continuous time function is constructed. The cubic B-spline function is selected because it has the excellent property of second-order derivative continuity, and the motion trajectory described by it is smooth in velocity and acceleration, which is consistent with the motion law of objects in the real world.

[0071] The B-spline function is a series of elements on the special Euclidean group SE(3). Each SE(3) group element is a 4 by 4 matrix, which can compactly encode both the 3D rotation information, i.e. a 3 by 3 rotation matrix, and the 3D translation information, i.e. a 3 by 1 translation vector, of the weapon. For example, for a 10-second sword dance, only 20 SE(3) control points need to be set. By optimizing the positions of these 20 control points, the B-spline function can generate a complete, smooth, 6-DOF motion trajectory covering the entire duration and queryable at any millisecond, greatly reducing the computational complexity and overfitting risk compared to optimizing thousands of independent poses.

[0072] In an optional embodiment, the obtaining of the 3D motion trajectory, pose, linear velocity, angular velocity and acceleration of the weapon during the entire training process comprises:

[0073] From the optimized B-spline function, the SE(3) matrix at any time t is obtained by interpolation as the 3D pose and position of the weapon at that time;

[0074] The first and second time derivatives of the B-spline function are calculated and mapped from the Lie algebra SE(3) space to the Euclidean space to analytically obtain the linear velocity, angular velocity, linear acceleration and angular acceleration of the weapon.

[0075] When the joint optimization of the entire system converges, a final version of the B-spline function that accurately describes the motion of the weapon is obtained. Using this function, detailed kinematic analysis can be performed. To obtain the accurate 3D pose and position of the weapon at any time t, for example at 2.51 seconds, only need to substitute t = 2.51 into the B-spline function for interpolation calculation to obtain an SE(3) matrix. The rotation part of the matrix gives the orientation of the weapon at that time, and the translation part gives its coordinates in three-dimensional space. By continuously querying the poses at different times, the complete motion trajectory of the weapon during the entire training process can be plotted.

[0076] Since the B-spline function is analytically derivable, its first and second time derivatives can be directly calculated to obtain high-precision velocity and acceleration information. The first derivative of the function gives a result in the Lie algebra SE(3) space, representing the instantaneous velocity. Through a standard mapping relationship, it can be converted into a six-dimensional vector, of which the first three dimensions are angular velocity in radians per second, and the last three dimensions are linear velocity in meters per second. Similarly, the second derivative of the function and the mapping can give the angular acceleration and linear acceleration of the weapon. For example, by this method, the peak speed of the gun tip can be accurately calculated from 0.1 seconds from rest to 15 meters per second, and the acceleration curve during the entire process, avoiding numerical errors and noise caused by difference calculation based on discrete frames

[0077] Those skilled in the art can clearly understand that the application can be implemented by means of software plus necessary universal hardware platforms from the description of the above embodiments. Based on such an understanding, the technical solutions of the application can be embodied in the form of a software product, and the computer software product can be stored in a storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, and the like) execute the methods described in various embodiments or some parts of the embodiments of the application.

[0078] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. In particular, for the system or the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the related parts can be referred to the part of the method embodiment. The system and the system embodiment described above are merely illustrative, and the units described as separate components can be or can not be physically separated, and the components displayed as units can be or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. According to the actual needs, some or all of the modules can be selected to achieve the purpose of the embodiment. Those skilled in the art can understand and implement it without creative labor.

[0079] The method for providing commodity object information and the electronic device provided by the application are described in detail above, and the principles and implementation manners of the application are described by applying specific examples. The above embodiment is only used to help understand the method of the application and its core idea; meanwhile, for those skilled in the art, according to the idea of the application, the specific implementation manner and application range can be changed. In conclusion, the content of the specification should not be understood as a limitation of the application.

Claims

1. A method for analyzing the training trajectories of long and short weapons based on three-dimensional reconstruction, characterized in that, Includes the following steps: Acquire multi-view synchronous video sequences and camera internal and external parameters for weapons training processes; The training scene is decomposed into a static background field and a dynamic foreground field. The dynamic foreground field includes the trainee and the weapon. The static background field is represented by a first neural radiation field and the dynamic foreground field is represented by a second neural radiation field. A parameterized six-DOF B-spline function is constructed to describe the rigid body motion of a weapon, and a deformable human body model is constructed to characterize the non-rigid deformation of the human body. For any spatial point in the dynamic foreground field at time t, the observation space is transformed to a unified normed space using a mapping parameterized by the B-spline function and the deformable human body model, and the coordinates of the spatial point in the normed space are used as features input to the second neural radiation field. The camera light rays are sampled using volume rendering technology. For each sampling point, the color and density are queried and synthesized by combining the first neural radiation field and the second neural radiation field. The rendered image is obtained by integrating along the light rays. The photometric rendering loss between the rendered image and the real observed image is constructed, and the first neural radiation field and the second neural radiation field, the weapon six-degree-of-freedom B-spline function and the deformable human body model are jointly optimized. After the joint optimization converges, the optimal six-degree-of-freedom B-spline function of the weapon is extracted, and the zeroth, first and second derivatives are calculated to obtain the three-dimensional motion trajectory, attitude, linear velocity, angular velocity and acceleration of the weapon throughout the training process.

2. The method according to claim 1, characterized in that, The acquisition of multi-view synchronized video sequences and camera intrinsic and extrinsic parameters for the weapons training process includes: Eight to twelve high-speed industrial cameras with a frame rate of no less than 100fps were deployed in a ring around the training area, and the intrinsic and extrinsic parameters of all cameras were calculated using a checkerboard calibration board and a joint beam adjustment algorithm. A hardware synchronization controller ensures that the start timestamp error of all camera exposures is less than 1 millisecond.

3. The method according to claim 1, characterized in that, The process of decomposing the training scene into a static background field and a dynamic foreground field includes: Pixel-level masks of the trainee and weapon in each frame of the image are extracted using an instance segmentation network. During the joint optimization process, the photometric loss for optimizing the first neural radiation field is calculated based on the pixel region outside the mask, and the photometric loss for optimizing the second neural radiation field is calculated based on the pixel region covered by the mask.

4. The method according to claim 1, characterized in that, The transformation from the observation space to a unified normed space using a mapping parameterized by the B-spline function and the deformable human body model includes: For any sampling point in the observation space at time t, the sampling point is transformed to a unified standard space. Specifically, the inverse transformation of the deformable human body model is applied to the points belonging to the trainee, and the inverse matrix of the transformation corresponding to the six-degree-of-freedom B-spline function is applied to the points belonging to the weapon. Within this unified, standardized space, trainees and weapons together constitute a static, standardized scenario.

5. The method according to claim 1, characterized in that, The joint optimization includes: Construct the total loss function ,in To compensate for the photometric error between the rendered image and the real image To render the cumulative opacity of the light rays and the binary cross-entropy loss between the foreground mask, For the parameter regularization term of the deformable human body model, The motion smoothness regularization term is applied to the B-spline function to penalize excessive velocity or acceleration; where the weighting coefficients α, β, and γ are preset hyperparameters.

6. The method according to claim 1, characterized in that, The construction of the parameterized six-DOF B-spline function for describing the rigid body motion of a weapon includes: A cubic B-spline function with a series of SE(3) group elements as control points is constructed. The continuous motion trajectory of the weapon in the time dimension is fitted by optimizing the control points. The SE(3) group elements encode the weapon's three-dimensional rotation matrix and translation vector.

7. The method according to claim 1, characterized in that, The acquisition of the weapon's three-dimensional motion trajectory, attitude, linear velocity, angular velocity, and acceleration throughout the training process includes: From the optimized and converged B-spline function, the SE(3) matrix at any time t is obtained by interpolation, which serves as the three-dimensional attitude and position of the weapon at that time. The first and second time derivatives of the B-spline function are obtained and mapped from the Lie algebra SE(3) space to the Euclidean space. The linear velocity, angular velocity, linear acceleration and angular acceleration of the weapon are obtained analytically.

8. A training trajectory analysis system for long and short weapons based on three-dimensional reconstruction, characterized in that, Includes the following modules: The scene decomposition module is used to acquire multi-view synchronized video sequences and camera intrinsic and extrinsic parameters for the weapon training process; The training scene is decomposed into a static background field and a dynamic foreground field. The dynamic foreground field includes the trainee and the weapon. The static background field is represented by a first neural radiation field and the dynamic foreground field is represented by a second neural radiation field. The mapping module is used to construct a parameterized six-DOF B-spline function to describe the rigid body motion of a weapon, and a deformable human body model to characterize the non-rigid deformation of the human body. For any spatial point in the dynamic foreground field at time t, the mapping parameterized by the B-spline function and the deformable human body model is used to transform the observation space to a unified normed space, and the coordinates of the spatial point in the normed space are used as features input to the second neural radiation field. The optimization module is used to sample along the camera rays using volume rendering technology. For each sampling point, it combines the first neural radiation field and the second neural radiation field to query and synthesize color and density, and integrates along the rays to obtain the rendered image. It constructs the photometric rendering loss between the rendered image and the real observed image, and jointly optimizes the first neural radiation field and the second neural radiation field, the weapon's six-degree-of-freedom B-spline function, and the deformable human body model. The analysis module is used to extract the optimal six-degree-of-freedom B-spline function of the weapon after the joint optimization has converged, and to calculate the zeroth, first and second derivatives to obtain the three-dimensional motion trajectory, attitude, linear velocity, angular velocity and acceleration of the weapon throughout the training process.

9. The system according to claim 8, characterized in that, The acquisition of multi-view synchronized video sequences and camera intrinsic and extrinsic parameters for the weapons training process includes: Eight to twelve high-speed industrial cameras with a frame rate of no less than 100fps were deployed in a ring around the training area, and the intrinsic and extrinsic parameters of all cameras were calculated using a checkerboard calibration board and a joint beam adjustment algorithm. A hardware synchronization controller ensures that the start timestamp error of all camera exposures is less than 1 millisecond.

10. The system according to claim 8, characterized in that, The process of decomposing the training scene into a static background field and a dynamic foreground field includes: Pixel-level masks of the trainee and weapon in each frame of the image are extracted using an instance segmentation network. During the joint optimization process, the photometric loss for optimizing the first neural radiation field is calculated based on the pixel region outside the mask, and the photometric loss for optimizing the second neural radiation field is calculated based on the pixel region covered by the mask.

Citation Information

Patent Citations

  • Monocular human body reconstruction method and device based on IMU (Inertial Measurement Unit) and forward deformation field

    CN114581571A

  • 4D scene characterization method combining pose and radiation field optimization in complex mine environment

    CN119006687A