Automatic track and field action recognition and capture method based on image data

By combining the polarized neural radiation field with the Lie group skeleton graph network, the posture drift problem caused by motion blur and occlusion in track and field action recognition is solved, the synchronous reconstruction of dense three-dimensional shapes and joint rotations in a single perspective is achieved, and the accuracy and stability of key frame capture are improved.

CN120673479AActive Publication Date: 2025-09-19SHENYANG SPORT UNIV

Patent Information

Application Number
CN202510837727.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-09-19
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

Existing track and field motion recognition technology suffers from false detection, missed detection and drift under real training conditions, and it is difficult to provide stable and interpretable key frames in outdoor running tracks and mobile shooting environments.

Method used

Polarized neural radiation field and Lie group skeleton graph network closed-loop are used to reconstruct three-dimensional scenes. Multi-source saliency curves are generated by combining topological path features and liquid state machines, and key frame indexing is achieved through policy gradient dynamic weighting.

Benefits of technology

It achieves the synchronous reconstruction of dense 3D shapes and legal joint rotations under single-view conditions, overcomes the posture drift caused by motion blur and occlusion, and improves the robustness of action recognition and the accuracy of key frame capture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673479A_ABST
    Figure CN120673479A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of sports biomechanics, in particular to an automatic track and field action recognition and capture method based on image data, which comprises the following steps: synchronizing a visible infrared image, a polarization image and an event stream according to a mutual information criterion and generating a polarization phase diagram; inputting a polarized neural radiation field network and a skeleton graph neural network based on Lie group message passing, and alternately optimizing to obtain a continuous voxel light field and a joint rotation sequence; performing Vietoris-Rips persistent homology and third-order path signature on the sequence to extract topological path features, encoding the topological path features into a pulse sequence to drive a liquid state machine, and generating a semantic confidence flow by combining bone potential tensor and action text embedding; and constructing a saliency curve by using the voxel light field time derivative, the neuromorphic difference, the semantic complementary value and the skeleton difference, carrying out strategy gradient on-line adjustment on the weight, taking curve minimum mapping as a key frame index, and outputting an image frame. The track and field action key frames can still be accurately captured in real time under complex illumination and shielding conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sports biomechanics, and in particular to a method for automatically identifying and capturing track and field movements based on image data. Background Art

[0002] Track and field training and event analysis increasingly rely on image data. Automatic movement recognition and keyframe capture can provide objective quantitative metrics for coaches, fast replay clips for referees and broadcast systems, and detailed posture assessment for sports rehabilitation institutions. However, existing technologies mostly rely on two-dimensional skeleton detection with threshold triggering or simple optical flow analysis: ① Convolutional network-based two-dimensional keypoint tracking is only reliable in unobstructed frontal scenes and is prone to mismatch when encountering motion blur caused by high-speed movement or occlusion by hurdles; ② Optical flow mutation methods are sensitive to specular highlights, shadows, and camera shake, resulting in a high false detection rate; ③ Multi-view 3D reconstruction schemes, while highly accurate, are expensive and complex to configure, making them unsuitable for outdoor tracks and mobile filming. These factors result in existing systems experiencing a combination of false detections, missed detections, and drift under real-world training conditions, making it difficult to generate stable and interpretable keyframes. Summary of the Invention

[0003] In response to the many problems existing in the above-mentioned existing technologies, the present invention provides a method for automatic recognition and capture of track and field movements based on image data. The present invention uses a unified time-base data packet to drive the polarized neural radiation field and Lie group skeleton graph network closed-loop to reconstruct the three-dimensional scene, and then uses the topological path features through a liquid state machine and semantic matching to generate a multi-source saliency curve. Through dynamic weighting of policy gradients, the key frame index is output in real time.

[0004] A method for automatically identifying and capturing track and field movements based on image data comprises the following steps: Time synchronization is performed on the visible infrared image, polarization image, and event stream. The first and second Stokes components are extracted from the synchronized polarization image, and the polarization phase map is generated using the arc tangent values ​​of the two to form a unified time-base data packet. A unified time-base data packet is input into a polarized neural radiation field network to generate a voxel light field. A skeleton graph neural network based on Lie group message passing is used to obtain a joint rotation sequence. The joint rotation sequence is subjected to matrix product state decomposition to form a skeleton latent tensor. Through alternating optimization, the voxel light field, joint rotation sequence, and skeleton latent tensor are coupled into a continuous spatiotemporal scene model. Based on the joint rotation sequence, Vietoris–Rips filtering and third-order path signature are used to extract topological path features. The topological path features are encoded as a pulse sequence and input into the liquid state machine to obtain a neuromorphic state stream. The semantic confidence stream is generated by combining the similarity between the skeleton latent tensor and the action text embedding. The target coherence value and the pulse peak time in the neuromorphic state stream are fed back into the continuous spatiotemporal scene model. Based on the amplitude of the voxel light field temporal derivative, neuromorphic state difference, semantic confidence complementarity and skeletal latent tensor difference, a hybrid saliency curve is constructed with preset and online adjustable weights. The local minimum of the hybrid saliency curve is selected and mapped to a set of keyframe indices according to the preset sampling rate.

[0005] Preferably, time synchronization is achieved by maximizing the mutual information between the edge of the event stream and the edge of the visible infrared image, and using a unified hardware pulse signal to calibrate the timestamps of the visible infrared image acquisition device, the polarization image acquisition device and the event stream acquisition device.

[0006] Preferably, the polarization phase map is generated by comparing the amplitude relationship between the first Stokes component and the second Stokes component to determine the phase of each pixel, so as to compensate for the polarization direction difference.

[0007] Preferably, each node of the skeletal graph neural network based on Lie group message passing corresponds to a preset human joint, and message passing realizes joint posture update by performing group logarithmic mapping and group exponential mapping on the rotation differences of adjacent nodes in the three-dimensional rotation group space.

[0008] Preferably, the matrix product state decomposition performs low-rank reconstruction on the joint rotation sequence within a sliding time window to generate the skeleton latent tensor and compress temporal redundant information.

[0009] Preferably, the topological path feature extracts persistent homology features by performing Vietoris–Rips filtering on the joint rotation sequence, and combines it with the third-order path signature feature to form a joint representation.

[0010] Preferably, the liquid state machine is composed of multiple layers of spiking neurons, the synaptic delays are randomly initialized within a preset range, and the read layer weights are periodically updated to output a neuromorphic state stream.

[0011] Preferably, the semantic confidence flow is generated by comparing the cosine similarity between the skeleton latent tensor mapping vector and the Chinese action text embedding vector, and selecting the text label with the highest similarity as the current action label.

[0012] Preferably, the weight coefficients of the mixed saliency curve are adjusted online by a policy gradient algorithm and normalized after each adjustment to keep the sum of all weight coefficients constant.

[0013] Preferably, the key frame time points are mapped to key frame indices according to the frame rate of the image acquisition sequence, and corresponding image frames are extracted from the visible light infrared image sequence based on the key frame indices and output as key frames.

[0014] Compared with the prior art, the advantages and beneficial effects of the present invention are: By coupling a polarized neural radiance field network with a Lie group skeletal graph neural network, we achieve synchronized reconstruction of dense 3D shapes and legal joint rotations from a single viewpoint, overcoming the pose drift caused by motion blur and occlusion. By jointly extracting topological path features using Vietoris–Rips persistent coherence and third-order path signatures, we achieve robust encoding of the global shape and temporal rhythm of the action, addressing the optical flow method's sensitivity to illumination noise. By fusing a liquid state machine with a semantic confidence stream, we achieve millisecond-level advance prediction of keyframes and automatic filtering of non-target actions, addressing the high false detection rate associated with threshold triggering. By adjusting saliency weights online through policy gradients, we achieve cross-scene adaptation and maintain a high recall rate without the need for recalibration. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 Schematic diagram of the process of the present invention; Figure 2 This is a schematic diagram of constructing the mixed significance curve in the present invention. DETAILED DESCRIPTION

[0016] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure.

[0017] like Figure 1 As shown, a method for automatically identifying and capturing track and field movements based on image data includes the following steps: Time synchronization is performed on the visible infrared image, polarization image, and event stream. The first and second Stokes components are extracted from the synchronized polarization image, and the polarization phase map is generated using the arc tangent values ​​of the two to form a unified time-base data packet. Visible-infrared images, polarized images, and event streams originate from three heterogeneous sensors, each with different frame rates and time bases. Without first establishing a unified time base, the temporal correspondence between the different modalities will drift, leading to skeleton misalignment or light field distortion during subsequent motion recognition and keyframe capture. This invention, in principle, treats time synchronization as a cross-modal registration problem, introducing a mutual information maximization criterion supplemented by hardware pulse calibration to achieve millisecond-level alignment.

[0018] Mutual information is a statistic that measures the joint uncertainty of two sets of random variables. For edges in the event stream and visible-infrared image, their positions strictly correspond in the time domain. When the two data streams slide to the optimal alignment, the joint entropy reaches a minimum and the mutual information reaches a maximum. In implementation, the event stream is integrated into a grayscale image within a short time window. Edge detection is then performed with the visible-infrared image within the corresponding window. The mutual information is calculated and the maximum point is found through a one-dimensional search. This search is performed within milliseconds, requiring minimal computation. To prevent mutual information distortion in scenarios with sudden changes in lighting, a common pulse signal is used at the hardware level to mark the timestamps of the three sensors. At algorithm startup, a coarse alignment is performed to the pulse boundary, followed by a fine-tuning of the mutual information. This dual mechanism ensures stable synchronization even under complex outdoor lighting conditions.

[0019] Polarization information can reveal the microscopic orientation of the surface of an object, and is particularly sensitive to muscle texture, clothing wrinkles, and metal reflections of equipment in track and field movements. The present invention uses the polarization phase map as an additional channel for subsequent neural radiation field reconstruction to constrain the voxel normal direction of the radiation field and suppress local reflection noise. The polarization sensor has a built-in four-point polarizer that outputs four intensity maps at different polarization angles. According to polarization measurement theory, the first Stokes component Describes the difference in polarization intensity between horizontal and vertical directions, the second Stokes component Describes the difference in polarization intensity between the 45° and 135° directions. The present invention uses only these two items to calculate the phase without introducing the third Stokes component, thereby reducing additional noise channels.

[0020] Polarization phase is calculated using the inverse tangent relationship. The core formula is:

[0021] in Pixels The polarization phase at ; Indicates the intensity difference of the first Stokes component of the pixel; The reason for using inverse tangent instead of simple ratio is that inverse tangent can convert The phase range is mapped to the grayscale range to achieve continuous transition of the cycle and avoid When it is close to 0, the value diverges. In order to improve the robustness, the present invention calculates and First, a three-by-three median filter is performed to suppress random noise spikes, and then a lookup table method is used to accelerate the arctangent operation to ensure that the additional delay of a single frame does not exceed 3 milliseconds in real-time scenarios.

[0022] After forming a unified time-base data packet, the present invention sends the visible-infrared image, polarization phase map and event stream together into the polarization neural radiation field network. In this process, the phase information mainly plays two roles: first, by inputting the phase as an additional voxel feature into the network, it helps the network to identify the true normal direction of small and medium texture areas at the same time, and reduce the stereo ambiguity caused by lens motion blur; second, when constructing the skeleton graph neural network later, the phase map can provide polarization consistency constraints for clothing reflections around the joints, reducing the bone drift caused by clothing wrinkles. In the experiment, on the hurdle jump test set, after the phase map was introduced, the average reprojection parallax of the skeleton posture was reduced by about 12%, and the timeliness of the key frame capture was increased by 8%.

[0023] Preferably, time synchronization is achieved by maximizing the mutual information between the edge of the event stream and the edge of the visible infrared image, and using a unified hardware pulse signal to calibrate the timestamps of the visible infrared image acquisition device, the polarization image acquisition device and the event stream acquisition device.

[0024] In the multimodal track and field motion acquisition system, visible-infrared images, polarization images, and event streams are output by three independent sensors, and their internal clocks are inconsistent at the microsecond to millisecond level. If the unaligned data is directly fed into the subsequent neural radiation field reconstruction and skeletal posture estimation module, the same action frame will be divided into different time slices, resulting in voxel density drift, skeletal node dislocation, and key frame misjudgment. The present invention first uses the common pulse synchronization signal at the hardware layer to complete the coarse calibration, and then uses mutual information maximization at the algorithm layer to complete the fine correction to ensure that the three-way data is strictly aligned within the millisecond time window.

[0025] The hardware pulse synchronization signal, via a distributed clock module, periodically sends rising-edge triggers to the three sensors. Each time a sensor detects a trigger edge, it records a local timestamp and marks the frame containing that pulse as the reference frame. Because all trigger edges originate from the same physical pin, the coarse time difference between the three sensors is determined by line delays and internal buffer delays, typically within tens of microseconds, allowing for initial alignment with a global reference.

[0026] After coarse alignment, the residual drift caused by frame reading delay, encoding cache and system bus delay still needs to be solved. The present invention uses mutual information as a fine-tuning indicator for cross-modal time alignment. Mutual information simultaneously quantifies the joint entropy and conditional entropy of the two sets of signals, and can measure the statistical dependence between quasi-synchronous signals without assuming a linear relationship. The event stream records the brightness change symbols within an extremely short exposure, which is essentially a high-frequency spatial edge excitation. The visible-infrared image contains spatial gradient information at the same time. When the two data channels are in the correct alignment position, the event stream activation point highly overlaps with the visible-infrared image edge, the coupling degree of the joint probability distribution is maximized, and the mutual information reaches a maximum. In order to make the mutual information calculation more stable, the present invention accumulates the event stream in a short time window to obtain grayscale fitting, and then performs Sobel edge detection on the visible-infrared image. The accumulation window length is automatically derived from the pulse period, which not only ensures the stability of the statistics, but also avoids motion blur caused by an overly wide window.

[0027] Mutual information maximization is achieved through one-dimensional search. The event accumulation edge graph is a random variable , the visible-infrared edge map is a random variable , the mutual information between the two is given by the formula:

[0028] Definition, where is the joint probability, 、 The present invention performs a limited range linear scan on the delay amount, and each delay amount corresponds to a pair of synchronization windows, and estimates the delay amount through fast histogram statistics. To reduce logarithmic computation overhead, the joint probability is multiplied by a fixed scaling factor and then approximated using a table lookup. The location with the maximum mutual information is used as the final time offset, which is then used to adjust the timestamps of the polarization image, visible-infrared image, and event stream to form a unified time-base data packet.

[0029] The theoretical optimality of mutual information search stems from the principle of maximizing statistical interdependence between signals, rather than relying on the consistency of signal amplitude or frequency. For high-speed arm swings and foot strikes in track and field movements, event stream edges often exhibit sparse and extreme activation. Compared to traditional alignment methods based on mean square error or cross-correlation, mutual information is more robust to amplitude ratios and noise. Experiments have shown that when athletes sprint or hurdle at speeds exceeding 10 meters per second, adjacent frames experience severe motion blur, and the cross-correlation peak exhibits multiple local extremes, while the mutual information curve maintains a single peak, significantly reducing the probability of mismatches.

[0030] To verify the effect, the embodiment uses actual shooting data of the 100-meter race in the training field, including visible-infrared images at 240 frames per second, polarization images at 90 frames per second, and event streams at 10 kHz. First, the three timestamps are roughly aligned using hardware pulses, and the mean square error of the initial time deviation is measured to be less than 0.1 milliseconds. Then, a mutual information fine search is used within each pulse interval, the search step is set to 0.1 milliseconds, and the search window is 4 milliseconds long. Finally, the remaining time error mean is 0.3 milliseconds and the standard deviation is 0.15 milliseconds. The aligned data is input into the subsequent network of the present invention, and the reprojection parallax is reduced by 15% compared to the case where only hardware pulse alignment is used, and the key frame recall rate is increased by 7%. Another set of hurdle jump data was shot in a backlit environment. Due to the increase in polarization noise caused by light reflection, the standard deviation of the error is maintained at 0.2 milliseconds after synchronization by the present invention, proving that the algorithm layer fine calibration is insensitive to lighting changes.

[0031] To further reduce computational complexity, the present invention incorporates an adaptive step-size strategy into the search process: a larger step-size is used in the initial scan to approximate the peak location, followed by a smaller step-size scan on either side of the peak. Furthermore, a parallel prefix sum is used to accelerate histogram normalization when calculating the joint histogram. Combined with GPU pipeline processing, this approach can achieve online synchronization at over 50 frames per second on a gaming-grade graphics card.

[0032] Preferably, the polarization phase map is generated by comparing the amplitude relationship between the first Stokes component and the second Stokes component to determine the phase of each pixel, so as to compensate for the polarization direction difference.

[0033] Polarization information is an important dimension for describing the vibration direction of the light field, and together with intensity and wavelength, it constitutes a complete radiation feature. In track and field scenarios, the mirror reflections of the athlete's muscle fibers, tights fabrics, and metal equipment surfaces will change the polarization direction of the incident light, resulting in visible light intensity not strictly corresponding to the true geometric normal. Relying solely on intensity texture to construct a neural radiation field or skeletal posture network is prone to local depth dislocation and joint rotation drift. The present invention introduces a polarization phase map as an additional input channel to compensate for polarization direction differences at the pixel level, significantly improving the accuracy and steady-state robustness of three-dimensional reconstruction and key frame determination.

[0034] The quad polarization camera outputs four intensity images taken at 0°, 45°, 90° and 135° directions at the same time by placing a micro-polarizer array in front of the photosensitive chip. The linear polarization state can be described by a three-dimensional Stokes vector, the first component of which is Indicates the intensity difference between the horizontal and vertical directions, the second component Represents the intensity difference in the diagonal direction. The present invention selects these two terms to determine the phase, because the phase is only related to the vibration direction and has nothing to do with the polarization degree. In order to maintain dimensional consistency with the density and color tensors in the subsequent network, the polarization phase is mapped to a continuous phase interval using the inverse tangent function. The core calculation expression is:

[0035] Where, Pixels The polarization phase at ; is the first Stokes component; The first Stokes component is obtained by subtracting the 0° and 90° images, and the second Stokes component is obtained by subtracting the 45° and 135° images. and Bilateral filtering is performed to suppress high-frequency noise, and then the inverse tangent is approximated using a lookup table to reduce hardware instruction delays.

[0036] This method does not use the third Stokes component and does not calculate the degree of polarization. On the one hand, the third component is close to zero under natural illumination and has a limited contribution to the phase. On the other hand, reducing the number of computation channels reduces memory bandwidth, ensuring online inference speed. Furthermore, the phase map is only concatenated with the color tensor at the first layer of the network and does not participate in the back propagation of the light field voxel density gradient, thus avoiding error amplification in the presence of unstructured noise such as raindrops and dust.

[0037] The direct effect of the phase map is to provide microsurface orientation information for each pixel, giving the neural radiation field explicit constraints when estimating the voxel normal. In running scenes, the athlete's calf muscles swing rapidly with the stride frequency. Traditional voxel reprojection based on color consistency is prone to producing streaking artifacts in the specular highlight area. After adding the phase channel, the network can use the position of the phase jump to determine the true reflection direction, thereby correctly estimating the calf surface shape. The indirect effect is reflected in the skeletal graph neural network. After taking phase consistency weighting for the pixels in the joint neighborhood, the joint heat map suppresses the false peaks at the folds of clothing and reduces the jitter of joint position.

[0038] Example: An athlete completing a hurdle jump was filmed outdoors under backlit conditions. After using the phase image compensation presented in this invention, the Euclidean angle error between the joint rotation matrix corresponding to the highest point of the hurdle jump and the laser motion capture reference value was reduced from 6.8 degrees to 4.1 degrees, and the dynamic time warping distance between the skeletal potential tensor and the reference trajectory was reduced by 19%.

[0039] In order to make full use of the phase information, the present invention also adds a "phase consistency loss" to the phase map during the training phase. and Take the absolute difference and weight it into the light field loss:

[0040] in is the phase consistency weight. This loss encourages the network to maintain local polarization consistency in consecutive frames and suppress the reverse noise introduced by slight camera shake. Training experiments show that when When choosing the same dimension as the main light field loss, the reprojection parallax is further reduced by about 6% and the number of network convergence iterations is reduced by about 12%.

[0041] The polarization phase calculation of the present invention does not require additional optical components on the hardware side and can be completed only through software post-processing, making it easy to directly deploy on existing multi-modal camera systems. Compared with common multi-view stereo solutions, the present invention uses phase to introduce normal constraints in a single-view configuration, solving the problem of track and field training venues being unable to arrange multiple cameras due to space limitations. This allows for deployment in university track and field halls, outdoor running tracks, and other venues at a lower cost. Actual tests have shown that a single fisheye lens plus a polarization camera can cover one to three lanes of a hurdle track, reducing the amount of hardware by half without sacrificing recognition accuracy.

[0042] like Figure 2 As shown, a unified time-base data packet is input into a polarized neural radiation field network to generate a voxel light field. A skeleton graph neural network based on Lie group message passing is used to obtain a joint rotation sequence. The joint rotation sequence is subjected to matrix product state decomposition to form a skeleton latent tensor. Through alternating optimization, the voxel light field, joint rotation sequence, and skeleton latent tensor are coupled into a continuous spatiotemporal scene model. The unified time base data packet contains a joint image of visible light and infrared light, a polarization phase map, and an event flow grayscale map. In the pipeline of the present invention, the data packet is first sent to the polarization neural radiation field network in batches. The neural radiation field network takes voxel coordinates and the direction of the sampled light as input, outputs voxel density and color through a multi-layer perceptron, and splices the polarization phase features in the first hidden layer, so that the network obtains additional directional constraints when solving the surface normal of each voxel. In order to avoid the periodicity of the polarization phase to produce discontinuous gradients, the network uses a sine-cosine dual-channel encoding for the phase at the input end, so that The range remains smooth. Unlike traditional radiation fields, the voxel features of the present invention not only contain 3D density and 3D color, but also store polarization normal encoding. This adds a phase consistency loss to the rendering loss, minimizing the phase residuals of adjacent frame voxels during the training phase and further improving the accuracy of voxel normal estimation.

[0043] The core integral of voxel rendering follows the volume rendering formula:

[0044] in For sampling light, represents the voxel density, is the distance between adjacent sampling points, is the voxel color vector, is the residual transmittance before the current voxel. Density, color, and phase share gradients during backpropagation, so any phase estimation error is propagated back along with the color error, forcing the network to correct both photometric and directional information during iteration.

[0045] In the human body motion modeling part, the present invention adopts a skeleton graph neural network based on Lie group message passing. Lie group message passing refers to performing convolution operations on the three-dimensional rotation group space to ensure that the network remains in the rotation group after updating without generating invalid postures. Its adjacent joints , define the relative rotation:

[0046] in 、 is the rotation matrix. The vector is obtained by converting the Lie group logarithm map into the Lie algebra Graph Neural Networks use As messages, after undergoing multiple layers of linear transformations and nonlinear activations, the updated values ​​are re-projected back into the rotation group using Lie group exponential mapping to obtain the new joint pose. This process ensures that message computation is performed in Euclidean space, while the final result remains valid in rotation space, preventing cumulative drift in joint poses during iterations.

[0047] The joint rotation sequence output by the skeletal graph neural network grows rapidly over time. This paper uses matrix product state decomposition to compress the long sequence. The rotation matrix of 40 consecutive frames is expanded along the time dimension into a high-order tensor. , and then decomposed into a chain tensor network:

[0048] in For the The kernel tensor of the frame, is the internal bonding index. By truncating the bonding dimension to 32, the main variation patterns of the pose sequence can be retained with a smaller storage capacity. The resulting one-dimensional latent vector is the skeleton latent tensor representation, which is used for subsequent topological feature, semantic matching, and hybrid saliency calculations.

[0049] In order to allow the voxel light field and the skeleton model to calibrate each other, the present invention adopts an alternating optimization strategy during the training phase. First, the skeleton parameters are fixed, and only the voxel network is updated to minimize the rendering error and phase consistency; then the voxel network is fixed, the rendering result is projected into the two-dimensional image space, the skeleton reprojection error in the camera frame is calculated, and then the skeleton graph network weights are updated. The two are performed alternately, and a spatiotemporal scene model is formed after multiple iterations: the voxel light field provides dense shape and texture, the skeleton sequence describes the motion skeleton, and the skeleton potential tensor compresses and transmits global temporal information. Since the voxel network uses the skeleton heat map to perform high-weight sampling on key parts during training, and the skeleton network uses the new voxel normal to re-estimate the Lie group message in each round of iteration, the two networks form a closed loop, so that the model remains stable under occlusion, mirror reflection and high-speed motion scenes.

[0050] Example: A hurdle jump was filmed outdoors under backlit conditions. After adding polarization phase, the joint rotation error at the highest point of the hurdle jump was reduced from 6.8° to 4.1°, and the dynamic time warping distance between the skeletal potential tensor and the laser reference trajectory was reduced by 19%.

[0051] After compressing the 40-frame rotation matrix to a bond dimension of 32 using matrix product state decomposition, the latent vector storage capacity is approximately 4120 bytes, while directly storing the rotation matrix exceeds 200,000 bytes. The ability to maintain the skeletal latent tensor in real time and interact with the voxel network in a single-chip microcontroller plus mobile graphics card system demonstrates the engineering deployability of this invention.

[0052] Through the alternating optimization of polarized neural radiation field networks, skeletal graph neural networks based on Lie group message passing, and matrix product state decomposition, the present invention achieves high-precision coupling of voxel light fields and skeletal sequences in the continuous spatiotemporal domain, significantly improving the stability and accuracy of automatic recognition of track and field movements and key frame capture.

[0053] Preferably, each node of the skeletal graph neural network based on Lie group message passing corresponds to a preset human joint, and message passing realizes joint posture update by performing group logarithmic mapping and group exponential mapping on the rotation differences of adjacent nodes in the three-dimensional rotation group space.

[0054] In the tasks of track and field motion recognition and keyframe capture, the core information of human posture is the rotation and displacement of each joint in three-dimensional space. Traditional joint rotation regression is usually expressed in Euler angles or quaternions, but Euler angles have universal locks, and quaternions need unit normalization to avoid singularity. Moreover, both parameterizations cannot guarantee that they still fall into the legal rotation set after update during the neural network propagation stage. The present invention adopts a skeletal graph neural network based on Lie group message passing to directly model the rotation in a special orthogonal group space. , naturally satisfies rotational closure, while using the graph structure to explicitly encode skeletal topology constraints.

[0055] The skeletal graph neural network consists of 22 nodes, each of which corresponds to a fixed human anatomical joint, including the ankle, knee, hip, shoulder, elbow, wrist, and thoracolumbar region. The edges between nodes are established according to the real bone connection relationship, for example, the tibia node is connected to the femur node, and the scapula node is connected to the humerus node. The network input is the initial rotation matrix sequence of each node in consecutive frames. In one message delivery cycle, for nodes Its adjacent nodes , the network first calculates the relative rotation:

[0056] The superscript Represents the current graph convolution layer. Rotation difference lie in space, but message aggregation needs to be done in vector space, so Through Lie group logarithmic mapping Projection to Lie algebra , we get the vector:

[0057] The geometric meaning is from the node Chaok Node The minimum rotation axis angle representation can be directly used for vector addition. The network will After multi-head attention weighting, and node The self-feature splicing is sent to the multi-layer perceptron to obtain the update vector To ensure that the update results still fall within , and then use the Lie group exponential mapping:

[0058] in is the Lie group exponential mapping, exist In space, this is efficiently achieved through the Rodrigues formula. The present invention sets up 6 layers of graph convolution, each of which executes the above-mentioned "rotation difference - logarithmic mapping - exponential mapping" cycle to propagate posture information layer by layer.

[0059] This Lie group message passing approach offers three fundamental advantages. First, the group logarithmic mapping transforms the nonlinearly dispersed rotation differences into a linear space, enabling deep networks to directly apply linear transformations while maintaining gradient stability. Second, the exponential mapping projects the updates back into the group space rather than simply adding them, preventing the network output from falling into regions with illegal rotations. Third, the rotation differences are inherently relatively invariant, making them robust against camera displacement and partial occlusion.

[0060] To improve training efficiency, the network initially uses 2D joint points extracted from visible and infrared images as weak supervision, and uses joint reprojection errors to guide convergence. Once the network stabilizes, a surface normal consistency loss derived from polarized neural radiance field rendering is added to align 3D rotations with true voxel normals. Because rotation information is accumulated in the graph structure, distal joints such as the wrist receive a global prior from the trunk nodes, which is particularly effective for predicting joint drift during high-speed arm swing.

[0061] In track and field applications, dozens of rapid arm swings or foot contact movements can occur within one second. If the Euler angle network is used, the error will accumulate over time and cause the wrist rotation to drift by more than 5 degrees; by using the Lie group message passing skeletal graph neural network of the present invention, the drift is suppressed to within 1.2 degrees. Specific embodiment: On a data set with 240 frames per second, taking the windless indoor 100-meter sprint as an example, 120 sprint clips of 10 athletes were recorded. Using a laser motion capture system as the true value, the baseline quaternion network has an average posture error of 4.9 degrees at the end of a 5-second clip; the error of the network of the present invention is 1.3 degrees, and the key frame recall rate is improved by 9%. In another set of hurdle data, the network still maintains a stable joint rotation differential in the stage of severe occlusion above the hurdle, proving that the relative rotation message has the ability to compensate for missing frames.

[0062] In terms of hardware deployment, the core computation of the Lie group logarithm / exponential mapping involves matrix diagonal block splitting and Rodriguez's formula, enabling simultaneous processing of all node pairs using thousands of threads in a GPU parallel environment. Compared to traditional joint regression, this network only increases inference time by approximately 8% within a single batch, while significantly improving pose stability and providing high-quality input for subsequent skeletal latent tensor compression and hybrid saliency curve calculations.

[0063] The alternating optimization strategy maps the rotation differences of the latest joint poses back to voxel sampling weights, prompting the voxel network to increase ray density in high-motion areas such as the arms and knees. The updated voxel network then outputs more accurate normals to provide new supervision for the skeletal network, achieving positive feedback between the two. In cross-scenario experiments, switching from an outdoor track to an indoor training gym required only 1,000 steps of fine-tuning for the voxel network and 400 steps for the skeletal network to achieve convergence, demonstrating transfer flexibility.

[0064] Preferably, the matrix product state decomposition performs low-rank reconstruction on the joint rotation sequence within a sliding time window to generate the skeleton latent tensor and compress temporal redundant information.

[0065] Matrix product state decomposition belongs to the category of tensor networks and was first used for low-rank approximation of quantum many-body systems. The present invention introduces this idea into track and field motion analysis to compress joint rotation sequences that grow rapidly over time, and thereby generate skeletal potential tensors. The rotation information can be represented as a high-order tensor stacked in sequence. If 40 frames are used as a group of time windows, a dense tensor of order 40 and dimension 22 is obtained. Direct storage requires a large amount of video memory, which is not conducive to deployment on mobile terminals or real-time systems. The matrix product state reconstructs high-order tensors through chained core tensors, retaining only the main change patterns, significantly reducing storage and computational complexity, while maintaining the ability to distinguish human body dynamics.

[0066] The processing flow first maps the joint rotation matrix of each frame in the sliding window to a Euclidean vector. The mapping operation uses the logarithmic form of Rodriguez's formula to convert the three-dimensional rotation matrix into a three-dimensional Lie algebraic rotation vector, which avoids the singularity of the Euler angle and maintains linear additivity. By connecting the rotation vectors of 22 joints in series, a frame-level feature of length 66 can be obtained. After collecting 40 consecutive frames with a sliding window, a second-order tensor with a shape of 66×40 can be constructed. In order to process in the matrix product state framework, the second-order tensor is converted into a high-order form in the time dimension. ,in Indicates the The index of the frame.

[0067] The present invention uses a fixed bond dimension to perform matrix product state decomposition on this high-order tensor. The core expression is:

[0068] Where, is a high-order rotation tensor, For the The kernel tensor corresponding to the frame, is the internal bonding index. By truncating the bonding dimension to 32, the main variation patterns of the pose sequence can be retained with a smaller storage capacity. The resulting one-dimensional latent vector is the skeleton latent tensor representation, which is used for subsequent topological feature, semantic matching, and hybrid saliency calculations.

[0069] To ensure real-time performance, the present invention uses block singular value decomposition on the graphics card to implement core tensor extraction. The decomposition time for each sliding window is approximately 2.3 milliseconds, which is much lower than the frame interval of track and field motion capture. Comparative experiments show that under the same hardware environment, directly inputting stacked rotation matrices into the semantic network will cause the inference video memory to occupy more than 300 megabytes, while using the matrix product state will reduce the occupancy to approximately 15 megabytes. Further cross-sectional analysis shows that most of the video memory comes from the core tensor gradient buffer. If the gradient is turned off during the inference phase, the video memory occupancy is less than 5 megabytes, which can meet the inference requirements of the embedded end.

[0070] The embodiment demonstrates the effect of the method. On the 100-meter sprint cross-scene dataset, the traditional long short-term memory network was used to process a 40-frame rotation vector sequence as a baseline, with a semantic recognition accuracy of 85% and a key frame recall rate of 83%. After replacing it with the matrix product state latent tensor, the accuracy was increased to 91% and the key frame recall rate was increased to 90% under the same network structure. Analysis found that the matrix product state latent tensor is insensitive to high-frequency jitter of the joints, but maintains a high response to the real action cycle, so that the noise and action inflection points can be separated on the mixed significance curve. In the hurdle dataset experiment, under the same window length and bonding dimension settings, the combination of the latent tensor and the topological persistent homology feature can predict the hurdle jump key frame one time window in advance, providing redundancy for real-time referee auxiliary evaluation.

[0071] The present invention employs alternating optimization during model training. When updating the voxel network, the temporal pattern encoded in the latent tensor is used to control ray sampling density, extending the integration duration for frames with intense action and shortening it for still frames. When updating the skeletal graph network, the integrated depth map is read from the voxel network to constrain joint rotations. After the alternating optimization converges, the skeletal latent tensor retains the dominant motion mode of each window, serving as a lightweight index for the overall scene model, facilitating subsequent retrieval or incremental fine-tuning.

[0072] The matrix product state is also highly interpretable. By examining the eigenvectors of the kernel tensor, the principal motion direction within the window can be intuitively identified. For example, in a hurdle jump, the first principal component corresponds to pelvic rotation during the takeoff phase, while the second principal component corresponds to arm swing amplitude. Coaches and athletes can use principal component analysis of the latent tensor to make targeted corrections to technical details and improve training efficiency.

[0073] Based on the joint rotation sequence, Vietoris–Rips filtering and third-order path signature are used to extract topological path features. The topological path features are encoded as a pulse sequence and input into the liquid state machine to obtain a neuromorphic state stream. The semantic confidence stream is generated by combining the similarity between the skeleton latent tensor and the action text embedding. The target coherence value and the pulse peak time in the neuromorphic state stream are fed back into the continuous spatiotemporal scene model. Joint rotation sequences are represented as high-dimensional curves in Euclidean space. Their essential structure encompasses both continuous temporal order and the overall shape topology. To capture both types of information simultaneously, this paper projects the rotation sequence into the topological and path algebraic domains before performing neuromorphic encoding. The resulting real-time state stream is then used to inversely constrain the spatiotemporal scene model, forming a closed-loop analysis framework. The following sections explain the principles, implementation, and resulting technical effects of each step.

[0074] First, in each sliding window, the rotation vector sequence output by the Lie group message passing skeleton graph network is subjected to Vietoris–Rips filtering. This filtering method constructs a simple complex with a gradually increasing distance threshold, tracking the generation and disappearance of one-dimensional connected loops. For a given rotation vector set in the window, , when the distance threshold is When the distance between two points is less than Then it is connected, and a higher order simplex is constructed. Starting from zero growth, high-order topological features gradually emerge and eventually fill and disappear. The longer the connected loop persists, the more it can reflect the overall movement rhythm of the skeleton. The present invention records the birth threshold and death threshold of each persistent entry, and uses a two-dimensional Gaussian kernel to project the birth and death coordinates into a persistent image. When a complete leg-lifting or arm-swinging cycle occurs during the movement process, the corresponding entry shows a high response in the persistent image; in the case of jitter or short-term occlusion, the corresponding entry disappears very quickly, thereby naturally filtering out noise.

[0075] Topological features only describe the overall shape but cannot express the direction and speed of movement. To describe the temporal sequence, the present invention calculates the rotation vector trajectory within the same window to a third-order path signature. The path signature is a series of iterative integrals that can be viewed as encoding the path into the coordinates of a free Lie algebra. The third-order truncation is sufficient to cover the velocity and acceleration information in track and field movements without excessive dimensionality expansion. Because the rotation vector is normalized to a fixed scale, the signature is numerically stable and easy to post-process using a neural network.

[0076] Persistent images differ from third-order path signatures in terms of dimension. This invention uses linear projection to concatenate the two into a topological path feature vector of uniform length, and then performs sinusoidal time coding on each vector channel to generate a pulse frequency. The pulse frequency drives the liquid state machine according to a Poisson distribution. The liquid state machine is a feedforward pulse neural network, with one thousand Izhikevich neurons randomly connected to form a high-dimensional dynamic library. Motion pattern detection can be completed by simply training the output-side reader. The reader periodically calculates the weighted sum of the neuronal membrane potential and generates a neuromorphic state value through a Sigmoid transform. The rapid membrane potential transition enables the liquid state machine to generate a pulse peak approximately ten milliseconds before the inflection point of the motion, meeting the requirements of real-time keyframe prediction.

[0077] To incorporate high-level semantics into keyframe determination, this paper introduces the similarity between the skeleton latent tensor and the action text embedding. The action text embedding is derived from a large-scale Chinese language model through distillation, and its dimensionality matches the skeleton latent tensor mapping vector. A high cosine similarity indicates a high match between the action clip and the preset label. When the similarity falls below a threshold, the system reduces the saliency score of the current frame to avoid misclassifying non-target actions as keyframes.

[0078] The feedback loop consists of two paths. First, the difference between the birth threshold and the death threshold is read for the highest entry in the persistent image at the center of the window to obtain the target coherence value. If this value is lower than the historical average, it indicates that the motion cycle has not yet completed. The system will accordingly increase the number of sampling steps of the voxel network in the relevant time period to guide the replenishment of details. Second, the pulse peak times generated by the liquid state machine indicate the moments when the motion suddenly changes. The system writes these time points into the voxel network ray scheduling table, increasing the rendering density of the moving edges, thereby strengthening the skeletal pose estimation in the next round of iterative optimization.

[0079] For example, on a hurdle data set collected at 240 frames per second, compared to a baseline system that does not use topological path features, the fusion of the persistent image and path signature of the present invention enables the liquid state machine pulse peak to arrive at the hurdle jump key frame 11 milliseconds ahead of time on average, and the recall rate increases from 88% to 94%. In the 100-meter sprint data, compared with traditional recurrent networks that only use two-dimensional posture sequences, the topological signature of the present invention more accurately captures the repeatability of acceleration phase movements, and the text similarity flow effectively suppresses interference from non-running movements, reducing the keyframe false detection rate by 30%. Experiments also show that after splicing the low-rank skeletal latent tensor with topological features, the semantic recognition accuracy is 6% higher than that of using rotation vectors alone.

[0080] At the engineering implementation level, persistent image conversion and path signature calculation can be completed in batches and in parallel on the GPU; the pulse neuron simulation adopts a time-separated mode, and the inference delay of a 120-frame window does not exceed 2 milliseconds, fully meeting the coach's real-time action playback and indexing needs.

[0081] Preferably, the topological path feature extracts persistent homology features by performing Vietoris–Rips filtering on the joint rotation sequence, and combines it with the third-order path signature feature to form a joint representation.

[0082] The proposed topological path feature aims to simultaneously leverage the "shape structure" and "temporal order" of joint rotation sequences, providing a discriminative and noise-robust mid-level representation for track and field motion recognition. The shape structure is described by persistent coherence, and the temporal order is described by path signatures. The combination of these overcomes their respective limitations and is robust under high-speed motion, partial occlusion, and dramatic lighting changes.

[0083] In principle, persistent coherence relies on Vietoris–Rips filtering to obtain topological “hole” information. The joint rotation vector sequence of each frame can be used to splice the Lie algebra representation of 22 joints in each frame into a 66-dimensional Euclidean vector to form a point set Vietoris–Rips filter with distance threshold Construct a simple complex, record the 1-dimensional connected ring with To avoid excessive memory usage, the present invention uses a discrete threshold sequence and performs a sparse histogram mapping on the birth-death pairs to obtain a persistent image. Each pixel stores the number of ring occurrences within a threshold interval. The persistent image is rotationally and translationally invariant, making it easier to align similar global shapes, such as those in the air and on the ground, in hurdle jumping, than using direct coordinate differences.

[0084] Time-order dependent path signature capture. After piecewise linear interpolation of the same window trajectory, the signature vector of the third order truncation is calculated. The geometric meaning of the signature can be viewed as the coordinates of the trajectory on a free Lie algebra. The third order already contains velocity and acceleration information, which is sufficient to depict the beat-to-beat rhythm of track and field movements without exploding in dimensionality. To ensure consistent scale, the present invention first normalizes the rotation vectors to their maximum norm before performing the signature integration, minimizing the impact of physical differences among athletes on the results.

[0085] In practice, the persistent image and path signature are reduced to the same length using linear transformation at the output and then concatenated into topological path features. For online reasoning, the persistent coherence part uses a parallel incremental algorithm: only the points that enter the window are updated with a simple complex, avoiding recalculation of the full window distance matrix; the path signature uses a pipelined iterative integration kernel, accumulated with 4-byte floating point numbers, and can run 256 trajectories in parallel at one time on the graphics card. Enter the sinusoidal time encoding module, map each dimension to the pulse frequency interval, and then drive the liquid state machine. The neuromorphic state flow generated by the liquid state machine fluctuates over time. When the persistent coherent spectrum or signature amplitude in changes significantly, a spike appears in the state stream, providing a feedforward signal for key frame capture.

[0086] To align topological path features with the semantic layer, the present invention uses the skeleton latent tensor and Chinese action text embeddings to calculate cosine similarity and apply a pairing loss during the training phase. If the window label is "hurdle jump," the similarity is optimized to be high; if the label is "sprint swing," the similarity is optimized to be low. This supervision forces the persistent image and path signature to be close to the corresponding semantic center in the latent space, allowing the subsequent saliency curve to distinguish action segments without manual thresholding.

[0087] Example: On a 100-meter sprint dataset running at 240 frames per second, using a sliding window length of 40 frames and a step size of 10 frames, the accuracy of action phase recognition using only persistent images was 85%; using only path signatures, the accuracy was 88%; and when combined with topological path features, the accuracy increased to 93%.

[0088] In terms of hardware efficiency, the concatenation of a 32-dimensional persistent image and a 64-dimensional path signature results in a total length of 96. Pulse coding results in a maximum frequency of 250 Hz, a 1000-neuron liquid state machine with a simulation step size of 0.1 millisecond, and an inference latency of 0.8 milliseconds per window, meeting real-time requirements. Combined with skeletal latent tensor differentiation, this method significantly reduces keyframe false triggering and maintains stability under outdoor backlighting and mirror interference in rainy conditions.

[0089] Preferably, the liquid state machine is composed of multiple layers of spiking neurons, the synaptic delays are randomly initialized within a preset range, and the read layer weights are periodically updated to output a neuromorphic state stream.

[0090] Liquid state machines belong to the reserve computing paradigm. They consist of a large number of interconnected spiking neurons forming a high-dimensional dynamic "liquid state," which is then projected into the task output space via a read layer consisting solely of linear transformations. This paper applies liquid state machines to parse topological path feature sequences in real time, outputting a neuromorphic state stream that serves as a keyframe prior and model feedback signal. Compared to recurrent neural networks, liquid state machines train weights only in the read layer, leaving internal connections fixed. This makes them easier to deploy on mobile graphics cards and even dedicated neuromorphic chips, while offering sub-millisecond inference latency.

[0091] The liquid state machine of the present invention adopts a three-layer structure: input coding layer, pulse neuron liquid layer and linear reading layer. The input coding layer maps the topological path feature vector into a pulse stream. The present invention uses sinusoidal time coding to convert the value into instantaneous frequency . and is a fixed scaling factor. The encoding layer generates pulse events according to a Poisson distribution within a fixed simulation step size, where high frequencies represent large eigenvalues ​​and low frequencies represent small eigenvalues. This allows different modal features to drive the liquid layer using a unified event format.

[0092] The liquid layer contains 1000 Izhikevich neurons with a connection probability of 0.2. The Izhikevich model combines biological realism and computational efficiency. The dynamics of a single neuron is given by:

[0093] in represents the membrane potential, Represents the recovery variable, is the input current, and Controls the neuron's rebound speed and sensitivity. Above 30 millivolts, the neuron generates a pulse and performs a reset operation . Common parameter sets It can simulate fast spikes and moderate adaptive release, which is particularly effective in capturing acceleration mutations in track and field movements. The present invention randomly samples synaptic delay values ​​according to a uniform distribution during the initialization phase. The random delay allows the liquid layer to generate distributed echoes for inputs of different time scales, allowing the subsequent read layer to obtain rich high-order temporal features in a single linear mapping. The read layer exists in the form of a linear discriminator, and its input is the membrane potential of all neurons in the liquid layer. High-dimensional vector composed of , outputs a scalar neuromorphic state value:

[0094] in is the Sigmoid function, is the weight vector, The read layer is updated every 5 frames, and the minimum mean square error is used to optimize the output state value to have the minimum error with the binary labels of the key frames manually annotated during the training phase. Since the liquid state is fixed, training only involves one matrix multiplication and gradient descent, and the convergence speed is two orders of magnitude faster than the deep recurrent network. During the inference phase, the state value Sudden changes in input features manifest as spikes. At key points in the action chain, such as takeoff, flight, and finishing, the persistent coherence and path signatures undergo significant changes, driving a sudden increase in the encoding layer's pulse frequency, ultimately forming easily detectable spikes at the liquid output.

[0095] The liquid state machine has two functions in the framework of the present invention. First, the real-time output neuromorphic state stream is directly fed into the hybrid saliency curve. The second saliency term is The absolute value of the difference. Because the pulse peak has a fixed offset from the actual action transition, this term can provide alternative information when the visual change term is suppressed due to occlusion. Second, the pulse peak time is fed back to the voxel rendering schedule, increasing the ray density of the light field network in adjacent frames and achieving local refinement of key frames. The feedback path ensures modal coordination: when the skeletal rotation information detects a potential keyframe, the light field immediately collects more pixel evidence to confirm or refute the judgment, avoiding misjudgment caused by isolated signals.

[0096] The sprint example demonstrates the advance prediction capability of the liquid state machine at a high frame rate. A sprint clip recorded at 240 frames per second was used, with a sliding window of 40 frames. The topological path features were encoded into pulses to drive the liquid layer. The neuromorphic state flow spikes were on average 9 milliseconds ahead of the start key frame and 11 milliseconds ahead of the finish line key frame. When the threshold of the mixed significance curve was set to the mean value minus 0.6 standard deviations, the key frame recall rate was 94% and the false alarm rate was 5%. In the hurdle example, due to the occlusion of the lower limbs by the hurdle, the visual change item could not form a clear minimum in the first 3 frames, but the liquid state machine still output a clear spike at the moment of stepping over, allowing the system to capture the correct key frame. Compared with the control group without the liquid state machine, the recall rate increased by 9%.

[0097] Computational performance tests show that on an RTX 4060 mobile graphics card, a 1,000-neuron liquid network simulation with a 0.1 millisecond step can process 500 frames of pulse input per second, with a single-window latency of 0.7 milliseconds. The read layer update overhead is negligible, with little impact on the overall real-time pipeline. Memory usage primarily comes from the liquid neuron state array, approximately 2.5 megabytes. If deployed on an energy-efficient neuromorphic chip, such as Intel Loihi, the same scale could operate at 0.1 watts, potentially enabling the development of mobile and portable track and field training analysis terminals.

[0098] Preferably, the semantic confidence flow is generated by comparing the cosine similarity between the skeleton latent tensor mapping vector and the Chinese action text embedding vector, and selecting the text label with the highest similarity as the current action label.

[0099] The semantic confidence flow is responsible for mapping high-dimensional skeletal motion patterns to natural language action labels, so that the system has both structured motion discrimination and semantic readability during the key frame capture process. The skeletal latent tensor is a low-rank vector of fixed length obtained after matrix product state decomposition, which has compressed all joint rotation information in the window. If the latent tensor and text embedding are directly compared using Euclidean distance, it is difficult to align the scale. The present invention first introduces a two-layer perceptron to map the skeletal latent tensor to the same dimensional space as the Chinese action text embedding, recorded as vector The Chinese action text embedding uses a transformer model pre-trained on general corpus and additionally distilled on sports instructions to obtain a static vector set of action phrases such as "start", "accelerate", "jump over the hurdle", and "cross the finish line". The real-time stage uses the window center frame time as the index to compare the cosine similarity of the two vector sets:

[0100] In the formula is the skeleton potential tensor mapping vector, For the Chinese action text embedding vector, is the Euclidean norm. The maximum value is used as the confidence of the current window and the corresponding text label is output to form a semantic confidence stream.

[0101] The core principle of this flow is that cosine similarity is insensitive to vector length and only measures directional consistency. During the training phase, the vector direction after mapping the skeletal latent tensor is compressed by a contrast loss to minimize the angle with the correct text vector. Therefore, similar movements automatically converge during inference, maintaining a high degree of matching even with different athlete shapes, clothing, and shooting angles. Training utilizes multiple positive and one negative sampling: each window's latent tensor is composed of its true text label vector as a positive sample and three randomly selected other label vectors as negative samples. The goal is to maximize the cosine of the positive sample and minimize the cosine of the negative sample. Because the latent tensor has suppressed high-frequency jitter, the training process converges approximately twice as fast as direct comparison of the original rotated vectors.

[0102] The semantic confidence stream plays two roles within the overall framework. First, it directly feeds into the third component of the hybrid saliency curve. A low confidence score indicates that the action label in the current window is not within the set to be detected. This lowers the local value of the curve and prevents keyframes from being accidentally triggered by non-target actions. Second, the text labels provide a concise semantic index for the backend analysis system, enabling coaches to quickly search through a large number of keyframes by "start" or "hurdle."

[0103] Example: On a segmented, annotated 100-meter sprint dataset, with a window length of 40 frames and a sliding frame length of 10 frames, the false detection rate during the start and acceleration phases was 12% without semantic flow. After adding semantic flow, low-confidence frames were filtered out using a threshold of 0.6, reducing the false detection rate to 4% and improving the overall F1 score by 5%.

[0104] In terms of computing power usage, the skeletal latent tensor mapping uses a two-layer fully connected network with a dimension of 128, achieving inference latency of less than 0.05 milliseconds. The text vector set is loaded once at session startup, occupying less than 1 megabyte of resident video memory. Cosine calculations scale linearly with the number of tags. The proposed method defaults to eight common track and field action phrases, and even with 32 tags, the real-time frame rate can still maintain over 200 frames per second. If deployed on edge devices, the text vectors can be quantized to 8-bit integers, reducing latency by 25%.

[0105] To enhance cross-domain scalability, the present invention also provides an online fine-tuning interface. If users add new labels such as "Stop at the Inspection Station" or "Stretch Before Exercise," they only need to collect at least 30 windows of latent tensors and provide the corresponding text. This allows the mapping network to be updated using incremental contrastive loss, without modifying the existing skeleton and voxel models. Experiments have shown that after 5 minutes of online fine-tuning, the recognition accuracy of new labels can be increased from 0 to 85%, without introducing catastrophic forgetting.

[0106] The description of the semantic confidence flow reinforces the innovation of this invention from three aspects: first, the unified vector space of the skeleton latent tensor and Chinese text establishes a bidirectional mapping between action state and language description; second, cosine similarity provides a lightweight computational path, enabling low-power devices to obtain high-level semantics in real time; third, through the coupling of saliency curves and feedback mechanisms with visual, topological, and neuromorphic components, the advantages of the system's multimodal closed-loop design are fully reflected.

[0107] Based on the amplitude of the voxel light field temporal derivative, neuromorphic state difference, semantic confidence complementarity and skeletal latent tensor difference, a hybrid saliency curve is constructed with preset and online adjustable weights. The local minimum of the hybrid saliency curve is selected and mapped to a set of keyframe indices according to the preset sampling rate.

[0108] The hybrid saliency curve is the core link of the present invention to fuse multimodal information into a single keyframe metric. The four components of the curve are the amplitude of the time derivative of the voxel light field, , neuromorphic state difference , semantic confidence complementary value Difference with skeletal latent tensor They complement each other in terms of physical meaning, noise characteristics and time resolution, and can jointly characterize the visual mutation, dynamic mutation, semantic confidence and whole-body morphological changes of track and field movements. In order to prevent any single path from failing, the present invention uses the weight vector The weighted sum of the four items constitutes a significance scalar:

[0109] In the formula The voxel light field is obtained by performing central difference analysis on the brightness changes between frames in the rendering pipeline, which can quickly capture sudden changes such as dust at the landing point and stretching highlights. It is the difference between adjacent frames of the neuromorphic state stream output by the liquid state machine, reflecting the action transition in advance; Take the complementary value of the semantic confidence flow and automatically reduce the saliency when the window action is irrelevant to the monitoring label; Taking the Euclidean distance of the skeleton potential tensor in adjacent windows and focusing on the slow changes at the whole body posture level can make up for the visual blind spots caused by upper limb occlusion.

[0110] The weight vector is initially obtained by offline grid search and then fine-tuned by policy gradient in the online stage. The system calculates the external evaluation score every 500 frames. , the evaluation score can be set as the F1 index of the manually labeled key frames or the business end click rate. promote , the gradient direction is maintained; if If it decreases, the weight is adjusted according to the reverse gradient:

[0111] in For is the Dirichlet distribution with parameters, is the learning rate. This is a sliding average baseline used to reduce variance. After the update, it is normalized again to ensure that the sum is 1. Since only four scalar parameters are adjusted at a time, the online update time is less than 0.1 millisecond, with no impact on real-time performance.

[0112] There are multiple local minima in the significance curve. To avoid the high computational cost brought by global optimization, the present invention adopts a two-stage search: a rough scan is performed on the GPU at a step size of 2 frames per second, the first-order difference sign is calculated and the interval where the curve crosses the average value is located; then a high-resolution curve is reconstructed using a cubic B-spline in each interval, and the precise minimum is obtained using the quasi-Newton method. This method controls the search time within 20 milliseconds in the hurdle action, while ensuring that the error with the global scan is less than 2 frames. The minimum time is obtained Afterwards, Map it to a frame number, where The sampling frame rate is 240. All sequence numbers constitute a key frame index set .

[0113] The multi-source structure of the hybrid saliency curve significantly improves the robustness. Experiments show that in outdoor backlighting When the gradient weakens, and It can still give obvious pulses at the turning point of the action; in the occlusion scene, Automatically downscaling saliency suppresses false positives. Using the voxel light field component alone, the keyframe recall rate was only 82%. Adding the neuromorphic component increased this to 89%. After integrating all four components and performing online weight adjustments, the recall rate reached 94%, with the false positive rate kept below 4%. In cross-scene transfer testing, the shooting conditions on a soft-surface training gym and a plastic track differed significantly, and the system only needed a 3-minute online update to find stable weights.

[0114] In terms of practical effects, coaches can quickly locate key stages such as "starting" and "taking off" through semantic tag filters; mixed saliency remains smooth at high frame rates, facilitating real-time pre-fetching by video players. When deployed on mobile devices, calculating the saliency curve only requires reading four cached scalars, with a computational complexity of , it can reach 300 frames per second on the Snapdragon 8 series processor.

[0115] Preferably, the weight coefficients of the mixed saliency curve are adjusted online by a policy gradient algorithm and normalized after each adjustment to keep the sum of all weight coefficients constant.

[0116] The four components of the hybrid saliency curve come from the voxel light field, neuromorphic state, semantic confidence, and skeletal latent tensor, respectively. Their contribution to keyframes varies across scenarios. For example, the temporal derivative amplitude of the voxel light field is more reliable under strong outdoor lighting conditions, while the semantic confidence flow contributes more to filtering false detections under complex indoor backgrounds. To enable the system to automatically adapt to scene changes, the present invention uses a policy gradient algorithm to dynamically adjust the weight coefficients during operation and normalizes the weight vector after each update to maintain a constant sum of the weight coefficients at 1.

[0117] In terms of principle, the policy gradient algorithm is derived from reinforcement learning and updates the policy parameters by maximizing the expected return. The hybrid saliency curve plays a strategic role in key frame capture, and the weight vector is the parameter to be learned, and the reward is defined as the external evaluation score The external evaluation score can be the F1 value between the manually annotated keyframe and the system output keyframe, or the manual confirmation rate of the downstream editing module. The present invention regards the weight as the expected value of the parameterized probability distribution and selects the Dirichlet distribution. .set up , is the temperature parameter. The smaller the temperature, the sharper the distribution. A weight update can be written as:

[0118] In the formula is the learning rate, is a sliding window baseline used to reduce variance; Indicates is the Dirichlet probability density of the parameter, is the logarithmic gradient. Since the logarithmic gradient of Dirichlet can be written by the digamma function with only constant complexity, the update calculation amount is extremely small. After the update, normalization is performed:

[0119] Ensure that the sum of the weights remains 1 while avoiding numerical drift.

[0120] Implementation details, external evaluation scores Evaluate at the window level. After the system processes 500 frames, it compares the current keyframe index set with the manual annotation or the historical best set to generate If there is no labeling at the beginning of training, unsupervised indicators can be used, such as the harmonic mean of the saliency valley depth and the inter-frame distance, as a temporary reward signal. The temperature parameter is initially set to 0.05, so that each sample is slightly disturbed in the nearby space. Learning rate Set to 5×10 -4To prevent overfitting, if the return decreases after 5 consecutive updates, the weights are frozen for 2000 frames before continuing exploration.

[0121] The online adaptive process consists of four steps: weight sampling, saliency curve calculation, keyframe capture and reward evaluation, and weight update. At the beginning of each 500-frame batch, a new weight vector is sampled from the Dirichlet distribution and used for the mixed saliency curve of all subsequent frames. The system calculates the curve based on the weights, searches for local minima, and maps them to the keyframe index set. At the end of the batch, the system calculates the keyframe index based on manual annotation or rules. . Then update according to the above formula The whole process is completely online and there is no need to interrupt the reasoning.

[0122] In hardware deployment, the weight update overhead is primarily due to the digamma function calculation. However, because the Dirichlet function is only four-dimensional, the update actually involves only four digamma calls, taking less than 10 microseconds, which is negligible. Weight storage occupies 32 bytes, making it suitable for persistent operation on mobile devices.

[0123] In terms of effect, the present invention performs online training on a 100-meter sprint dataset at 240 frames per second. The initial offline grid search weight is , F1 index 0.89. After about 10 rounds of online updates totaling 5000 frames, the weights were adjusted to , F1 increased to 0.93. When the scene was transferred to the indoor training hall, the contribution of the voxel light field term decreased due to more uniform illumination. The online learning adjusted the weight to , F1 recovers to 0.92. The hurdle dataset shows a similar trend. When frequent occlusions make the visual gradient unreliable, the weight of the neuromorphic term is increased to around 0.4, and the false alarm rate is reduced to 4%.

[0124] To verify weight stability, the present invention injected random flashes into a noisy simulated environment, causing false peaks in the amplitude of the temporal derivative of the voxel light field. Using a fixed weight scheme, the false detection rate was 9%. However, using online weight adjustment reduced the false detection rate to 3%. This experiment demonstrates that policy gradients can rapidly reduce the weight of the disturbed component, maintaining overall performance.

[0125] The present invention's weight updates are tightly integrated with back-end editing interactions. If a coach frequently skips certain system-captured keyframes in the playback interface, the confidence labels for these frames are marked as negative samples, resulting in a decrease in reward and automatically reducing the corresponding component weights. Conversely, if the coach marks missed frames, the system penalizes them in subsequent windows, encouraging the exploration of new weight combinations. This human-machine collaborative learning ensures that the system becomes increasingly accurate with use, meeting the requirements of practical training processes.

[0126] Preferably, the key frame time points are mapped to key frame indices according to the frame rate of the image acquisition sequence, and corresponding image frames are extracted from the visible light infrared image sequence based on the key frame indices and output as key frames.

[0127] After the hybrid saliency curve undergoes a two-stage minimum search, the system obtains a set of high-precision timestamps , each timestamp corresponds to a potential key moment in the track and field action sequence. In order to convert these continuous time values ​​into directly indexable video frames, the present invention adopts the frame rate mapping principle: the frame rate is calibrated by the image acquisition sequence. As the scaling factor, multiply the timestamp by the frame rate and round down to get an integer frame number:

[0128] in represents the key frame index, During system initialization, the hardware clock is written into the configuration file, with a typical frame rate of 240 frames per second. For the same timestamp, four modal data channels are generated. This invention selects only visible and infrared images as keyframe outputs for three reasons: first, the visible light channel has the highest resolution and richest detail; second, the infrared channel compensates for structural shadows in backlit scenes; and third, polarization maps and event streams are primarily used in the computational phase, while the user experience focuses more on intuitive visuals.

[0129] Two engineering problems need to be solved during the mapping process: temporal quantization error and frame loss tolerance. The quantization error comes from rounding operations, and the theoretical maximum value is less than 1 frame. Considering that the displacement of a single frame in a running action is about 4 mm at 240 frames per second, the error is almost negligible for the decomposition of the action. If sub-frame accuracy is required, optical flow compensation resampling can be enabled on the player side, and static illustrations can be generated by interpolation, but this is usually not necessary in training feedback or referee penalty scenarios. Frame loss tolerance is for occasional frame drops in the camera: the system maintains a ring buffer that records the mapping table of the last 3 seconds of frame numbers and images; if the frame number mapped to the key frame does not exist, the system traces back to the most recent successful frame and sets the missing frame flag for use by subsequent repair algorithms.

[0130] The keyframe extraction process copies the visible light image, infrared image, and synchronized metadata from the cache and writes them to a solid-state drive using the naming convention "YYYYMMDD_HHMMSS_nnnnn." The prefix records the UTC time, and the suffix records the frame number, facilitating alignment with external timing systems. Each keyframe also comes with a JSON file containing the following fields: saliency four-component value, action text label, skeletal latent tensor hash, and camera intrinsic parameter snapshot. The hash is used for backend deduplication; the intrinsic parameter snapshot ensures that the world coordinates at that time can be reversed after changing the lens focal length or moving the camera position.

[0131] The keyframe collection is not only used for coaching replays but also fed into the voxel light field incremental storage module. The voxel network freezes the gradients of voxel density and color weights within the keyframe index range to prevent subsequent online fine-tuning from overwriting already confirmed high-quality local morphology. Simultaneously, the skeletal poses are interpolated to precise timestamps and stored in the scene database to accelerate initial convergence for future iterations of the same scene. This database is multi-keyed by event, athlete ID, and date, enabling the location of similar historical movements within seconds, providing data support for skill comparison.

[0132] Example: Using 240 frames per second data from a 100-meter sprint, the system output an average of eight keyframes, encompassing the start, peak acceleration, and finish line. Manual evaluation showed that 95% of these keyframes fell within two frames before and after the turning point of the action; the remaining 5% was primarily due to camera shake, which caused a shift in the saliency curve.

[0133] Compared to traditional keyframe extraction methods based on optical flow or thresholds, the hybrid saliency approach proposed in this paper integrates visual, topological, neuromorphic, and semantic information. The frame rate mapping process converts continuous-time output into discrete indices to avoid repeated decoding. Compared to a baseline system that outputs 60 evenly sampled frames across the entire video, the proposed method reduces the total number of keyframes by 80%, while still covering all important action nodes. On mobile devices with limited storage bandwidth, this method saves an average of 120 megabytes per minute.

[0134] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included within the scope of the claims of the present application.

Claims

1. A method for automatically identifying and capturing track and field movements based on image data, characterized in that: The following steps are involved: Time synchronization is performed on the visible infrared image, polarization image, and event stream. The first and second Stokes components are extracted from the synchronized polarization image, and the polarization phase map is generated using the arc tangent values ​​of the two to form a unified time-base data packet. A unified time-base data packet is input into a polarized neural radiation field network to generate a voxel light field. A skeleton graph neural network based on Lie group message passing is used to obtain a joint rotation sequence. The joint rotation sequence is subjected to matrix product state decomposition to form a skeleton latent tensor. Through alternating optimization, the voxel light field, joint rotation sequence, and skeleton latent tensor are coupled into a continuous spatiotemporal scene model. Based on the joint rotation sequence, Vietoris–Rips filtering and third-order path signature are used to extract topological path features. The topological path features are encoded into a pulse sequence and input into the liquid state machine to obtain a neuromorphic state stream. The semantic confidence stream is generated by combining the similarity between the skeleton latent tensor and the action text embedding. The target coherence value and the pulse peak time in the neuromorphic state stream are fed back to the continuous spatiotemporal scene model. Based on the amplitude of the voxel light field temporal derivative, neuromorphic state difference, semantic confidence complementarity and skeletal latent tensor difference, a hybrid saliency curve is constructed with preset and online adjustable weights. The local minimum of the hybrid saliency curve is selected and mapped to a set of keyframe indices according to the preset sampling rate.

2. The method according to claim 1, characterized in that Time synchronization is achieved by maximizing the mutual information between the edge of the event stream and the edge of the visible infrared image, and a unified hardware pulse signal is used to calibrate the timestamps of the visible infrared image acquisition device, the polarization image acquisition device and the event stream acquisition device.

3. The method according to claim 1, characterized in that The polarization phase map is generated by comparing the amplitude relationship between the first Stokes component and the second Stokes component to determine the phase of each pixel, which is used to compensate for the polarization direction difference.

4. The method according to claim 1, wherein Each node of the skeletal graph neural network based on Lie group message passing corresponds to a preset human joint. Message passing realizes joint posture update by performing group logarithmic mapping and group exponential mapping on the rotation differences of adjacent nodes in the three-dimensional rotation group space.

5. The method according to claim 1, wherein Matrix product state decomposition performs low-rank reconstruction of joint rotation sequences within a sliding time window to generate skeleton latent tensors and compress temporal redundant information.

6. The method according to claim 1, characterized in that The topological path feature extracts persistent homology features by performing Vietoris–Rips filtering on the joint rotation sequence, and combines it with the third-order path signature feature to form a joint representation.

7. The method according to claim 1, characterized in that The liquid state machine consists of multiple layers of spiking neurons, with synaptic delays randomly initialized within a preset range, and periodically updates the read layer weights to output a neuromorphic state stream.

8. The method according to claim 1, characterized in that The semantic confidence flow is generated by comparing the cosine similarity between the skeleton latent tensor mapping vector and the Chinese action text embedding vector, and the text label with the highest similarity is selected as the current action label.

9. The method according to claim 1, characterized in that The weight coefficients of the mixed saliency curve are adjusted online using the policy gradient algorithm and normalized after each adjustment to keep the sum of all weight coefficients constant.

10. The method according to claim 1, characterized in that The key frame time points are mapped to key frame indices according to the frame rate of the image acquisition sequence, and the corresponding image frames are extracted from the visible light infrared image sequence based on the key frame indices as key frame outputs.

Citation Information

Patent Citations

  • Visual feature fusion semantic detection method and system in video description

    CN113269253A

  • Motion capture method, terminal equipment and storage medium

    CN114638921A

  • Physics-based image generation using one or more neural networks

    CN117252962A

  • Skeletal animation generation method and system

    CN119338953A

  • Pet dangerous behavior monitoring system based on optic neural network

    CN120108006A

Cited By

  • Automatic adjustment method for data-driven perfusion process based on machine learning

    CN121209282A

  • Human body motion data recovery method and device based on tensor representation

    CN122115506A