Athletic action automatic recognition and capture method based on image data

By combining polarized neural radiation field with Lie group skeletal graph network, the problems of posture drift and illumination noise in single-viewpoint track and field action recognition are solved, achieving stable keyframe capture and action recognition under single-viewpoint conditions, and improving the automated analysis capability of track and field training.

CN120673479BActive Publication Date: 2025-12-12SHENYANG SPORT UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510837727.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-12-12
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

Existing track and field motion recognition technologies suffer from false detections, missed detections, and drift under real training conditions, making it difficult to achieve stable and interpretable keyframe capture under single-view conditions, especially in high-speed motion and occlusion scenarios.

Method used

A closed-loop reconstruction of a 3D scene is achieved using polarized neural radiation field and Lie group skeletal graph network. Multi-source saliency curves are generated by combining topological path features and semantic matching, and keyframe indexing is realized through dynamic weighting of policy gradient.

Benefits of technology

It achieves synchronous reconstruction of dense 3D shapes and joint rotations under single-view conditions, overcomes pose drift caused by motion blur and occlusion, improves the robustness of global shape and temporal rhythm of motion, reduces the sensitivity of optical flow to lighting noise, and achieves high recall rate with cross-scene adaptation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673479B_ABST
    Figure CN120673479B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of sports biomechanics, and particularly relates to a track and field action automatic recognition and capture method based on image data, which comprises the following steps: synchronizing visible infrared images, polarization images and event streams according to mutual information criteria and generating a polarization phase image; inputting a polarization neural radiance field network and a Lie group message passing based skeleton graph neural network, and alternately optimizing to obtain a continuous voxel light field and a joint rotation sequence; performing Vietoris-Rips persistent homology and third-order path signature extraction topology path feature on the sequence, encoding into a pulse sequence driven liquid state machine, and combining with a skeleton latent tensor and an action text embedding to generate a semantic confidence stream; constructing a saliency curve with a voxel light field time derivative, a neuromorphic difference, a semantic complementary value and a skeleton difference, adjusting a weight online through a strategy gradient, mapping a curve minimum to a key frame index and outputting an image frame. The present application can still capture track and field action key frames in real time and accurately under complex lighting and shielding conditions.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of sports biomechanics, and particularly relates to a track and field action automatic recognition and capture method based on image data. BACKGROUND

[0002] Track and field training and competition increasingly rely on image data, and automatic recognition and key frame capture of actions can provide objective quantitative indicators for coaches, provide fast playback clips for referees and broadcast systems, and provide fine posture evaluation for sports rehabilitation agencies. However, most of the existing technologies use two-dimensional skeleton detection plus threshold triggering or simple optical flow analysis: ① two-dimensional key point tracking based on convolution network is only reliable in the front unobstructed scene, and is prone to mismatch when encountering motion blur caused by high-speed motion or barrier obstruction; ② the light flow mutation method is sensitive to mirror highlights, shadows and camera jitter, and has a high false detection rate; ③ the multi-view three-dimensional reconstruction scheme has high accuracy, but is expensive and complex to arrange, which does not meet the requirements of outdoor tracks and mobile shooting. The above reasons cause the existing system to coexist with false detection, missed detection and drift under real training conditions, and it is difficult to give stable and interpretable key frames. SUMMARY

[0003] In view of the many problems existing in the prior art, the present application provides a track and field action automatic recognition and capture method based on image data. The present application drives the polarization neural radiance field and the Lie group skeleton graph network closed loop reconstruction three-dimensional scene with a unified time base data package, and then generates a multi-source saliency curve through the topological path feature liquid state machine and semantic matching. Through strategy gradient dynamic weighting, the key frame index is output in real time.

[0004] A track and field action automatic recognition and capture method based on image data, comprising the following steps:

[0005] Time synchronization is performed on the visible infrared image, the polarization image and the event stream, the first Stokes component and the second Stokes component are extracted in the synchronized polarization image, and the polarization phase image is generated by taking the inverse tangent values of the two, to form a unified time base data package;

[0006] The unified time base data package is input into the polarization neural radiance field network to generate a voxel light field, a joint rotation sequence is obtained by using a skeleton graph neural network based on Lie group message passing, the joint rotation sequence is subjected to matrix product state decomposition to form a skeleton latent tensor, and the voxel light field, the joint rotation sequence and the skeleton latent tensor are coupled into a continuous space-time scene model through alternating optimization;

[0007] Topological path features are extracted based on Vietoris-Rips filtering and third-order path signatures of joint rotation sequences, and are encoded into a pulse sequence input liquid state machine to obtain a neuromorphic state stream, a semantic confidence stream is generated by combining the similarity of the skeletal latent tensor and the action text embedding, and the target coherence value and the pulse peak time in the neuromorphic state stream are fed back to the continuous spatio-temporal scene model.

[0008] A hybrid saliency curve is constructed by a preset and online adjustable weight according to the amplitude of the time derivative of the voxel light field, the neuromorphic state difference, the semantic confidence complementary value and the skeletal latent tensor difference, and a local minimum value of the hybrid saliency curve is selected and mapped into a key frame index set according to a preset sampling rate.

[0009] Preferably, time synchronization is completed by maximizing the mutual information of the event stream edge and the visible infrared image edge, and the timestamps of the visible infrared image acquisition device, the polarized image acquisition device and the event stream acquisition device are calibrated by using a unified hardware pulse signal.

[0010] Preferably, the generation of the polarization phase map determines the phase of each pixel by comparing the amplitude relationship between the first Stokes component and the second Stokes component, which is used to compensate for the difference in polarization direction.

[0011] Preferably, each node of the skeletal graph neural network based on Lie group message passing corresponds to a preset human joint, and the message passing is implemented by group logarithmic mapping and group exponential mapping of the rotation difference between adjacent nodes in a three-dimensional rotation group space to update the joint posture.

[0012] Preferably, the matrix product state decomposition performs low-rank reconstruction on the joint rotation sequence in a sliding time window to generate the skeletal latent tensor and compress the timing redundancy information.

[0013] Preferably, the topological path features extract persistent homology features by Vietoris-Rips filtering on the joint rotation sequence, and form a joint representation combined with the third-order path signature features.

[0014] Preferably, the liquid state machine is composed of multiple layers of pulse neurons, the synaptic delay is randomly initialized within a preset range, and the read layer weight is periodically updated to output the neuromorphic state stream.

[0015] Preferably, the semantic confidence stream is generated by comparing the cosine similarity of the skeletal latent tensor mapping vector and the Chinese action text embedding vector, and the text label with the highest similarity is selected as the current action label.

[0016] Preferably, the weight coefficients of the hybrid saliency curve are adjusted online by a policy gradient algorithm, and are normalized after each adjustment to keep the sum of all weight coefficients constant.

[0017] Preferably, the key frame time points are mapped to key frame indexes according to a frame rate of the image acquisition sequence, and corresponding image frames are extracted from the visible-infrared image sequence as key frame outputs based on the key frame indexes.

[0018] Compared with the prior art, the application has the advantages and beneficial effects that:

[0019] By coupling the polarized neural radiance field network and the Lie group skeletal graph neural network, synchronous reconstruction of dense three-dimensional shapes and legal joint rotations under single-view conditions is realized, and the attitude drift defect caused by motion blur and occlusion is overcome. By jointly extracting topological path features through Vietoris-Rips persistent homology and third-order path signatures, robust coding of global shape and timing rhythm of the motion is realized, and the problem of sensitivity of the optical flow method to light noise is solved. By fusing the liquid state machine and the semantic confidence flow, key frames are predicted milliseconds in advance and non-target motions are automatically filtered, filling the short board of high false detection rate of threshold triggering. By adjusting the saliency weight online through policy gradient, cross-scene adaptation is realized, and high recall rate can be maintained without recalibration. BRIEF DESCRIPTION OF DRAWINGS

[0020] Figure 1 is a flowchart of the method of the application;

[0021] Figure 2 is a schematic diagram of constructing a mixed saliency curve in the application. DETAILED DESCRIPTION

[0022] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the embodiments of the present disclosure.

[0023] As Figure 1 shown, a track and field motion automatic recognition and capture method based on image data includes the following steps:

[0024] Time synchronization is performed on the visible-infrared image, the polarization image and the event stream, the first Stokes component and the second Stokes component are extracted in the synchronized polarization image, and the polarized phase image is generated by taking the inverse tangent values of the two components, forming a unified time-base data packet;

[0025] Visible-infrared images, polarization images, and event streams originate from three different heterogeneous sensors, each with varying frame rates and time bases. Without establishing a unified time base first, temporal correspondences between different modes will drift, leading to skeleton misalignment or light field distortion in subsequent motion recognition and keyframe capture. This invention treats time synchronization as a cross-modal registration problem in principle, introducing a mutual information maximization criterion and supplementing it with hardware pulse calibration to achieve millisecond-level alignment.

[0026] Mutual information is a statistic that measures the joint uncertainty of two sets of random variables. For the edges of an event stream and visible-infrared images, the edge positions have a strict correspondence in the time domain. When the two data streams slide to the optimal alignment position, the joint entropy reaches its minimum, and the mutual information reaches its maximum. In implementation, the event stream is integrated into a grayscale image within a short time window, and then edge detection is performed with the visible-infrared image within the corresponding window. The mutual information is calculated, and a one-dimensional search is used to find the maximum point. This search is performed only within milliseconds, with relatively low computational cost. To avoid distortion of mutual information under sudden changes in lighting conditions, a common pulse signal is used at the hardware level to mark the timestamps of the three sensors. At algorithm startup, coarse alignment to the pulse boundaries is performed first, followed by fine-tuning of the mutual information. This dual mechanism ensures stable synchronization even under complex outdoor lighting conditions.

[0027] Polarization information can reveal the microscopic orientation of an object's surface, and is particularly sensitive to muscle texture, clothing wrinkles, and metallic reflections from equipment during track and field movements. This invention uses the polarization phase map as an additional channel for subsequent neural radiation field reconstruction, constraining the voxel normal direction of the radiation field and suppressing local reflection noise. The polarization sensor incorporates a quarter polarizer, outputting four intensity maps at different polarization angles. According to polarization measurement theory, the first Stokes component... Describing the difference in polarization intensity between the horizontal and vertical directions, the second Stokes component This describes the polarization intensity difference between the 45° and 135° directions. This invention uses only these two terms to calculate the phase, eliminating the need to introduce a third Stokes component and thus reducing additional noise channels.

[0028] Polarization phase calculation uses the arctangent relationship. The core formula is:

[0029]

[0030] in For pixels The polarization phase at that location; This indicates the intensity difference of the first Stokes component of the pixel; This represents the intensity difference of the second Stokes component of the pixel. The reason for using the arctangent instead of a simple ratio is that the arctangent can... Phase mapping of the range to the grayscale range enables a continuous periodic transition, avoiding [further issues]. Numerical divergence occurs when approaching 0. To improve robustness, the present application performs a median filter of 3x3 before formula calculation and First, a 3x3 median filter is performed to suppress random noise spikes, and then a lookup table method is used to accelerate the arctangent operation, ensuring that the additional delay of a single frame does not exceed 3 milliseconds in a real-time scenario.

[0031] After forming a unified time base data packet, the present application sends the visible-infrared image, polarization phase map and event stream into the polarization neural radiance field network. In this process, the phase information mainly plays two roles: first, by inputting the phase as an additional voxel feature into the network, it helps the network identify the true normal direction in small texture areas at the same time, reducing the stereo ambiguity caused by lens dynamic blur; second, when constructing the skeleton map neural network later, the phase map can provide polarization consistency constraints for clothing reflections around the joints, reducing the skeleton drift caused by clothing wrinkles. In the experiment, on the hurdle take-off test set, after introducing the phase map, the average reprojection error of the skeleton pose is reduced by about 12%, and the key frame capture and timing rate is increased by 8%.

[0032] Preferably, time synchronization is completed by maximizing the mutual information of the event stream edge and the visible-infrared image edge, and the time stamps of the visible-infrared image acquisition device, the polarization image acquisition device and the event stream acquisition device are calibrated using a unified hardware pulse signal.

[0033] In a multi-modal track and field action acquisition system, visible-infrared images, polarization images and event streams are output by three independent sensors, and the internal clocks of each are inconsistent by microseconds to milliseconds. If the unaligned data is directly sent to the subsequent neural radiance field reconstruction and skeleton pose estimation module, the same action frame will be divided into different time slices, causing voxel density drift, skeleton node misplacement and key frame misjudgment. The present application first uses a hardware layer common pulse synchronization signal to complete coarse calibration, and then uses mutual information maximization to complete fine calibration at the algorithm layer, to ensure that the three-way data is strictly aligned within a millisecond time window.

[0034] The hardware pulse synchronization signal periodically triggers the rising edge to the three sensors through a distribution clock module, and whenever the sensor detects the trigger edge, it records the local timestamp and marks the frame where the pulse is located as the reference frame. Since all trigger edges come from the same physical pin, the coarse time difference of the three sensors is determined by the line delay and internal cache delay, and the typical value does not exceed tens of microseconds, which can be considered as an initial alignment of the global reference.

[0035] After the coarse alignment, the remaining drift caused by frame read delay, encoding buffer delay and system bus delay still needs to be solved. The mutual information is selected as the fine adjustment index of cross-modal time alignment. The mutual information quantifies the joint entropy and conditional entropy of two groups of signals, and can measure the statistical dependence between the pseudo-synchronous signals without assuming a linear relationship. The event stream records the sign of brightness change in a very short exposure, which is essentially a high-frequency spatial edge trigger. The visible-infrared image contains spatial gradient information at the same time. When the two data are at the correct alignment position, the event stream activation points and the visible-infrared image edges are highly overlapped, the coupling degree of the joint probability distribution is maximum, and the mutual information reaches the maximum. In order to make the mutual information calculation more stable, the event stream is accumulated in a short time window to obtain a gray fitting, and then the Sobel edge detection is performed on the visible-infrared image. The length of the accumulation window is automatically derived from the pulse period, which not only ensures the stability of the statistics, but also avoids the motion blur caused by the too wide window.

[0036] The mutual information maximization is realized by one-dimensional search. Let the event accumulated edge map be a random variable , the visible-infrared edge map be a random variable , and the mutual information of the two be

[0037]

[0038] , where is the joint probability, , are the edge probabilities of the edge map gray scale respectively. The present application performs a linear scan in a limited range on the delay amount, and each delay amount corresponds to a pair of synchronization windows. The is estimated by fast histogram statistics. In order to reduce the logarithmic calculation overhead, the joint probability is multiplied by a fixed scaling factor before using the table method to approximate the logarithmic function. The position of the maximum mutual information is taken as the final time offset, and the time stamps of the polarized image, the visible-infrared image and the event stream are adjusted by the offset to form a unified time base data package.

[0039] The theoretical optimality of mutual information search is derived from the principle of maximum statistical interdependence between signals, rather than relying on the consistency of signal amplitude or frequency. For the high-speed swing and foot landing moment in track and field movements, the event stream edge often presents sparse and extreme activation. Compared with the traditional alignment method based on mean square deviation or cross-correlation, the mutual information is more robust to amplitude ratio and noise. Experiments show that when the athlete sprints or performs hurdle at a speed of more than 10 meters per second, the motion blur between adjacent frames is serious, and the cross-correlation peak value appears multiple local extreme values, while the mutual information curve still maintains the single peak feature, which can significantly reduce the probability of false matching.

[0040] To verify the effect, the embodiment selects the actual shooting data of the 100-meter run in the training ground, including 240 frames per second of visible-infrared images, 90 frames per second of polarization images, and 10 kilohertz of event streams. First, the three time stamps are roughly aligned by using hardware pulses, and it is measured that the mean square deviation of the initial time deviation is less than 0.1 millisecond. Then, the mutual information is used for fine search in each pulse interval, and the search step is set to 0.1 millisecond, and the search window is 4 milliseconds long. Finally, the residual time error mean is 0.3 milliseconds, and the standard deviation is 0.15 milliseconds. The aligned data is input into the subsequent network of the application, and the re-projection error is reduced by 15% compared with the case of using only hardware pulse alignment, and the key frame recall rate is increased by 7%. Another group of hurdle flying data is shot in a backlight environment, and the polarization noise increases due to light reflection, and the standard deviation of the error is maintained at 0.2 milliseconds after synchronization by the application, which proves that the algorithm layer fine calibration is not sensitive to light changes.

[0041] In order to further reduce the operation amount, the application introduces an adaptive step strategy in the search process: a larger step is used in the initial scanning to obtain the approximate peak position, and a small step is used for secondary scanning on both sides of the peak; and the parallel prefix is used to accelerate the histogram normalization when calculating the joint histogram. With the cooperation of the GPU pipeline processing, the online synchronization capability of more than 50 frames per second can be maintained on a game-level graphics card.

[0042] Preferably, the generation of the polarization phase map determines the phase of each pixel by comparing the amplitude relationship between the first Stokes component and the second Stokes component, which is used to compensate for the difference in polarization direction.

[0043] Polarization information is an important dimension for describing the vibration direction of the light field, which is parallel to intensity, wavelength and constitutes a complete radiation feature. In the track and field scene, the muscle fibers of the athletes, the fabric of the tight clothes and the mirror reflection on the surface of the metal equipment will change the polarization direction of the incident light, resulting in that the visible light intensity does not strictly correspond to the real geometric normal. Only relying on the intensity texture to construct the neural radiation field or the skeletal pose network, local depth layering errors and joint rotation drifts are prone to occur. The application introduces the polarization phase map as an additional input channel to compensate for the difference in polarization direction at the pixel level, which significantly improves the accuracy and steady-state robustness of three-dimensional reconstruction and key frame determination.

[0044] The quarter polarization camera outputs four intensity images shot at 0°, 45°, 90° and 135° directions at the same time by setting a micro-polarization plate array in front of the photosensitive chip. The linear polarization state can be described by a three-dimensional Stokes vector, and the first component represents the intensity difference between the horizontal and vertical directions, and the second component denotes the difference in intensity between the diagonal directions. The present application selects these two items to determine the phase, because the phase is only related to the vibration direction, and is not related to the degree of polarization. In order to maintain the dimensionality consistent with the density and color tensor in the subsequent network, the polarization phase is mapped to the continuous phase interval using the arctangent function. The core calculation expression is:

[0045]

[0046] wherein, is the polarization phase at pixel ; is the first stokes component; is the second stokes component. The first stokes component is obtained by image subtraction of 0° and 90° directions, and the second stokes component is obtained by image subtraction of 45° and 135° directions. Since the trigonometric function is numerically sensitive when the denominator is close to zero, the algorithm first performs bilateral filtering on and to suppress high-frequency noise, and then uses a lookup table method to approximate the arctangent to reduce hardware instruction delay.

[0047] The present application does not use the third stokes component, nor does it calculate the degree of polarization. On the one hand, the third component is close to zero under natural light, and its contribution to the phase is limited; on the other hand, reducing the operation channel can reduce the memory bandwidth and ensure the online inference speed. In addition, the phase map is only spliced with the color tensor at the first layer of the network, and does not participate in the back propagation of the density gradient of the light field voxel, avoiding amplifying errors when there are raindrops, dust and other non-structural noise.

[0048] The direct effect of the phase map is to provide the micro-surface orientation information of each pixel, allowing the neural radiance field to have explicit constraints when estimating the voxel normal. In a running scene, the calf muscles of an athlete swing rapidly with the step frequency, and traditional voxel reprojection based on color consistency is prone to produce stripe artifacts in mirror highlights. After adding the phase channel, the network can determine the true reflection direction using the phase jump position, thereby correctly estimating the shape of the calf curve. The indirect effect is reflected in the skeletal graph neural network. After taking the phase consistency weighted for the pixels in the joint neighborhood, the joint heat map will suppress the false peaks at the folds of the clothing lines, reducing the jitter of the joint position.

[0049] Embodiment: An athlete completes a hurdle action under outdoor backlight conditions. After using the phase map compensation of the present application, the Euclidean angular error of the joint rotation matrix corresponding to the highest point of the hurdle flight and the reference value of the laser motion capture is reduced from 6.8 degrees to 4.1 degrees, and the dynamic time warping distance of the skeletal latent tensor and the reference trajectory is reduced by 19%.

[0050] In order to fully utilize the phase information, the present application also adds a "phase consistency loss" to the phase map during the training stage. The phase values of the same voxel rendered in the previous two frames and Take the absolute difference and weight into the light field loss:

[0051]

[0052] where is the phase consistency weight. This loss encourages the network to keep local polarization consistent in consecutive frames, suppressing the reverse noise introduced by slight camera shaking. Training experiments show that when When choosing the same level of dimension as the main loss of the light field, the reprojection disparity is further reduced by about 6%, and the number of network convergence iterations is reduced by about 12%.

[0053] The polarization phase calculation of the present application does not require additional optical components on the hardware side, and can be completed only by software post-processing, which is convenient for direct deployment on existing multi-modal camera systems. Compared with the common multi-view stereo scheme, the present application introduces a normal constraint using phase in a single-view configuration, solving the problem that multiple cameras cannot be arranged due to space limitations in track training venues, thereby deploying in college track stadiums, outdoor tracks and other occasions at a lower cost. Practical tests show that a single fisheye lens plus a polarization camera can cover one to three lanes of a hurdle track, reducing the number of hardware by half without losing recognition accuracy.

[0054] As Figure 2 shown, the unified time base data packet is input into the polarization neural radiance field network to generate a voxel light field, a skeletal graph neural network based on Lie group message passing is used to obtain a joint rotation sequence, the joint rotation sequence is subjected to matrix product state decomposition to form a skeletal latent tensor, and the voxel light field, the joint rotation sequence and the skeletal latent tensor are coupled into a continuous space-time scene model through alternating optimization;

[0055] The unified time base data packet includes a visible light and infrared light combined image, a polarization phase map and an event stream grayscale map. In the pipeline of the present application, the data packet is sent into the polarization neural radiance field network in batches. The neural radiance field network takes voxel coordinates and sampling ray directions as inputs, outputs voxel density and color through a multilayer perceptron, and simultaneously splices polarization phase features in the first hidden layer, so that the network obtains additional directional constraints when solving the surface normal of each voxel. In order to avoid the periodicity of the polarization phase causing discontinuous gradients, the network encodes the phase using a sine and cosine dual channel at the input end, so that it remains smooth in the interval. Unlike traditional radiance fields, the voxel features of the present application not only include three-dimensional density and three-dimensional color, but also store polarization normal encoding, and a phase consistency loss is added to the rendering loss, so that the phase residuals of adjacent frames of voxels are minimized in the training stage, further improving the accuracy of voxel normal estimation.

[0056] The core integral of voxel rendering follows the volume rendering formula:

[0057]

[0058] wherein is the sampled light ray, denotes the voxel density, is the distance between adjacent sampling points, is the voxel color vector, is the residual transmittance before the current voxel. Density, color and phase share gradients in backpropagation, so any phase estimation error will be backpropagated together with color error, prompting the network to correct photometric and directional information simultaneously in iterations.

[0059] In the human motion modeling part, the present application adopts a skeletal graph neural network based on Lie group message passing. Lie group message passing refers to convolution operation in a three-dimensional rotation group space, which ensures that the network is still in the rotation group after updating and will not produce invalid poses. For joints and its adjacent joints , the relative rotation is defined as:

[0060]

[0061] wherein , is the rotation matrix. The Lie group logarithm mapping is used to convert into a vector in the Lie algebra. The graph neural network uses as the message, and after a multi-layer linear transformation and a nonlinear activation, the updated value is mapped back to the rotation group by Lie group exponential mapping to obtain a new joint pose. This process ensures that the message calculation is performed in Euclidean space, and the final result is still legal in the rotation space, avoiding the accumulation of drift in the iteration of the joint pose.

[0062] The joint rotation sequence output by the skeletal graph neural network grows rapidly over time, and the present application uses matrix product state decomposition to compress long sequences. The continuous 40-frame rotation matrix is unfolded along the time dimension into a high-order tensor , which is then decomposed into a chain tensor network:

[0063]

[0064] wherein is the core tensor of the frame, is the internal bonding index. By truncating the bonding dimension to 32, the main variation mode of the pose sequence can be preserved with a small amount of storage. The resulting one-dimensional latent vector, i.e., the skeletal latent tensor representation, is used for subsequent topological feature, semantic matching and hybrid saliency calculation.

[0065] In order to correct the voxel light field and the bone model, the application adopts an alternating optimization strategy in the training stage. First, fix the bone parameters, only update the voxel network, minimize the rendering error and phase consistency; then fix the voxel network, project the rendering result to the two-dimensional image space, calculate the bone re-projection error in the camera frame, and update the bone graph network weight. The two are alternately carried out, and after multiple iterations, a space-time scene model is formed: the voxel light field provides dense shape and texture, the bone sequence describes the moving skeleton, and the bone latent tensor compresses and transmits global timing information. Since the voxel network samples the key parts with high weight by means of the bone heat map during training, and the bone network re-estimates the Lie group message using the new voxel direction in each iteration, the two networks form a closed loop, so that the model still maintains stability in the presence of occlusion, mirror reflection and high-speed motion scenes.

[0066] Embodiment: Take the hurdling action in outdoor backlight conditions. After adding the polarization phase, the joint rotation error at the highest point of the hurdling flight is reduced from 6.8° to 4.1°, and the dynamic time warping distance of the bone latent tensor and the laser reference track is reduced by 19%.

[0067] After the matrix product state decomposition compresses the 40-frame rotation matrix to the key dimension of 32, the latent vector storage is about 4120 bytes, while the direct storage of the rotation matrix is more than 200000 bytes. In a single-chip microcomputer plus mobile graphics card system, the bone latent tensor can still be maintained in real time and interacted with the voxel network, proving that the application has engineering deployability.

[0068] Through the alternating optimization of the polarized neural radiation field network, the bone graph neural network based on Lie group message passing and the matrix product state decomposition, the application realizes the high-precision coupling of the voxel light field and the bone sequence in the continuous space-time domain, and significantly improves the stability and precision of the automatic recognition and key frame capture of track and field actions.

[0069] Preferably, each node of the bone graph neural network based on Lie group message passing corresponds to a preset human joint, and the message passing is implemented by group logarithmic mapping and group exponential mapping on the rotation difference between adjacent nodes in the three-dimensional rotation group space to update the joint posture.

[0070] In the track and field action recognition and key frame capture task, the core information of human posture is the rotation and displacement of each joint in three-dimensional space. Traditional joint rotation regression is usually expressed in Euler angles or quaternions, but Euler angles have gimbal lock, quaternions avoid singularity but need unit normalization, and both parameterizations cannot guarantee that the updated parameters still fall within the legal rotation set during neural network propagation. The application adopts a bone graph neural network based on Lie group message passing to directly model the rotation in the special orthogonal group space , which naturally satisfies the rotation closure, and simultaneously uses the graph structure to explicitly encode the bone topological constraint.

[0071] The skeleton graph neural network is composed of 22 nodes, each of which is fixed to correspond to an anatomical joint of a human body, including ankle, knee, hip, shoulder, elbow, wrist and thoracolumbar segment. The edges between the nodes are established according to the real bone connection relationship, for example, the tibia node is connected with the femur node, and the scapula node is connected with the humerus node. The network input is the initial rotation matrix sequence of each node in the continuous frame In a message passing cycle, for a node and its adjacent nodes , the network first calculates the relative rotation:

[0072]

[0073] Wherein the superscript represents the current graph convolution layer. The rotation difference is located in space, but message aggregation needs to be carried out in vector space, so is projected to Lie algebra by Lie group logarithm mapping , to obtain a vector:

[0074]

[0075] The geometric meaning of is the minimum rotation axis angle from the node to the node , which can be directly used for vector addition. The network carries out multi-head attention weighting on , splices the self-features of the node , and sends the updated vector to a multilayer perceptron to obtain an updated vector . In order to ensure that the updated result still falls in

[0076] , Lie group exponential mapping is used again:

[0077] Wherein is the Lie group exponential mapping, is realized in space by the Rodrigues formula. The present application sets 6 layers of graph convolution, each of which performs the above-mentioned “rotation difference-logarithm mapping-exponential mapping” cycle, and propagates the posture information layer by layer.

[0078] This Lie group message passing has three principle advantages. First, the group logarithm mapping first converts the nonlinear dispersion rotation difference into a linear space, so that the deep network can directly apply linear transformation and keep the gradient stable; second, the exponential mapping maps the updated quantity back to the group space instead of simple addition, avoiding the network output falling into an illegal rotation area; third, the rotation difference naturally has relative invariance, which can resist camera displacement and local occlusion.

[0079] To improve training efficiency, the network uses the two-dimensional joint nodes extracted from visible light and infrared light images as weak supervision in the initial stage, and uses joint reprojection error to guide convergence; when the network enters the stable stage, the surface normal consistency loss obtained by polarized neural radiance field rendering is added to make the three-dimensional rotation and the real voxel normal cooperate. Due to the accumulation of rotation messages in the graph structure, remote joints such as wrists can obtain global priors from torso nodes, which is particularly effective for joint drift during high-speed arm swinging.

[0080] In track and field applications, dozens of rapid arm swings or foot-ground contact actions can occur within one second. If an Euler angle network is used, the error will accumulate over time, causing the wrist rotation to drift more than 5 degrees; using the Lie group message passing skeletal graph neural network of the application, the drift is suppressed within 1.2 degrees. Specific embodiments: On a 240 frames per second dataset, taking 100m indoor sprint in a windless room as an example, 120 sprint clips of 10 athletes were recorded. Using a laser motion capture system as the true value, the average pose error of the baseline quaternion network at the end of a 5-second clip is 4.9 degrees; the error of the network of the application is 1.3 degrees, and the key frame recall rate is increased by 9%. In another set of hurdle data, the network still maintains the stability of the joint rotation difference in the stage where the barrier is severely occluded above the barrier, proving that the relative rotation message has the ability to compensate for missing frames.

[0081] In terms of hardware deployment, the core calculation of Lie group logarithm / exponential mapping is matrix diagonal block splitting and Rodrigues formula, which can simultaneously process all node pairs in thousands of threads in a GPU parallel environment. Compared with traditional joint regression, the network only increases about 8% of the inference time in a single batch, but significantly improves the pose stability, providing high-quality input for subsequent skeletal latent tensor compression and hybrid saliency curve calculation.

[0082] The alternating optimization strategy maps the rotation difference of the latest joint pose back to the voxel sampling weight, prompting the voxel network to increase the ray density in high motion areas such as arms and knees; after the voxel network is updated, the more accurate normal is output to provide new supervision for the skeletal network, realizing positive feedback between the two. In cross-scene experiments, it only takes 1,000 steps of fine-tuning of the voxel network and 400 steps of fine-tuning of the skeletal network to re-converge from an outdoor track to an indoor training hall, showing flexibility in migration.

[0083] Preferably, matrix product state decomposition performs low-rank reconstruction on the joint rotation sequence within a sliding time window to generate a skeletal latent tensor and compress the temporal redundancy information.

[0084] Matrix product state decomposition belongs to the category of tensor network, which was first used for low-rank approximation of quantum many-body systems. The present invention introduces this idea into the analysis of track and field movements, which is used to compress the joint rotation sequence that grows rapidly over time, and generate the skeletal latent tensor. The rotation information can be represented as a high-order tensor by stacking in order. If a time window of 40 frames is taken, a dense tensor with an order of 40 and a dimension of 22 is obtained, which requires a large amount of video memory to store directly, which is not conducive to deployment on mobile terminals or real-time systems. The matrix product state reconstructs the high-order tensor through the chain core tensor, only retains the main variation mode, significantly reduces the storage and calculation complexity, and at the same time maintains the ability to distinguish human dynamics.

[0085] The processing flow first maps the joint rotation matrix of each frame in the sliding window to the Euclidean vector. The mapping operation uses the logarithmic form of the Rodrigues formula to convert the three-dimensional rotation matrix into a three-dimensional Lie algebra rotation vector, which avoids the singularity of Euler angles and maintains linear additivity. By concatenating the rotation vectors of 22 joints, a frame-level feature with a length of 66 can be obtained. After collecting 40 consecutive frames with a sliding window, a second-order tensor with a shape of 66x40 can be constructed. In order to process under the matrix product state framework, the second-order tensor is converted to a high-order form in the time dimension , where represents the index of the frame.

[0086] The present invention uses a fixed bonding dimension to perform matrix product state decomposition on the high-order tensor. The core expression is:

[0087]

[0088] In the formula, is a high-order rotation tensor, is the core tensor corresponding to the frame, is the internal bonding index. By truncating the bonding dimension to 32, the main variation mode of the posture sequence can be retained with a small amount of storage, and the resulting one-dimensional latent vector is the skeletal latent tensor representation, which is used for subsequent topological feature, semantic matching and mixed saliency calculation.

[0089] In order to ensure real-time performance, the present invention uses block singular value decomposition on the graphics card to extract the core tensor, and the decomposition time of each sliding window is about 2.3 milliseconds, which is much lower than the frame interval of track and field movement collection. Comparative experiments show that under the same hardware environment, directly inputting the stacked rotation matrix into the semantic network will cause the inference video memory occupancy to exceed 300 megabytes, while using the matrix product state reduces the occupancy to about 15 megabytes. Further profile analysis shows that most of the video memory comes from the core tensor gradient buffer, and if the gradient is turned off during inference, the video memory occupancy is less than 5 megabytes, which can meet the inference requirements of embedded terminals.

[0090] The embodiment shows the method effect. On the 100-meter sprint cross-scene dataset, using a traditional long short-term memory network to process 40 frames of rotation vector sequences as a baseline, the semantic recognition accuracy is 85%, and the key frame recall rate is 83%. After replacing the matrix product state latent tensor, the accuracy is improved to 91% under the same network structure, and the key frame recall rate is improved to 90%. Analysis shows that the matrix product state latent tensor is not sensitive to high-frequency jitter of joints, but maintains high response to real action cycles, so it can separate noise and action inflection points on the mixed significance curve. In the hurdle dataset experiment, under the same window length and key binding dimension settings, the combination of the latent tensor and the topological persistent homology feature can predict the hurdle key frame one time window earlier, providing redundancy for real-time judge-assisted evaluation.

[0091] The application adopts alternating optimization during model training. When the voxel network is updated, the temporal pattern encoded in the latent tensor is used to control the ray sampling density, prolonging the integration time for action-intensive frames and shortening the integration time for static frames; when the skeleton graph network is updated, the integrated depth map is read from the voxel network to constrain joint rotation. After the alternating optimization converges, the skeleton latent tensor retains the main action mode of each window, which can be used as a lightweight index for the overall scene model, facilitating subsequent retrieval or incremental fine-tuning.

[0092] The matrix product state also has good interpretability. By viewing the eigenvectors of the kernel tensor, the main motion direction within the window can be intuitively discovered, for example, the first principal component of the hurdle action corresponds to the pelvis rotation during the flight phase, and the second principal component corresponds to the arm swing amplitude. Coaches and athletes can use the principal component analysis of the latent tensor to correct technical details and improve training efficiency.

[0093] Based on the Vietoris-Rips filter and the third-order path signature extraction of the joint rotation sequence, the topological path feature is encoded into a pulse sequence input liquid state machine to obtain a neuromorphic state stream. The similarity between the skeleton latent tensor and the action text embedding is combined to generate a semantic confidence stream. The target homology value and the pulse peak time in the neuromorphic state stream are fed back to the continuous spatiotemporal scene model.

[0094] The joint rotation sequence appears as a high-dimensional curve in Euclidean space, and its essential structure contains both continuous time order and overall shape topology. To capture both types of information, the rotation sequence is projected into the topological domain and the path algebra domain before being neuromorphically encoded. The real-time state stream obtained in this way is used to constrain the spatiotemporal scene model in reverse, forming a closed-loop analysis framework. The principles, implementations, and technical effects of each link are described below.

[0095] First, within each sliding window, a Vietoris-Rips filter is applied to the rotation vector sequence output by the Lie group message-passing skeleton graph network. This filtering method constructs a simple complex by progressively increasing the distance threshold, tracking the generation and disappearance of one-dimensional connected cycles. For a given set of rotation vectors within the window... When the distance threshold is When the distance between two points is less than Then it becomes connected, allowing for the construction of higher-order simplexes. With... Starting from zero, higher-order topological features gradually appear and eventually disappear. Connected loops that persist longer better reflect the overall rhythm of skeletal movement. This invention records the birth and death thresholds for each persistent entry and projects the birth and death coordinates into a persistent image using a two-dimensional Gaussian kernel. When a complete leg-raising or arm-swinging cycle occurs during movement, the corresponding entry shows a high response in the persistent image; in jittery or short-term occlusion scenarios, the corresponding entry disappears extremely quickly, thus naturally filtering noise.

[0096] Topological features only characterize the overall shape but cannot express the direction and speed of movement. To characterize the temporal sequence, this invention calculates a third-order path signature for the rotation vector trajectory within the same window. The path signature is a series of iterative integrals, which can be viewed as encoding the path into the coordinates of a free Lie algebra. The third-order truncation is sufficient to cover the velocity and acceleration information in track and field movements without excessively expanding the dimensionality. Since the rotation vector is normalized to a fixed scale, the signature values ​​are stable and easy for neural network post-processing.

[0097] Persistent images and third-order path signatures differ in dimension. This invention employs linear projection to concatenate them into a topological path feature vector of uniform length, and performs sinusoidal time encoding on each vector channel to generate pulse frequencies. The pulse frequencies drive a liquid state machine according to a Poisson distribution. The liquid state machine is a feedforward spiking neural network, with one thousand Izhikevich neurons randomly connected to form a high-dimensional dynamic library. Only the output-side reader needs to be trained to complete action pattern detection. The reader periodically calculates a weighted sum of the neuronal membrane potentials, which is then transformed by a sigmoid transform to generate neuromorphic state values. Rapid membrane potential transitions enable the liquid state machine to provide a pulse peak approximately ten milliseconds before the motion inflection point, meeting the requirements for real-time keyframe prediction.

[0098] To incorporate high-level semantics into keyframe determination, this invention introduces the similarity between the skeletal latent tensor and the action text embedding. The action text embedding is obtained by distillation from a large-scale Chinese language model, and its dimension is consistent with the skeletal latent tensor mapping vector. A high cosine similarity indicates a high degree of matching between the action segment and the preset label. When the similarity is below a threshold, the system reduces the saliency score of the current frame to avoid misclassifying non-target actions as keyframes.

[0099] The feedback link includes two paths. One is to read the difference between the birth threshold and the death threshold of the highest entry of the persistent image of the window center moment, to obtain the target coherence value; if the value is lower than the historical average, it means that the action cycle has not been completed, and the system will increase the sampling steps of the voxel network in the relevant time period, guiding the supplement of details. The second is that the pulse peak time table generated by the liquid state machine indicates the moment when the motion mutates, and the system writes these time points into the voxel network ray scheduling table to improve the rendering density of the motion edge, thereby strengthening the skeletal pose estimation in the next round of alternating optimization.

[0100] For example: on a 240-frame-per-second hurdling dataset, compared with the baseline system without using topological path features, the liquid state machine pulse peak of the persistent image and path signature fusion of the invention is 11 milliseconds earlier than the hurdling key frame, and the recall rate is increased from 88% to 94%. In the 100-meter sprint data, compared with the traditional recurrent network using only two-dimensional pose sequence, the topological signature of the invention captures the repeatability of the acceleration stage action more accurately, and the text similarity flow effectively suppresses the interference of non-running actions, reducing the key frame false detection rate by 30%. Experiments also show that the semantic recognition accuracy of the low-rank skeletal latent tensor spliced with the topological feature is improved by 6% compared with the rotation vector alone.

[0101] In the engineering implementation layer, both the persistent image conversion and the path signature calculation can be completed in batch parallel on the GPU; the pulse neuron simulation adopts time separation mode, and the inference delay of a 120-frame window is not more than 2 milliseconds, fully meeting the real-time action playback and indexing needs of coaches.

[0102] Preferably, the topological path feature is extracted by Vietoris-Rips filtering on the joint rotation sequence to obtain the persistent homology feature, and combined with the third-order path signature feature to form a joint representation.

[0103] The topological path feature is proposed to simultaneously utilize the "shape structure" and "time sequence" information of the joint rotation sequence, providing a discriminative and noise-robust middle-level representation for track and field action recognition. The shape structure is described by persistent homology, and the time sequence is described by path signature, which overcomes the limitations of each other and can remain stable under high-speed motion, local occlusion, and dramatic changes in light.

[0104] At the principle level, persistent homology relies on Vietoris-Rips filtering to obtain topological "hole" information. Given a continuous frame joint rotation vector sequence, the Lie algebra representation of 22 joints in each frame can be spliced into a 66-dimensional Euclidean vector to form a point set Vietoris-Rips filtering constructs a simplicial complex with a distance threshold to record 1-dimensional connected rings as Variable birth and death time. To avoid excessive memory, the invention selects a discrete threshold sequence and sparsely maps the birth-death pair to a persistent image , each pixel of which stores the number of occurrences of a certain pair of threshold intervals. The persistent image has rotational and translational invariance, and is more easily aligned than direct use of coordinate difference for repeated occurrences of similar global patterns such as take-off and landing in hurdle.

[0105] Time-ordered dependence path signature capture. After piecewise linear interpolation of the same window trajectory, the signature vector is calculated to the 3rd order truncation . The geometric meaning of the signature can be regarded as the coordinates of the trajectory on the free Lie algebra, and the 3rd order has included velocity and acceleration information, which is sufficient to describe the rhythm of one beat and one fall of track and field movements, while the dimension does not explode. To ensure scale consistency, the invention first normalizes the rotation vector according to the maximum norm, and then integrates the signature, so that the size difference of different athletes has the least impact on the results.

[0106] In implementation, the persistent image and the path signature are spliced into a topological path feature after being reduced to the same length using linear transformation at the output end . In order to infer online, the persistent cohomology part uses a parallel incremental algorithm: only the simplex complex of the points newly entering the window is updated, avoiding the recalculation of the full window distance matrix; the path signature uses a pipeline iterative integral kernel, which accumulates in 4-byte floating-point numbers, and can be parallelized for 256 trajectories at a time in a graphics card. The spliced enters the sinusoidal time encoding module, maps each dimension to the pulse frequency interval, and then drives the liquid state machine. The neuro-morphic state stream generated by the liquid state machine fluctuates over time, and when there is a significant change in the amplitude of the persistent cohomology spectrum or the signature, a spike appears in the state stream, providing a feedforward signal for key frame capture.

[0107] In order to align the topological path feature with the semantic layer, the invention uses the bone latent tensor and the Chinese action text embedding to calculate the cosine similarity and apply the pairing loss in the training stage. If the window label is "hurdle take-off", the optimization makes the similarity high; if the label is "sprint swing", the similarity is low. This supervision prompts the persistent image and the path signature to be close to the corresponding semantic center in the latent space, so that the subsequent saliency curve can distinguish the action paragraphs without manual thresholding.

[0108] Embodiment: On a 240-frame-per-second 100-meter sprint dataset, a sliding window length of 40 frames and a step size of 10 frames are used. Using only the persistent image, the action phase recognition accuracy is 85%; using only the path signature, the accuracy is 88%; after combining the topological path feature, the accuracy is improved to 93%.

[0109] In terms of hardware efficiency, the 32-dimensional persistent image and the 64-dimensional path signature are spliced into a total length of 96, the maximum frequency is 250 Hz after pulse coding, the simulation step of 1000 neurons of the liquid state machine is 0.1 ms, and the inference delay of each window is 0.8 ms, which meets the real-time requirement. In combination with the bone potential tensor difference, the application significantly reduces the key frame false triggering and remains stable under outdoor backlight, rainy day mirror interference.

[0110] Preferably, the liquid state machine is composed of multiple layers of pulse neurons, the synaptic delay is randomly initialized within a preset range, and the read layer weight is periodically updated to output the neuromorphic state stream.

[0111] The liquid state machine belongs to the reservoir computing paradigm, is composed of a large number of interconnected pulse neurons to form a high-dimensional dynamic 'liquid state', and projects the liquid state into a task output space through a read layer containing only linear transformation. The application uses the liquid state machine to analyze the topological path feature sequence in real time, outputs a neuromorphic state stream, and uses the neuromorphic state stream as a key frame priori and model feedback signal. Compared with the recurrent neural network, the liquid state machine only trains the weight of the read layer, and the internal connection remains fixed, so it is easier to deploy on a mobile graphics card or even a special neuromorphic chip, and has a sub-millisecond inference delay.

[0112] The liquid state machine of the application adopts a three-layer structure: an input encoding layer, a pulse neuron liquid layer and a linear read layer. The input encoding layer maps the topological path feature vector into a pulse stream. For the feature vector , the application uses sinusoidal time encoding to convert the numerical value into an instantaneous frequency . And are fixed scaling coefficients. The encoding layer generates pulse events according to the Poisson distribution within a fixed simulation step, and a high frequency indicates a large feature value, and a low frequency indicates a small feature value. In this way, different modal features can be driven in a unified event format to drive the liquid layer.

[0113] The liquid layer includes 1000 Izhikevich neurons with a connection probability of 0.2. The Izhikevich model combines biological authenticity and computational efficiency, and the single neuron dynamics is given by the following formula:

[0114]

[0115] Wherein represents the membrane potential, represents the recovery variable, is the input current, and control the rebound speed and sensitivity of the neuron. Whenever exceeds 30 millivolts, the neuron generates a pulse and performs a reset operation . The commonly used parameter set is The model can simulate fast spikes and moderate adaptation firing, which is particularly effective for capturing acceleration mutations in athletic movements. In the initialization phase, the synaptic delay values are randomly sampled according to a uniform distribution fall between 2 and 12 milliseconds, forming a time-expanding loop. The random delay allows the liquid layer to produce a distributed echo for inputs of different time scales, enabling the subsequent reading layer to obtain rich high-order temporal features in a single linear mapping. The reading layer exists in the form of a linear discriminator, with the input being the membrane potential of all neurons in the liquid layer The high-dimensional vector composed of The output scalar neuromorphic state value is:

[0116]

[0117] where is the Sigmoid function, is the weight vector, is the bias. The reading layer is updated every 5 frames, with the output state value being optimized to have the smallest error with respect to the binary label of the key frame annotated by humans in the training phase. Due to the internal fixation of the liquid, the training only involves one matrix multiplication and gradient descent, with a convergence speed two orders of magnitude faster than that of deep recurrent networks. In the inference phase, the state value shows a spike in response to a mutation in the input feature. When key nodes such as takeoff, flight, and sprint appear in the action chain, the persistent coherence and path signature undergo significant changes, driving the pulse frequency of the encoding layer to surge, and finally forming an easily detectable spike at the output end of the liquid.

[0118] The liquid state machine has two functions in the framework of the invention. First, the real-time output of the neuromorphic state stream directly enters the mixed saliency curve. The second item takes the absolute difference of the difference. Since the pulse peak has a fixed offset from the real action transition, this item can provide alternative information when the visual change item is suppressed due to occlusion. Second, the pulse peak time is fed back to the voxel rendering schedule table, increasing the ray density of the light field network in adjacent frames, and achieving local refinement of key frames. The feedback path ensures modal collaboration: when the skeletal rotation information detects a potential key frame, the light field immediately collects more pixel evidence to support or refute the judgment, avoiding false positives caused by isolated signals.

[0119] The sprint embodiment illustrates the advance prediction capability of the liquid state machine at high frame rate. Using a 240 frames per second recorded sprinting clip, a sliding window of 40 frames. After topological path feature to pulse encoding, the neuromorphic state stream spikes advance the key frame of starting 9 milliseconds and the key frame of crossing the line 11 milliseconds. When the mixed saliency curve threshold is set to the mean minus 0.6 standard deviation, the key frame recall rate is 94% and the false positive rate is 5%. In the hurdle embodiment, the visual change item cannot form a clear minimum value in the first 3 frames due to the barrier blocking the lower limbs, but the liquid state machine still outputs a clear spike at the moment of crossing the hurdle, enabling the system to capture the correct key frame. Compared with the control group without the liquid state machine, the recall rate is improved by 9%.

[0120] The calculation performance test shows that on an RTX4060 mobile graphics card, when the 1000 neuron liquid state network simulation step is 0.1 milliseconds, 500 frames of pulse input can be processed per second, and the single window delay is 0.7 milliseconds; the reading layer update overhead can be ignored, and has almost no effect on the overall real-time pipeline. The memory occupation mainly comes from the liquid neuron state array, about 2.5 megabytes. If deployed to an energy-saving neuromorphic chip such as Intel Loihi, the same size can run at 0.1 watt power consumption, providing a possibility for mobile and portable track and field training analysis terminals.

[0121] Preferably, the semantic confidence stream is generated by comparing the cosine similarity of the skeleton latent tensor mapping vector and the Chinese action text embedding vector, and the text label with the highest similarity is selected as the current action label.

[0122] The semantic confidence stream undertakes the task of mapping high-dimensional skeletal motion patterns to natural language action labels, enabling the system to have both structured motion discrimination and semantic readability during key frame capture. The skeleton latent tensor is a fixed-length low-rank vector obtained after matrix multiplication state decomposition, which has compressed all joint rotation information within the window. If the latent tensor and the text embedding are compared directly in Euclidean distance, it is difficult to align the scale. The invention first introduces a two-layer perception machine to map the skeleton latent tensor to the same dimension space as the Chinese action text embedding, denoted as vector The Chinese action text embedding uses a transformer model pre-trained on general corpus and distilled with additional sports domain instructions to obtain a static vector set of action phrases "start", "accelerate", "hurdle jump", "cross the line", etc. In the real-time stage, the cosine similarity of the two vector sets is compared with the window center frame time as the index:

[0123]

[0124] wherein is the skeleton latent tensor mapping vector, is the th Chinese action text embedding vector, is the Euclidean norm. Let The maximum value is taken as the confidence of the current window and the corresponding text label is output, constituting a semantic confidence stream.

[0125] The core principle of the stream is that the cosine similarity is not sensitive to the length of the vector, only measuring the consistency of the direction; the direction of the vector after the skeletal latent tensor mapping is compressed to the minimum angle with the correct text vector by the contrast loss in the training stage, so that in the inference stage, the same type of action will automatically converge, even if the shape of the athlete, the clothes and the shooting angle are different, the matching degree can be kept high. The training adopts multi-positive and negative sampling: each window latent tensor and its true text label vector form a positive sample, and three other label vectors are randomly selected to form a negative sample, the goal is to maximize the positive sample cosine and minimize the negative sample cosine. Since the latent tensor has suppressed high-frequency jitter, the convergence speed in the training process is about twice faster than that of the original rotation vector directly compared.

[0126] The semantic confidence stream plays two roles in the overall framework. First, it directly enters the third component of the mixed saliency curve, and a low confidence indicates that the action label of the current window is not in the set to be detected, which can reduce the local value of the curve and avoid false triggering of key frames by non-target actions. Second, the text label can provide a simple semantic index for the back-end analysis system, making it easy for coaches to quickly search for “starting” and “hurdle” in a large number of key frames.

[0127] Embodiment: On a segmented annotated 100m sprint dataset, the window length is 40 frames and the sliding step is 10 frames. Without the semantic stream, the false detection rate of the starting and accelerating stages is 12%. After adding the semantic stream, the confidence low frames are filtered out by a threshold of 0.6, the false detection rate is reduced to 4%, and the overall F1 score is improved by 5%.

[0128] In terms of computing power occupation, the skeletal latent tensor mapping uses a 2-layer fully connected network with a dimension of 128, and the inference delay is less than 0.05 milliseconds; the text vector set is loaded once at the start of the session, and the resident memory is less than 1 megabyte. The cosine calculation is linearly expanded to the number of labels, and the invention defaults to 8 commonly used track and field action phrases. Even if the number of labels is expanded to 32, the real-time frame rate can still be maintained above 200 frames per second. If deployed on an edge device, the text vector can be quantized to 8-bit integers, reducing the delay by 25%.

[0129] To enhance cross-domain expandability, the invention also provides an online fine-tuning interface. If the user adds labels such as “stop at the check-in station” and “stretch after the exercise”, they only need to collect at least 30 window latent tensors and provide the corresponding text, and then update the mapping network through incremental contrast loss without modifying the original skeletal and voxel models. Experiments show that after online fine-tuning for 5 minutes, the recognition accuracy of the new label can be improved from 0 to 85% without introducing catastrophic forgetting.

[0130] The description of the semantic confidence stream enhances the innovation of the application from three aspects: first, the skeleton latent tensor and the unified vector space of Chinese text establish a bidirectional mapping between action state and language description; second, the cosine similarity provides a lightweight calculation path, enabling low-power devices to obtain high-level semantics in real time; third, the coupling of the saliency curve, feedback mechanism, visual, topological, and neuromorphic components fully reflects the advantages of the system's multi-modal closed-loop design.

[0131] A hybrid saliency curve is constructed by the preset and online-adjustable weights of the voxel light field time derivative amplitude, neuromorphic state difference, semantic confidence complementary value, and skeleton latent tensor difference. The local minimum value of the hybrid saliency curve is selected and mapped to the key frame index set according to the preset sampling rate.

[0132] The hybrid saliency curve is the core link of the application in which multi-modal information is fused into a single key frame metric. The four components that constitute the curve are the voxel light field time derivative amplitude , neuromorphic state difference , semantic confidence complementary value , and skeleton latent tensor difference . They are complementary in physical meaning, noise characteristics, and time resolution, and can collectively represent visual mutations, kinetic mutations, semantic confidence, and overall morphological changes of track and field actions. To prevent failure of any single path, the application uses a weight vector to weight the sum of the four components to form a saliency scalar:

[0133]

[0134] In the formula , the central difference of inter-frame brightness changes of the voxel light field in the rendering pipeline can quickly capture dramatic changes such as dust and high light; is the difference between adjacent frames of the neuromorphic state stream output by the liquid state machine, which reflects the action transition in advance; takes the complementary value of the semantic confidence stream, which automatically reduces the saliency when the window action is unrelated to the monitoring label; takes the Euclidean distance of the skeleton latent tensor in adjacent windows, focusing on slow changes at the overall posture level, which can compensate for visual blind spots caused by upper limb occlusion.

[0135] The weight vector is initially obtained by offline grid search, and then fine-tuned in the online stage using policy gradient. The system calculates an external evaluation score every 500 frames, and the evaluation score can be set as the F1 indicator of the manually labeled key frames or the business side click rate. If the latest saliency curve minimum value set increases , the gradient direction remains; if If the price decreases, adjust the weights according to the reverse gradient:

[0136]

[0137] in For Dirichlet distribution with parameter , This is the learning rate. The moving average baseline is used to reduce variance. After updating, it is normalized again to ensure the sum is 1. Since only four scalar parameters are adjusted each time, the online update time is less than 0.1 milliseconds, having no impact on real-time performance.

[0138] The saliency curve exhibits multiple local minima. To avoid the high computational cost of global optimization, this invention employs a two-stage search: a coarse scan is performed on the GPU at a step size of 2 frames per second to calculate the first-order difference sign and locate the interval where the curve crosses below the average value; then, within each interval, a high-resolution curve is reconstructed using cubic B-splines, and the precise minimum value is obtained using a quasi-Newton method. This method controls the search time to within 20 milliseconds during hurdling motion, while ensuring that the error with the global scan is less than 2 frames. Minimum value acquisition time. After that, with Map it to a frame number, where The sampling frame rate is 240. All sequence numbers form the keyframe index set. .

[0139] The multi-source structure of the mixed significance curve significantly improves robustness. Experiments show that outdoor backlighting leads to... When gradient weakening, and A clear pulse can still be given at the inflection point of the action; in occluded scenes, Automatic saliency reduction suppresses false alarms. Using only the voxel light field component, the keyframe recall rate is only 82%; adding the neuromorphic component raises it to 89%; after complete four-component fusion and online weight adjustment, the recall rate reaches 94%, with the false alarm rate controlled below 4%. In cross-scene transfer testing, the shooting conditions in the soft-surface training hall and the synthetic track differ significantly; the system only needs to update online for 3 minutes to find stable weights again.

[0140] In terms of practical effectiveness, coaches can quickly locate key stages such as "start" and "takeoff" using semantic tag filters; the mixed saliency remains smooth at high frame rates, facilitating real-time prefetching by the video player. When deployed on mobile devices, calculating the saliency curve only requires reading four cached scalars, significantly reducing computational complexity. On Snapdragon 8 series processors, it can reach 300 frames per second.

[0141] Preferably, the weight coefficients of the mixed saliency curve are adjusted online by a policy gradient algorithm and normalized after each adjustment to keep the sum of all weight coefficients constant.

[0142] The four components of the mixed saliency curve come from the voxel light field, the neuro-morphological state, the semantic confidence and the skeletal latent tensor respectively. Their contribution proportions to the key frame are not constant in different scenarios. For example, the time derivative amplitude of the voxel light field is more reliable in strong outdoor light conditions, while the semantic confidence stream contributes more to filtering false positives in complex indoor backgrounds. To enable the system to automatically adapt to changes in the scene, the present invention uses a policy gradient algorithm to dynamically adjust the weight coefficients during operation, and normalizes the weight vector after each update to keep the sum of the weight coefficients constant at 1.

[0143] In principle, the policy gradient algorithm is derived from reinforcement learning, which updates policy parameters by maximizing expected returns. The mixed saliency curve plays the role of policy in key frame grabbing, and the weight vector is the parameter to be learned, and the return is defined as the external evaluation score . The external evaluation score can be the F1 value between the manually labeled key frame and the system output key frame, or the manual confirmation rate of the downstream editing module. The present invention regards the weight as the expected value of the parameterized probability distribution, and selects the Dirichlet distribution . Let , be the temperature parameter, the smaller the temperature, the sharper the distribution. A weight update can be written as:

[0144]

[0145] where is the learning rate, is the sliding window baseline used to reduce variance; represents the Dirichlet probability density with as the parameter, is the log gradient. Since the log gradient of Dirichlet can be written by the digamma function, and only has constant complexity, the update calculation is very small. After updating, normalization is performed:

[0146]

[0147] to ensure that the sum of the weights is still 1, while avoiding numerical drift.

[0148] In terms of implementation details, the external evaluation score is evaluated at the window level. After the system processes 500 frames, the current key frame index set is compared with the manually labeled or historical best set to generate If no labels are available at the beginning of training, unsupervised metrics such as the harmonic mean of saliency valley depth and inter-frame distance can be used as temporary reward signals. The temperature parameter is initially set to 0.05 to allow slight perturbations in the vicinity of the current weights at each sampling. The learning rate is set to 5 x 10 -4 To prevent overfitting, if the reward decreases for 5 consecutive updates, the weights are frozen for 2000 frames before continuing exploration.

[0149] The online adaptation procedure consists of four steps: weight sampling, saliency curve computation, keyframe extraction and reward evaluation, and weight update. At the beginning of each 500-frame batch, a new weight vector is sampled from the Dirichlet distribution, which is used to compute the mixed saliency curve for all subsequent frames. The system computes the curve based on the weights, searches for local minima, and maps them to a set of keyframe indices. At the end of the batch, the reward is computed based on human annotations or rules . The weights are then updated according to the formula and normalized. The entire process is completely online and does not require interrupting inference.

[0150] In hardware deployment, the weight update overhead is mainly the computation of the digamma function, but because the Dirichlet has only 4 dimensions, the update actually involves only 4 digamma calls, which takes less than 10 microseconds and can be ignored. The weight storage occupies 32 bytes, which is suitable for mobile devices to run persistently.

[0151] In terms of effectiveness, the invention is trained online on a 240-frame-per-second dataset for 100-meter sprint. The initial offline grid search weights are , and the F1 metric is 0.89. After about 10 rounds of online updates for a total of 5000 frames, the weights are adjusted to , and the F1 is improved to 0.93. When the scene is transferred to an indoor training hall, the contribution of the voxel light field term decreases due to more uniform lighting, and the online learning adjusts the weights to in 3 minutes, and the F1 is restored to 0.92. The hurdle dataset shows similar trends, and when frequent occlusions cause unreliable visual gradients, the neuromorphic term weight increases to about 0.4, and the false positive rate drops to 4%.

[0152] To verify the stability of the weights, the invention injects random flash interference into the noisy simulation environment, causing false peaks in the amplitude of the time derivative of the voxel light field. When using the fixed weight scheme, the false detection rate is 9%. After using the online weight adjustment, the false detection rate is compressed to 3%, and the experiment shows that the strategy gradient can quickly reduce the weight of the disturbed component to maintain overall performance.

[0153] The weight update of the application is closely combined with the interaction of the back end clipping. If the coach frequently skips some key frames captured by the system in the playback interface, the confidence labels of these frames are marked as negative samples, resulting in a decrease in return, thereby automatically reducing the corresponding component weight. On the contrary, if the coach marks the missed frames, the system gives a penalty in the subsequent window, encouraging the exploration of new weight combinations. This human-computer collaborative learning makes the system more and more accurate, meeting the needs of the actual training process.

[0154] Preferably, the key frame time point is mapped to a key frame index according to the frame rate of the image acquisition sequence, and the corresponding image frame is extracted from the visible light infrared image sequence based on the key frame index as the key frame output.

[0155] After two-stage minimum value search of the mixed saliency curve, the system obtains a set of high-precision timestamps Each timestamp corresponds to a potential key moment in the track and field action sequence. In order to convert these continuous time values into directly indexable video frames, the application adopts the frame rate mapping principle: the frame rate of the image acquisition sequence is taken as the proportion coefficient The timestamp is multiplied by the frame rate and rounded down to obtain an integer frame number:

[0156]

[0157] Among them The key frame index is represented by The frame rate is written in the configuration file by the hardware clock at the system initialization, and 240 frames per second is a typical value. Four-way modal data is generated for the same timestamp, and the application only selects the visible light infrared image as the key frame output for three reasons: first, the visible light channel has the highest resolution and rich details; second, the infrared channel compensates for structural shadows in backlight scenes; third, the polarization map and event stream are mainly used in the calculation stage, and the experience end pays more attention to intuitive pictures.

[0158] Two engineering problems need to be solved in the mapping process: time quantization error and frame loss tolerance. The quantization error comes from the rounding operation, and the theoretical maximum value is less than 1 frame. Considering that the single-frame displacement of the running action is about 4 mm at 240 frames per second, the error can be almost ignored for action decomposition. If sub-frame level accuracy is required, light flow compensation resampling can be enabled on the player side to generate static illustrations by interpolation, but this is usually unnecessary in the training feedback or referee penalty scene. Frame loss tolerance is aimed at occasional frame loss of the camera: the system maintains a ring buffer to record the mapping table of the frame number and the image in the last 3 seconds; if the frame number mapped by the key frame does not exist, the system backtracks to the latest successful frame and sets a missing frame flag bit for subsequent repair algorithm.

[0159] The key frame extraction process copies the visible light image, infrared image and synchronization metadata from the cache, writes them to the solid state disk, and names them according to the rule "YYYYMMDD_HHMMSS_nnnnn". The prefix records the UTC time, and the suffix records the frame number, which facilitates alignment with external timing systems. Each key frame is also accompanied by a JSON file that saves the following fields: saliency four-component value, action text label, skeletal latent tensor hash, and camera intrinsic snapshot. The hash is used for backend deduplication; the intrinsic snapshot ensures that the world coordinates can be inversely calculated after changing the lens focal length or moving the camera position.

[0160] The key frame set is not only used for the coach to play back the video, but also fed back to the voxel light field incremental saving module. The voxel network freezes the gradient of the voxel density and color weight within the key frame index range to prevent subsequent online fine-tuning from covering the local form that has been confirmed to be of high quality; at the same time, the skeletal pose is interpolated to the precise timestamp and written into the scene database to speed up the initial convergence of the same scene in the future. The database is indexed by multiple keys of events, athlete IDs and dates, and can locate the same action in history within seconds, providing data support for technical comparison.

[0161] Embodiment: On 240 frames per second of 100-meter sprint data, the system outputs an average of 8 key frames, including the start, maximum acceleration, and finish line stages. Artificial evaluation shows that 95% of the key frames fall within 2 frames before and after the action turning point; the remaining 5% is mainly due to camera jitter causing the saliency curve to shift.

[0162] Compared with traditional key frame extraction methods based on optical flow or threshold, the hybrid saliency of the present invention takes into account visual, topological, neuro-morphological and semantic information, and the frame rate mapping link converts continuous time output into discrete index, avoiding repeated decoding. Compared with the baseline system outputting 60 uniformly sampled frames from the full video, the total number of key frames of the present invention is reduced by 80%, but covers all important action nodes; on mobile devices with limited storage bandwidth, an average of 120 megabytes per minute is saved.

[0163] The above is only an embodiment of the present application and is not intended to limit the present application. The present application can have various modifications and changes for those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the scope of the claims of the present application.

Claims

1. An automatic identification and capture method of track and field action based on image data, characterized in that, The method comprises the following steps: Time synchronization is performed on the visible infrared image, the polarization image and the event stream, the first Stokes component and the second Stokes component are extracted in the synchronized polarization image, and a polarization phase image is generated by taking the inverse tangent values of the first Stokes component and the second Stokes component, thereby forming a unified time base data packet; The unified time base data packet is input into a polarization neural radiation field network to generate a voxel light field, a skeletal graph neural network based on Lie group message passing is used to obtain a joint rotation sequence, matrix product state decomposition is performed on the joint rotation sequence to form a skeletal latent tensor, and the voxel light field, the joint rotation sequence and the skeletal latent tensor are coupled into a continuous space-time scene model through alternating optimization; Topological path features are extracted from the Vietoris-Rips filtering and third-order path signature of the joint rotation sequence, the topological path features are encoded into a pulse sequence, a liquid state machine is used to obtain a neuromorphic state stream, a semantic confidence stream is generated by comparing the similarity between the skeletal latent tensor mapping vector and the Chinese action text embedding vector, and the target homology value and the pulse peak time in the neuromorphic state stream are fed back to the continuous space-time scene model; A hybrid saliency curve is constructed by taking the time derivative amplitude of the voxel light field, the neuromorphic state difference, the semantic confidence complementary value and the skeletal latent tensor difference as weights, a local minimum value of the hybrid saliency curve is selected, and the local minimum value is mapped into a key frame index set according to a preset sampling rate.

2. The method of claim 1, wherein, The time synchronization is completed by maximizing the mutual information between the event stream edge and the visible infrared image edge, and the time stamps of the visible infrared image acquisition device, the polarization image acquisition device and the event stream acquisition device are calibrated by using a unified hardware pulse signal.

3. The method of claim 1, wherein, The polarization phase image is generated by comparing the amplitude relationship between the first Stokes component and the second Stokes component to determine the phase of each pixel, and used to compensate for the difference in polarization direction.

4. The method of claim 1, wherein, Each node of the skeletal graph neural network based on Lie group message passing corresponds to a preset human joint, and the message passing is implemented by group logarithmic mapping and group exponential mapping on the rotation difference between adjacent nodes in a three-dimensional rotation group space to update the joint posture.

5. The method of claim 1, wherein, The matrix product state decomposition performs low-rank reconstruction on the joint rotation sequence in a sliding time window to generate the skeletal latent tensor and compress the time sequence redundancy information.

6. The method of claim 1, wherein, The topological path features are extracted by Vietoris-Rips filtering on the joint rotation sequence to obtain persistent homology features, and combined with the third-order path signature features to form a joint representation.

7. The method of claim 1, wherein, The liquid state machine is composed of multiple layers of pulse neurons, the synaptic delay is randomly initialized within a preset range, and the read layer weight is periodically updated to output the neuromorphic state stream.

8. The method of claim 1, wherein, The semantic confidence stream is generated by comparing the cosine similarity between the skeletal latent tensor mapping vector and the Chinese action text embedding vector, and the text label with the highest similarity is selected as the current action label.

9. The method of claim 1, wherein, The weight coefficients of the hybrid saliency curve are adjusted online by using a policy gradient algorithm, and are normalized after each adjustment to keep the sum of all weight coefficients constant.

10. The method of claim 1, wherein, The key frame time points are mapped into key frame indexes according to the frame rate of the image acquisition sequence, and the corresponding image frames are extracted from the visible light infrared image sequence based on the key frame indexes as key frame outputs.

Citation Information

Patent Citations

  • Motion capture method, terminal equipment and storage medium

    CN114638921A

  • Pet dangerous behavior monitoring system based on optic neural network

    CN120108006A