Sports event real-time score statistical method and system based on AI visual identification

By constructing a dynamic skeleton diagram structure and timing constraint model, combining phase consistency measurement and dual-flow feature projection, the accuracy and reliability of the intelligent scoring system in sports events are solved, and accurate scoring statistics and data support are achieved.

CN120298950AActive Publication Date: 2025-07-11ZHONGSHIYUN (BEIJING) TECH CO LTD

Patent Information

Application Number
CN202510438314.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-11
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The existing intelligent scoring system is difficult to accurately capture the athlete's complex movement details in sports events. The changes in equipment status are affected by occlusion and lighting, and lack of modeling the timing correlation between athlete's movements and equipment status, resulting in insufficient scoring accuracy and reliability.

Method used

Using a real-time score statistics method of multimodal feature fusion, a dynamic skeleton graph structure and timing constraint model is constructed, and a phase consistency measurement and anomaly correction mechanism are used, combined with dual-flow feature projection and multi-head attention mechanism, the timing correlation between athletes' movements and equipment status is established.

Benefits of technology

It improves the accuracy and reliability of score judgments, reduces the misjudgment rate, realizes accurate identification and matching of athletes' movements and equipment status, and supports event management, playback review and data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298950A_ABST
    Figure CN120298950A_ABST
Patent Text Reader

Abstract

The invention provides a sports event real-time score statistical method and system based on AI visual identification, and relates to the technical field of intelligent scoring of sports events, and the method comprises the steps: collecting images through multiple cameras, and carrying out target detection to obtain data of athletes, equipment and scoring areas; generating a joint thermodynamic diagram by using the feature pyramid network, and constructing a dynamic skeleton diagram structure to extract action features; performing periodic detection and abnormity correction on the equipment; and performing feature matching and score calculation through a double-flow feature projection network and a multi-head attention mechanism, generating a score data packet including time, athlete numbers, types and video clips, and sending the score data packet to a scoring system. According to the invention, intelligent real-time scoring of sports events is realized, and the scoring accuracy and efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent scoring for competitions, and in particular, to a real-time score statistics method and system for sports competitions based on AI visual recognition. Background Art

[0002] In sports competitions, accurately recording and statistics the scores of the game is an important part of the referee's work. The traditional manual scoring method mainly relies on the on-site judgment of the referee. During the game, the referee needs to observe in real time the completion of the athletes' movements, the changes in the state of the equipment, and the effective triggering of the scoring area, and make score judgments according to the competition rules. With the development of artificial intelligence and computer vision technologies, intelligent scoring systems based on visual recognition have gradually been applied to various sports competitions. The competition pictures are collected by camera devices, and deep learning algorithms are used to assist the referee in making score judgments.

[0003] However, there are still some problems in the actual application of existing intelligent scoring systems. Due to the high speed and complexity of athletes' movements, a single object detection and pose estimation model is difficult to accurately capture the movement details, and it is easy to miss detections and misjudgments. The determination of the state changes of sports equipment and the effective triggering of the scoring area are often affected by factors such as occlusion and lighting, reducing the scoring accuracy. The existing systems lack the modeling of the temporal correlation between athletes' movements and the state of equipment, and cannot effectively distinguish valid scoring actions from invalid actions, affecting the reliability of scoring.

[0004] In summary, the present invention aims to solve the above technical problems, and proposes a real-time score statistics method based on multi-modal feature fusion. By constructing a dynamic skeleton graph structure and a temporal constraint model, the accurate modeling of athletes' movements is realized; the phase consistency metric and the anomaly correction mechanism are used to improve the robustness of equipment state recognition; the two-stream feature projection and the multi-head attention mechanism are adopted to establish the temporal correlation between action features and equipment states, thereby improving the accuracy and reliability of score judgments. Summary of the Invention

[0005] An embodiment of the present invention provides a real-time score statistics method and system for sports competitions based on AI visual recognition, which can solve the problems in the prior art.

[0006] In the first aspect of the embodiment of the present invention, A real-time score statistics method for sports competitions based on AI visual recognition is provided, including: Collecting image data of a sports competition venue through multiple cameras; performing object detection on the image data to obtain athlete image data, equipment image data, and scoring area image data; Generate the athlete joint heatmaps for the athlete image data using a Feature Pyramid Network (FPN), construct a dynamically connected skeleton graph structure through temporal smoothing constraints, and extract the athlete action feature vectors via a Graph Convolutional Network (GCN) and Recurrent Neural Ordinary Differential Equation (RNODE). Perform periodic detection and anomaly correction on the equipment image data using a Phase Congruency Metric Network to obtain the equipment status data. Construct a dynamic graph structure through a Two-Stream Feature Projection Network, and perform feature interaction matching on the athlete action feature vectors and the equipment status data using a multi-head attention mechanism and temporal consistency constraints to calculate the effective score data. Generate a score data packet containing the scoring time, the athlete number, the scoring type, and the scoring video segment based on the effective score data. Send the score data packet to the competition scoring system and store the scoring video segment in the video storage server.

[0007] In an alternative embodiment, Generating the athlete joint heatmaps for the athlete image data using a Feature Pyramid Network (FPN), constructing a dynamically connected skeleton graph structure through temporal smoothing constraints, and extracting the athlete action feature vectors via a Graph Convolutional Network (GCN) and Recurrent Neural Ordinary Differential Equation (RNODE) includes: Perform downsampling and feature mapping on the athlete image data through the Feature Pyramid Network to obtain feature branches at different scales; map the feature branches through convolution operations and activation functions to determine the spatial position distribution of the joints and generate the athlete joint heatmaps. Perform weighted integral calculation on the response values of each joint point in the athlete joint heatmap to obtain the joint spatial coordinates, and perform smoothing constraint calculation in combination with the coordinate information of adjacent temporal frames to obtain the joint smoothed coordinate sequence. Establish a node set based on the joint smoothed coordinate sequence, calculate the node connection probability through the temporal displacement change between nodes, and form a dynamically connected skeleton graph structure. Input the dynamically connected skeleton graph structure into a multi-layer stacked Graph Convolutional Network, perform normalization processing and weight update on the node features in each layer of the network, calculate the difference in velocity vectors and position vectors between nodes, and obtain the node motion correlation features. Decompose and map the node motion correlation features to obtain multiple basic motion representation vectors, calculate the corresponding combined weight coefficients, and fuse them to form the global action feature vector. Input the global action feature vector into the Recurrent Neural Ordinary Differential Equation for temporal modeling, extract multi-scale temporal features through multi-layer dilated convolution, and output the athlete action feature vectors.

[0008] In an alternative embodiment, Establish a node set based on the smooth joint coordinate sequence, calculate the node connection probability through the temporal displacement changes between nodes, and form a dynamic connection skeleton graph structure, including: Construct the joint smooth coordinate sequence into an initial node set, calculate the temporal displacement difference between adjacent nodes in the initial node set, and construct a node state vector; Establish a Bayesian network model based on the node state vector, and encode the spatial position relationship and temporal displacement difference between nodes into a conditional probability distribution; Construct a Markov random field to describe the spatial constraint relationship of nodes. The spatial constraint relationship includes node distance constraint, angle constraint, and topological connection constraint, and generate a spatial potential function of nodes; Combine the conditional probability distribution and the spatial potential function, and use the variational expectation maximization algorithm to iteratively calculate the node connection probability; Construct a probability graph network based on the node connection probability, calculate the mutual information and conditional entropy between each pair of nodes, use the weighted sum of the mutual information and the conditional entropy as the connection uncertainty, and set a node connection threshold according to the connection uncertainty; Determine the connection relationship between nodes according to the node connection threshold, and perform online update on the connection relationship based on the temporal displacement difference between nodes to form a dynamic connection skeleton graph structure.

[0009] In an alternative embodiment, Combining the conditional probability distribution and the spatial potential function, using the variational expectation maximization algorithm to iteratively calculate the node connection probability includes: Multiply the conditional probability distribution by the spatial potential function and normalize it to construct a joint probability distribution model of node states and connection relationships; Construct an auxiliary probability distribution to represent the connection relationship between nodes, and use the auxiliary probability distribution as a variational distribution to approximate the posterior distribution of the joint probability distribution model; Construct a variational objective function based on the joint probability distribution model and the variational distribution. The variational objective function includes an observation data log-likelihood term, a KL divergence term of the posterior distribution, and a spatial potential function term; Perform expectation maximization iterative optimization on the variational objective function. In the expectation step of each iteration, fix the model parameters to optimize the variational distribution. In the maximization step, fix the variational distribution, obtain an adaptive learning rate by calculating the first moment and second moment of the gradient, use the product of the adaptive learning rate and the gradient as the parameter update amount and perform truncation processing, optimize the model parameters, and determine the optimized variational distribution; Calculate the initial connection probability between node pairs based on the optimized variational distribution, and perform softmax normalization in combination with a temperature parameter to obtain a normalized connection probability; Monitor the difference of the variational objective function in adjacent iterations. When the difference is less than a preset difference threshold, determine that the iteration converges and output the final node connection probability.

[0010] In an alternative embodiment, Use a phase consistency metric network to perform periodic detection and anomaly correction on the equipment image data, and obtain equipment status data including: Perform multi-directional and multi-scale Gabor transforms on the equipment image data to obtain the complex response of the equipment image data. Calculate the local phase value based on the complex response, and calculate the phase consistency metric by combining the amplitude of the complex response and the local phase value; Perform a two-dimensional Fourier transform on the phase consistency metric to obtain a power spectrum. Calculate the radial average value of the power spectrum, and determine the periodic characteristics of the equipment image data based on the maximum response of the radial average value; Construct a reference phase consistency template, calculate the difference between the phase consistency metric and the reference phase consistency template as the local anomaly metric, and construct an adaptive threshold based on the mean and standard deviation of the local anomaly metric; Compare the local anomaly metric with the adaptive threshold to identify mutant anomaly regions. Calculate the spatial gradient of the local anomaly metric and compare it with a preset gradient threshold to identify gradual anomaly regions; Perform temporal correction and spatial correction on the mutant anomaly regions and the gradual anomaly regions, and fuse the first result of the temporal correction and the second result of the spatial correction based on an adaptive weight to obtain a corrected phase consistency metric; Convert the corrected phase consistency metric into local state features through a non-linear mapping, and perform weighted aggregation on the local state features to obtain equipment status data.

[0011] In an alternative embodiment, Construct a dynamic graph structure through a two-stream feature projection network, and use a multi-head attention mechanism and temporal consistency constraints to perform feature interaction matching on the athlete's action feature vector and the equipment status data, and calculate and obtain effective score data including: Temporally collect the athlete's action feature vector and the equipment status data based on an adaptive sampling rate, and perform dimension unification processing through a two-stream feature projection network constructed by a multi-layer perceptron and a recurrent neural network to obtain action features and equipment features; Construct the action features and the equipment features as node representations of the dynamic graph structure, calculate the edge weight values between the node representations through a sparse attention weight matrix based on local sensitive hashing, and construct a weighted graph structure; The node representation of the graph structure generates a query matrix, a key matrix, and a value matrix through a multi-scale convolutional network and positional encoding, and performs multi-head parallel calculation using a scaled dot-product attention mechanism adjusted by feature gradients. The fused feature is determined through residual connection and layer normalization; Calculate the temporal attention weights of the fused features through a gated recurrent unit, calculate the attention coefficients by combining the current node hidden state and the context vector, weight-aggregate the features corresponding to the neighboring nodes in the time series based on the attention coefficients, and update the node hidden state using skip connections; Calculate the similarity matrix between the action features and the equipment features, and calculate the Mahalanobis distance of the similarity matrix at adjacent temporal positions through a sliding window to determine the temporal consistency constraint; The similarity matrix and the temporal consistency constraint are non-linearly transformed to obtain an initial score, and the exponential moving average method is used to calculate the maximum similarity mean of the initial score in the temporal dimension to determine the score confidence; Based on the initial score and the score confidence, perform anomaly detection and smoothing filtering through an adaptive threshold to generate effective score data.

[0012] In an alternative embodiment, The node representation of the graph structure generates a query matrix, a key matrix, and a value matrix through a multi-scale convolutional network and positional encoding, and performs multi-head parallel calculation using a scaled dot-product attention mechanism adjusted by feature gradients. Determining the fused feature through residual connection and layer normalization includes: Extract features from the node representation of the graph structure through a multi-scale convolutional network, and perform feature concatenation to obtain a multi-scale feature representation; Introduce positional encoding, dynamically adjust the positional encoding through an adaptive weight matrix to obtain enhanced positional encoding, and perform feature enhancement operation on the enhanced positional encoding and the multi-scale feature representation to obtain enhanced features; Based on the enhanced features, generate a query matrix, a key matrix, and a value matrix through a query weight matrix, a key weight matrix, and a value weight matrix respectively. Calculate the feature gradients of the enhanced features, and input the feature gradients into a feature adjustment network to obtain an attention adjustment factor; Perform multi-head parallel calculation on the query matrix, the key matrix, and the value matrix through a scaled dot-product attention mechanism to obtain multiple attention head features, dynamically adjust the multiple attention head features through the attention adjustment factor to obtain an attention output feature, and perform residual connection on the attention output feature and the multi-scale feature representation to obtain a feature fusion result; Input the feature fusion result into a layer normalization network to generate a normalization coefficient and a bias parameter, and process to obtain the final fused feature.

[0013] In the second aspect of the embodiments of the present invention, Provided is a real-time score statistics system for sports competitions based on AI vision recognition, including: A first unit, configured to collect image data of a sports competition venue through multiple cameras; perform object detection on the image data to obtain athlete image data, equipment image data, and scoring area image data; A second unit, configured to generate an athlete joint heat map from the athlete image data by using a feature pyramid network, construct a dynamic connection skeleton graph structure through temporal smoothing constraints, and extract an athlete action feature vector through a graph convolutional network and a recurrent neural ordinary differential equation; A third unit, configured to perform periodic detection and anomaly correction on the equipment image data by using a phase consistency metric network to obtain equipment status data; A fourth unit, configured to construct a dynamic graph structure through a two-stream feature projection network, perform feature interaction matching on the athlete action feature vector and the equipment status data by using a multi-head attention mechanism and temporal consistency constraints, and calculate and obtain effective score data; A fifth unit, configured to generate a score data packet including score time, score athlete number, score type, and score video segment according to the effective score data; A sixth unit, configured to send the score data packet to a competition scoring system and store the score video segment in a video storage server.

[0014] In a third aspect of the embodiments of the present invention, Provided is an electronic device, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0015] In a fourth aspect of the embodiments of the present invention, Provided is a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0016] In the embodiments of the present invention, through multi-camera acquisition and advanced AI vision recognition technology, the automatic statistics of real-time scores in sports competitions are realized, significantly improving the accuracy and efficiency of scoring. It can accurately capture the complex motion characteristics of athletes, effectively reducing human judgment biases; the phase consistency metric network and the two-stream feature projection network are introduced, which can accurately identify the equipment status and match it with the athletes' movements, significantly improving the accuracy of score determination, enhancing the system's understanding ability of complex competition scenarios, and effectively reducing the misjudgment rate; automatically generate score data packets containing detailed information and transmit them to the competition scoring system in real-time, while saving relevant video clips, providing comprehensive support for event management, playback review, and data analysis, not only improving the fairness and viewing experience of the competition, but also providing strong support for the digital development of sports events. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a flowchart of the method for real-time score statistics of sports events based on AI vision recognition in the embodiments of the present invention; Figure 2 is a comparison diagram of the simulation effects of the dynamic skeleton connection structure under complex high-speed movements; Figure 3 is a heat map of the parameter sensitivity of the spatial potential energy function; Figure 4 is a flowchart of the feature enhancement and dynamic attention mechanism. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0019] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0020] Figure 1 is a flowchart of the method for real-time score statistics of sports events based on AI vision recognition in the embodiments of the present invention, as Figure 1 shown, the method includes: Collect image data of the sports event venue through multiple cameras; perform object detection on the image data to obtain athlete image data, equipment image data, and scoring area image data; Generate the athlete joint heatmaps from the athlete image data using a Feature Pyramid Network (FPN), construct a dynamically connected skeleton graph structure through temporal smoothing constraints, and extract the athlete action feature vectors via a Graph Convolutional Network (GCN) and Recurrent Neural Ordinary Differential Equation (RNODE). Perform periodic detection and anomaly correction on the equipment image data using a Phase Congruency Metric Network (PCMN) to obtain the equipment status data. Construct a dynamic graph structure through a Two-Stream Feature Projection Network (TSFPN), and perform feature interaction matching on the athlete action feature vectors and the equipment status data using a multi-head attention mechanism and temporal consistency constraints to calculate the effective score data. Generate a score data packet containing the scoring time, the athlete number, the scoring type, and the scoring video segment based on the effective score data. Send the score data packet to the competition scoring system and store the scoring video segment in the video storage server.

[0021] In an optional embodiment, generating the athlete joint heatmaps from the athlete image data using a Feature Pyramid Network (FPN), constructing a dynamically connected skeleton graph structure through temporal smoothing constraints, and extracting the athlete action feature vectors via a Graph Convolutional Network (GCN) and Recurrent Neural Ordinary Differential Equation (RNODE) includes: Perform downsampling and feature mapping on the athlete image data through the Feature Pyramid Network (FPN) to obtain feature branches at different scales; map the feature branches through convolution operations and activation functions to determine the spatial position distribution of the joints and generate the athlete joint heatmaps. Perform weighted integral calculation on the response values of each joint point in the athlete joint heatmaps to obtain the joint spatial coordinates, and perform smoothing constraint calculation in combination with the coordinate information of adjacent temporal frames to obtain the joint smooth coordinate sequence. Establish a node set based on the joint smooth coordinate sequence, calculate the node connection probability through the temporal displacement changes between the nodes, and form a dynamically connected skeleton graph structure. Input the dynamically connected skeleton graph structure into a multi-layer stacked Graph Convolutional Network (GCN), perform normalization processing and weight update on the node features in each layer of the network, calculate the velocity vector difference and position vector difference between the nodes, and obtain the node motion correlation features. Perform decomposition mapping on the node motion correlation features to obtain multiple basic motion representation vectors, calculate the corresponding combined weight coefficients, and fuse them to form the global action feature vector. Input the global action feature vector into a Recurrent Neural Ordinary Differential Equation (RNODE) for temporal modeling, extract multi-scale temporal features through multi-layer dilated convolution, and output the athlete action feature vector.

[0022] In a specific embodiment, the collected athlete image data is input into a Feature Pyramid Network, which consists of five descending structures. The sampling rates of each layer structure are 1 / 1, 1 / 2, 1 / 4, 1 / 8, and 1 / 16 of the original size respectively. In each downsampling process, a 3×3 convolutional operation with a stride of 2 is used to halve the image size. The number of feature channels output by each scale layer is set to 256. In this way, the original input image with a resolution of 640×480 can generate five feature branches of different scales, namely feature maps of 640×480×256, 320×240×256, 160×120×256, 80×60×256, and 40×30×256 respectively.

[0023] The feature branches are mapped to joint heatmaps through 1×1 convolutional operations. Each joint point corresponds to one channel. For example, for 17 human joint points, the number of output channels is 17. The convolutional feature maps are processed by the ReLU activation function to generate heatmaps with non-negative values, where the value of each pixel represents the probability of the existence of a specific joint point at that position. For the image data of gymnasts, the generated heatmaps show high-response value regions at key joint points such as the head, shoulders, and knees, and the value range is between 0 and 1.

[0024] The response values in each joint heatmap are weighted and integrated according to the coordinates. The calculation method is to multiply the coordinate value of each pixel position in the heatmap by the corresponding response value, then sum over the entire heatmap, and finally divide by the sum of all response values in the heatmap to obtain the floating-point coordinates of the joint point. For example, for a shoulder joint heatmap, if the highest response region is concentrated near (320, 150), the calculated exact coordinates may be (319.78, 151.23).

[0025] To reduce jitter, a temporal smoothing constraint is introduced. For each joint point, the coordinates of the current frame and the previous three frames are weighted and averaged, where the weight of the current frame is 0.5, the previous frame is 0.3, the second previous frame is 0.15, and the third previous frame is 0.05. For example, if the coordinates of a certain joint in four frames are (319.78, 151.23), (318.92, 150.87), (320.15, 152.01), and (319.56, 151.45) respectively, the smoothed coordinates are approximately (319.53, 151.24).

[0026] Based on the acquired joint smooth coordinate sequence, each joint point is regarded as a node in the graph, and the connection probability is calculated according to the displacement change of the node in adjacent frames. If the displacement change synergy between two nodes is high, the connection probability increases. Specifically, the cosine similarity between the node pairs in two adjacent frames is calculated, and a connection relationship is established when the cosine similarity is greater than 0.8. For the action sequence of a diving athlete, during the falling process, the connection probability of the arm and torso nodes is about 0.95, while the connection probability of the arm and leg nodes is about 0.3, thereby forming a dynamic skeleton diagram that conforms to the biomechanical characteristics of the human body.

[0027] The constructed dynamic connection skeleton graph is input into a three-layer stacked graph convolution network for feature extraction. In each layer of graph convolution, the node features are normalized by subtracting the center node value from the average value of the neighboring nodes and dividing by the standard deviation. For the first layer of graph convolution, the input is the original coordinates and displacement information of the joint points, and the output feature dimension is 64. For example, for a shoulder joint node, the input features are the three-dimensional coordinates (319.53, 151.24, 0) and the displacement vector (0.75, 0.37, 0), which are converted into a 64-dimensional feature vector after graph convolution.

[0028] During the graph convolution process, the motion correlation features between nodes are calculated. For connected node pairs, the velocity vector difference and position vector difference are calculated. For example, the velocity difference vector between the shoulder joint and the elbow joint is (-1.2, 0.8, 0.5), and the position difference vector is (35.6, 12.3, 2.1). These differences reflect the relative motion patterns between joints. For gymnasts' horizontal bar movements, the arm and torso nodes show a highly coordinated velocity difference pattern during rotation, and the difference range is usually between (-2, 2).

[0029] The attention mechanism is used to decompose the node features into 8 basic motion representation vectors, each with a dimension of 32. Each basic vector represents a basic motion mode, such as flexion and extension, rotation, translation, etc. The weight coefficient of each basic vector is determined by calculating the attention score. The weight value ranges from 0 to 1, and the sum is 1. For example, for the flipping action of a diving athlete, the weight coefficient of the rotation basic vector is about 0.6, while the weight coefficient of the translation basic vector is about 0.1. The basic vectors are weighted and fused according to the weight coefficients to form a 256-dimensional global motion feature vector.

[0030] The global motion feature vector is input into the recursive neural ordinary differential equation network for time series modeling. The recursive neural ordinary differential equation network contains three layers of recursive structure, and each layer uses four dilated convolutions with dilation rates of 1, 2, 4, and 8 to extract multi-scale time series features. For a gymnast's motion sequence of 90 frames in length, a 512-dimensional motion feature vector is generated after time series modeling, which contains the temporal and spatial information of the athlete's complete motion.

[0031] In this embodiment, through the Feature Pyramid Network and downsampling, feature branches of different scales are extracted, providing multi-level information with both fine-grained and global background for subsequent joint positioning; based on the smooth coordinate sequence, a node set is constructed and the temporal displacement and connection probability between nodes are calculated, thus forming a skeleton graph structure that can reflect the dynamic posture changes of athletes, providing rich structural information for action recognition; through a multi-layer stacked graph convolutional network, the node features are normalized and the weights are updated. By calculating the difference between the velocity and position vectors, the motion correlation between each joint is effectively captured, enhancing the expression ability of local motion features; the node motion correlation features are decomposed and mapped into multiple basic motion representation vectors, and combined with the corresponding combined weight coefficients to fuse into a global action feature vector, thereby realizing the comprehensive description of the entire action process and improving the robustness of action recognition.

[0032] In an alternative embodiment, a node set is established based on the joint smooth coordinate sequence, and the node connection probability is calculated through the temporal displacement change between nodes. The formation of the dynamic connection skeleton graph structure includes: The joint smooth coordinate sequence is constructed into an initial node set, and the temporal displacement difference between adjacent nodes in the initial node set is calculated to construct a node state vector; Based on the node state vector, a Bayesian network model is established, and the spatial position relationship and temporal displacement difference between nodes are encoded as conditional probability distributions; A Markov random field is constructed to describe the spatial constraint relationship of nodes. The spatial constraint relationship includes node distance constraint, angle constraint, and topological connection constraint, and a spatial potential function of nodes is generated; Combining the conditional probability distribution and the spatial potential function, the variational expectation-maximization algorithm is used to iteratively calculate the node connection probability; Based on the node connection probability, a probability graph network is constructed, the mutual information and conditional entropy between each pair of nodes are calculated, and the weighted sum of the mutual information and the conditional entropy is used as the connection uncertainty. According to the connection uncertainty, a node connection threshold is set; According to the node connection threshold, the connection relationship between nodes is determined, and the connection relationship is updated online based on the temporal displacement difference between the nodes to form a dynamic connection skeleton graph structure.

[0033] In a specific embodiment, three-dimensional coordinate data of human body bone joint points are obtained through a depth camera or a multi-view camera system to form a time-series coordinate sequence {P_t}, where P_t represents the set of all joint point coordinates at time t. To eliminate noise interference, a bilateral filter is used to smooth the original coordinate sequence, with the time-domain window size set to 5 frames, the spatial distance parameter set to 0.05 meters, and the value-domain similarity parameter set to 0.1. Taking the human body bones as an example, a total of 18 joint points are identified, including the head, neck, shoulders, elbows, wrists, hips, knees, and ankles. These joint points are used as the initial node set N = {n_1, n_2, ..., n_18}.

[0034] Calculate the time-series displacement differences between adjacent nodes in the initial node set to construct a node state vector. For each node n_i, record the displacement change amount within 10 consecutive frames, and calculate the relative displacement difference D_ij with the adjacent node n_j. The relative displacement difference is calculated through the difference between the displacement vectors of the two nodes. For example, the average relative displacement difference between the head node n_1 and the neck node n_2 within 10 consecutive frames is 0.015 meters. Combine the displacement differences of each node with all its potential adjacent nodes to form a state vector S_i = {D_i1, D_i2, ..., D_iN}.

[0035] Based on the node state vector, establish a Bayesian network model, and encode the spatial position relationship and time-series displacement difference between nodes into a conditional probability distribution. Construct a directed probability graph G = (V, E), where V is the node set and E is the edge set. For each pair of nodes (n_i, n_j), calculate the conditional probability P(n_j|n_i) according to their state vectors S_i and S_j, which represents the state distribution of n_j under the condition that the state of n_i is known. The conditional probability is estimated through a Gaussian mixture model, and the expectation-maximization algorithm is used to determine the model parameters. In practice, set the number of Gaussian mixture components to 3, the upper limit of the number of iterations to 100, and the convergence threshold to 0.001. For example, the conditional probability P(n_6|n_5) between the elbow node n_5 and the wrist node n_6 is 0.87, indicating that the elbow state has a strong predictive ability for the wrist state.

[0036] Construct a Markov random field to describe the spatial constraint relationships among nodes, and define an undirected graph \(H=(V, E')\), where \(V\) is the same as the set of nodes in the Bayesian network, and \(E'\) represents the set of undirected edges between nodes. The spatial constraint relationships include node distance constraints, angle constraints, and topological connection constraints. The distance constraint limits the spatial distance range between two nodes. For example, the normal distance range from the shoulder to the elbow is from 0.25 meters to 0.35 meters. The angle constraint specifies the range of joint movement angles. For example, the normal bending angle range of the elbow joint is from 0 to 150 degrees. The topological connection constraint ensures the overall structural rationality of the skeleton. For example, a node is connected to at most 4 other nodes. Based on these constraints, a spatial potential function \(U(n_i, n_j)\) of the nodes is generated, which represents the compatibility between the node pair \((n_i, n_j)\). The lower the value of the potential function, the higher the likelihood of connection between the two nodes. For example, the potential function value of adjacent joints such as the shoulder and the elbow is 0.12, while the potential function value of non - adjacent joints such as the shoulder and the ankle joint is 0.89.

[0037] Combining the conditional probability distribution and the spatial potential function, the variational expectation - maximization algorithm is used to iteratively calculate the node connection probability. Define the objective function \(F = \alpha\cdot P(n_j|n_i)-\beta\cdot U(n_i, n_j)\), where \(\alpha\) and \(\beta\) are balance parameters. Through cross - validation, \(\alpha = 0.6\) and \(\beta = 0.4\) are determined. The variational expectation - maximization algorithm iterates by alternately performing the expectation step and the maximization step: In the expectation step, the current model parameters are fixed, and the posterior distribution of the latent variables is calculated; in the maximization step, the updated posterior distribution is used to optimize the model parameters. Set the number of iterations to 20 and the convergence threshold to 0.01. After the iteration is completed, the connection probability \(C_{ij}\) between the node pair \((n_i, n_j)\) is obtained. For example, the connection probability between the shoulder node and the elbow node is 0.92, indicating that these two nodes are very likely to be directly connected.

[0038] Based on the node connection probability, a probabilistic graph network is constructed, and the mutual information and conditional entropy between each pair of nodes are calculated. For the node pair \((n_i, n_j)\), the mutual information \(I(n_i;n_j)\) represents the amount of information they share, and the conditional entropy \(H(n_i|n_j)\) represents the uncertainty of \(n_i\) given \(n_j\). The weighted sum of the mutual information and the conditional entropy \(\lambda\cdot I(n_i;n_j)+(1 - \lambda)\cdot H(n_i|n_j)\) is used as the connection uncertainty \(U_{ij}\), where \(\lambda\) is determined to be 0.7 through experiments. For example, the mutual information of the shoulder - elbow node pair is 0.85, the conditional entropy is 0.25, and the calculated connection uncertainty is 0.67. According to the connection uncertainty, a node connection threshold \(\tau\) is set. In practice, \(\tau = 0.75\) is determined through receiver operating characteristic curve analysis.

[0039] Determine the connection relationship between nodes according to the node connection threshold. When C_ij > τ, establish a connection between nodes n_i and n_j. Based on the temporal displacement difference between nodes, online update the connection relationship to form a dynamic connection skeleton graph structure. During the online update process, use the sliding window method with the window size set to 30 frames. Within each window, recalculate the temporal displacement difference of the nodes, update the conditional probability distribution and the spatial potential function, and then update the node connection probability. When it is detected that the change in the connection probability exceeds the threshold δ = 0.15, dynamically adjust the connection relationship. For example, during normal walking, the connection between the hip and knee nodes remains stable, while during difficult movements such as aerial rotation, the connection between certain joints may be temporarily adjusted to adapt to the special posture.

[0040] Existing skeleton graph construction techniques mainly adopt methods based on template matching or directly connecting predefined joint points, which are difficult to adapt to action changes and individual differences. For example, traditional methods use fixed skeleton templates and map the detected joint points onto the templates, unable to handle unconventional actions or special body types. Another type of method uses deep learning to directly predict joint connections, but it relies on a large amount of labeled data and lacks physical constraints.

[0041] This embodiment introduces a probabilistic graph model combined with temporal dynamic analysis to adaptively generate a skeleton structure; captures the joint motion correlation through the temporal displacement difference, avoiding relying solely on static spatial positions; combines Bayesian networks and Markov random fields, considering both directed dependency relationships and undirected constraint relationships; introduces information theory metrics to evaluate the connection uncertainty, and realizes the dynamic adjustment of the skeleton structure.

[0042] The experimental results show that the method of this embodiment improves the skeleton recognition accuracy by 15.3% when dealing with complex action sequences, enhances the adaptability to unconventional actions by 23.7%, and at the same time improves the real-time performance of skeleton tracking from 25 frames per second to 40 frames per second, significantly improving the skeleton structure recognition effect in a dynamic environment.

[0043] Such as Figure 2As shown, it presents a comparison of the simulation effects of three different skeleton graph construction techniques when dealing with high-speed rotational jumping movements (angular velocity up to 235° / second). The figure respectively shows the skeleton connection structures generated by this technical solution, the template matching method, and the deep learning prediction method, and the corresponding connection probability values are marked on each connection line. It can be clearly seen from the figure that in the same high-speed action scenario, the skeleton structure generated by this technical solution is the most accurate and complete, with a connection accuracy rate of 87.5%, a processing time of only 18 milliseconds, and an adaptability score of 84.2 points (out of 100). The connection relationships between all 15 joint points are correctly recognized, and the connection probabilities are generally high, ranging from 0.98 from the head to the neck to 0.89 at the ankle joint, indicating that even in a high-speed motion state, this method can still maintain a highly reliable skeleton structure recognition.

[0044] In contrast, the template matching method performs poorly under the same conditions, with a connection accuracy rate of only 57.8%, a processing time of 32 milliseconds, and an adaptability score of only 53.6 points. It can be seen from the figure that there are many obvious errors in the template matching method: the left wrist (n_7) is wrongly directly connected to the shoulder (n_3) instead of the elbow (n_5); the connection of the right wrist (n_8) is completely missing; the left ankle joint (n_14) is wrongly connected to the right knee (n_13); the connection of the right ankle joint (n_15) is also completely missing. These errors are mainly due to the fact that the template matching method cannot effectively adapt to the rapid changes in joint positions during high-speed rotational movements, resulting in incorrect matching or matching failures.

[0045] The deep learning prediction method performs between the two, with a connection accuracy rate of 69.4%, a processing time of 27 milliseconds, and an adaptability score of 67.8 points. Although this method can recognize all joint connection relationships, the connection probabilities are generally low, especially for the connections at the ends of the limbs. For example, the connection probabilities of the left and right wrists are only 0.68 and 0.67, and the connection probabilities of the left and right ankle joints are 0.63 and 0.62, indicating that this method has insufficient confidence in judging joint connections in high-speed action scenarios. This uncertainty may lead to instability in connection relationships in continuous action sequences, affecting the coherence of skeleton tracking.

[0046] Through this comparison of simulation effects, it can be intuitively seen the significant advantages of this technical solution in dealing with high-speed and complex movements. Especially its method of capturing the correlation of joint movements by combining temporal displacement differences enables the skeleton structure to better adapt to rapidly changing action postures while maintaining a high processing speed and connection accuracy. These results fully verify the technical innovation and practical value of this solution in dynamic skeleton structure recognition.

[0047] In this embodiment, by combining the Bayesian network and the Markov random field, the temporal dynamic relationship and spatial structure constraints of the nodes are comprehensively considered, effectively enhancing the stability and expression ability of the skeleton graph structure under different motion states; the conditional probability is used to model the state dependence relationship of the nodes, and the spatial potential energy function is used to evaluate the structural compatibility between the nodes, so that the final node connection not only has statistical support but also conforms to the physiological and motion structure logic, significantly improving the credibility of the connection relationship; the sliding window and online update mechanism are introduced, enabling the skeleton graph to adjust the connection relationship in real time according to the action changes, which is particularly suitable for dealing with complex or sudden posture changes, such as high-dynamic actions like jumping and rotating; by introducing mutual information and conditional entropy to measure the uncertainty of the connection and setting connection thresholds for screening, the phenomena of redundant connections and misconnections are significantly reduced, improving the accuracy and discriminability of the skeleton graph structure; using the high-dimensional three-dimensional bone data provided by the depth camera or multi-view system, combined with bilateral filtering and displacement difference modeling, enhances the robustness of the model to multi-source data noise and ensures the stability of node state estimation.

[0048] In an alternative embodiment, combining the conditional probability distribution and the spatial potential energy function, the variational expectation-maximization algorithm is used to iteratively calculate the node connection probability, including: Multiply the conditional probability distribution by the spatial potential energy function and normalize it to construct a joint probability distribution model of the node state and the connection relationship; Construct an auxiliary probability distribution to represent the connection relationship between the nodes, and use the auxiliary probability distribution as the variational distribution to approximate the posterior distribution of the joint probability distribution model; Based on the joint probability distribution model and the variational distribution, construct a variational objective function, which includes an observation data log-likelihood term, a KL divergence term of the posterior distribution, and a spatial potential energy function term; Perform expectation-maximization iterative optimization on the variational objective function. In the expectation step of each iteration, fix the model parameters to optimize the variational distribution. In the maximization step, fix the variational distribution, obtain the adaptive learning rate by calculating the first moment and the second moment of the gradient, and use the product of the adaptive learning rate and the gradient as the parameter update amount and truncate it to optimize the model parameters and determine the optimized variational distribution; Based on the optimized variational distribution, calculate the initial connection probability between node pairs, and perform softmax normalization with the temperature parameter to obtain the normalized connection probability; Monitor the difference of the variational objective function in adjacent iterations. When the difference is less than the preset difference threshold, determine that the iteration converges and output the final node connection probability.

[0049] In a specific embodiment, a joint probability distribution model of node states and connection relationships is constructed. The joint probability distribution model is obtained by multiplying a conditional probability distribution by a spatial potential energy function and then performing a normalization process. The conditional probability distribution describes the dependence of node states on connection relationships, while the spatial potential energy function characterizes the relative position relationships of nodes in space. For example, for a network containing 10 nodes, a 10×10 matrix can be used to represent the conditional probability distribution, and each element in the matrix represents the conditional probability between a pair of nodes. The spatial potential energy function can be defined based on the Euclidean distance between nodes, and the closer the distance, the lower the potential energy. Multiplying these two parts and normalizing them yields the joint probability distribution model that describes the entire network structure.

[0050] Next, an auxiliary probability distribution is introduced as a variational distribution to approximate the posterior distribution of the joint probability distribution model. This auxiliary probability distribution can also be represented by a 10×10 matrix, and each element in the matrix represents the probability of the existence of a connection between the corresponding pair of nodes. Initially, all elements can be set to 0.5, indicating an equal possibility of connection.

[0051] Then, a variational objective function is constructed. This function includes three main parts: the log-likelihood term of the observed data, the KL divergence term of the posterior distribution, and the spatial potential energy function term. The log-likelihood term reflects the fitting degree of the model to the observed data, the KL divergence term measures the difference between the variational distribution and the true posterior distribution, and the spatial potential energy function term introduces the constraint of spatial position information. The weighted sum of these three terms constitutes the final variational objective function.

[0052] Based on the constructed variational objective function, an expectation-maximization iterative optimization process is performed. In the expectation step of each iteration, the model parameters are fixed, and the variational distribution is optimized. Specifically, the gradient of the variational objective function with respect to the variational distribution can be calculated, and the variational distribution can be updated using the gradient ascent method. In the maximization step, the variational distribution is fixed, and the model parameters are optimized. Here, an adaptive learning rate method is adopted, and the learning rate is determined by calculating the first moment and the second moment of the gradient. For example, exponential moving average can be used to estimate the first moment and the second moment of the gradient, and then the learning rate can be dynamically adjusted according to these two statistics. The product of the adaptive learning rate and the gradient is used as the parameter update amount, and a truncation process is performed to prevent an overly large update step. Assuming the initial learning rate is 0.01, it can be adjusted to between 0.001 and 0.1 according to the gradient statistics.

[0053] After each iteration, the initial connection probabilities between node pairs are calculated based on the optimized variational distribution. This can be achieved by directly reading the corresponding element values in the variational distribution matrix. To obtain the normalized connection probabilities, a temperature parameter is introduced and softmax normalization is performed. The temperature parameter controls the smoothness of the probability distribution. A lower temperature makes the probability distribution steeper, while a higher temperature makes the distribution flatter. For example, the temperature parameter can be chosen as 0.1, which will make the high-probability connections more prominent.

[0054] Finally, monitor the difference between the variational objective functions in adjacent iterations to determine whether the iteration has converged. When this difference is less than a preset threshold, it is considered that the iteration has converged, and the final node connection probabilities can be output. For example, the difference threshold can be set to 0.001. If the change in the objective function for 5 consecutive iterations is less than this threshold, it is considered that the algorithm has converged.

[0055] To better illustrate the practical application of this method, we can consider a specific example. Suppose there is a small network consisting of 5 nodes, and the initial conditional probability distribution between the nodes is as follows: ; The spatial potential function can be defined based on the distance between nodes. Suppose the coordinates of the nodes are (0, 0), (1, 1), (2, 0), (1, -1), (3, 1) respectively.

[0056] Initialize the variational distribution as a matrix of all 0.5, indicating equal connection possibilities between all node pairs. Set the initial learning rate to 0.01, the temperature parameter to 0.1, and the convergence threshold to 0.001.

[0057] In the first iteration, the expectation step updates the variational distribution, and the maximization step updates the model parameters. Suppose the updated variational distribution becomes: ; After softmax normalization, the obtained normalized connection probabilities may be: ; This process will be iterated continuously until the change in the variational objective function is less than 0.001. The finally output connection probability matrix reflects the connection strength between nodes in the network and can be used for subsequent network analysis and applications.

[0058] Existing skeleton connection probability calculation techniques mainly use deterministic methods or simple probability models, such as threshold judgment based on geometric distance or basic Bayesian inference. For example, traditional methods usually predefine the skeleton structure, rigidly connect the joint points to a fixed template, or determine the connection relationship based on a simple distance threshold. These methods perform poorly in dealing with complex actions or non-standard body shapes and lack the ability to model dynamic connection relationships.

[0059] The method of this embodiment constructs a more accurate joint probability model by introducing a variational inference framework and an adaptive optimization algorithm; constructs a comprehensive joint probability model by fusing conditional probability distributions and spatial potential functions, while considering node state dependencies and spatial constraints; introduces a variational distribution to approximate the complex posterior distribution to solve the computational complexity of directly calculating the posterior distribution; and uses an expectation maximization algorithm with an adaptive learning rate, combined with temperature-regulated softmax normalization to improve the optimization efficiency and result stability.

[0060] Experimental results show that, compared with traditional methods, the method of this embodiment has a 21.5% increase in the accuracy of connection relationships in complex action recognition, an 18.3% improvement in robustness against abnormal postures, and especially in scenarios of fast movement and partial occlusion, the accuracy of node connection prediction is increased by 25.7%. At the same time, the variational optimization framework reduces the computational complexity, increases the processing speed by 35%, and realizes more efficient real-time skeleton connection relationship inference.

[0061] As Figure 3As shown, it demonstrates the influence of the parameter λ of the spatial potential energy function on the accuracy of networks of different scales. The depth of color represents the level of accuracy (the darker the color, the higher the accuracy). It can be clearly seen from the heat map that the changing trends of accuracy with respect to λ are similar for various network scales: it reaches the highest point near λ = 1.0 and then gradually decreases as the parameter deviates from the optimal value. Specifically, for the small-scale network (20 nodes), the accuracy reaches the highest value of 92.4% at λ = 1.0. When λ drops to 0.1, the accuracy decreases to 86.5%, and when λ increases to 10.0, it drops to 82.1%. The medium-scale network (100 nodes) also performs best at λ = 1.0 with an accuracy of 82.6%, and it drops to 76.2% and 72.8% at the extreme parameter points (λ = 0.1 and λ = 10.0) respectively. The optimal parameter for the large-scale network (200 nodes) is also λ = 1.0 with an accuracy of 76.3%, while the ultra-large-scale network (500 nodes) reaches the best accuracy of 69.8% under the same parameter settings. It can also be seen from the heat map that the larger the network scale, the higher the parameter sensitivity, which is manifested as more obvious color changes in the heat map. For example, the accuracy drop of the 500-node network from the optimal parameter (λ = 1.0) to the extreme parameter (λ = 10.0) reaches 10.7 percentage points (69.8% → 59.1%), while the corresponding drop of the 20-node network is only 10.3 percentage points (92.4% → 82.1%). This indicates that the present technical solution has a certain sensitivity to parameter settings, but the sensitive interval is relatively wide (it performs relatively stably within the range of λ = 0.5 - 2.0), providing a large parameter tuning space for practical applications.

[0062] In this embodiment, by constructing a joint probability distribution model, the conditional probability distribution and the spatial potential energy function are organically integrated, taking into account both the temporal dependence and spatial structure constraints between nodes, making the modeling of node connection relationships more accurate and reasonable; using the auxiliary probability distribution to approximate the posterior distribution and constructing a variational objective function, while retaining the model expressiveness, it greatly reduces the computational difficulty of directly solving the complex posterior distribution, improving the inference efficiency and controllability; adopting the first-order and second-order moment estimates to dynamically adjust the learning rate, making the parameter update more stable and faster, and at the same time suppressing abnormal updates through gradient truncation processing, ensuring the stability of the training process; automatically judging the convergence state by monitoring the change of the objective function, stopping the iteration without manual intervention, and improving the practicality of the algorithm and the convenience of engineering deployment.

[0063] In an alternative embodiment, a phase consistency metric network is used to perform periodic detection and anomaly correction on the equipment image data, and the obtained equipment status data includes: Perform multi - directional and multi - scale Gabor transforms on the equipment image data to obtain the complex response of the equipment image data. Calculate the local phase value based on the complex response, and calculate the phase congruency metric by combining the amplitude of the complex response and the local phase value. Perform a two - dimensional Fourier transform on the phase congruency metric to obtain a power spectrum. Calculate the radial average value of the power spectrum, and determine the periodic characteristics of the equipment image data based on the maximum response of the radial average value. Construct a reference phase congruency template, calculate the difference between the phase congruency metric and the reference phase congruency template as the local anomaly metric, and construct an adaptive threshold based on the mean and standard deviation of the local anomaly metric. Compare the local anomaly metric with the adaptive threshold to identify the mutant anomaly regions. Calculate the spatial gradient of the local anomaly metric and compare it with a preset gradient threshold to identify the gradual anomaly regions. Perform temporal correction and spatial correction on the mutant anomaly regions and the gradual anomaly regions, and fuse the first result of the temporal correction and the second result of the spatial correction based on an adaptive weight to obtain a corrected phase congruency metric. Convert the corrected phase congruency metric into local state features through a non - linear mapping, and perform weighted aggregation on the local state features to obtain equipment state data.

[0064] In a specific embodiment, select Gabor filters with 8 directions and 4 scales, perform convolution operations on the input equipment image data, and obtain 32 complex response maps. Each complex response map contains a real part and an imaginary part, corresponding to the cosine and sine components of the Gabor filter respectively.

[0065] For each pixel position, calculate the phase angle using the real part and the imaginary part of the complex response to obtain the local phase value. The local phase value reflects the geometric features of the local structure of the image.

[0066] Within the local neighborhood of each pixel position, statistically analyze the distribution of the phase values and calculate the phase congruency metric. The phase congruency metric reflects the degree of phase consistency within the local region, and the larger the value, the more regular the local structure.

[0067] Perform a two - dimensional Fourier transform on the phase congruency metric to obtain a power spectrum. Calculate the radial average value of the power spectrum, that is, accumulate and average the two - dimensional power spectrum along the radial direction to obtain a one - dimensional radial average curve. Determine the periodic characteristics of the equipment image data based on the maximum response of the radial average curve. Exemplarily, if there is an obvious peak at a certain frequency in the radial average curve, the period corresponding to this frequency is the main periodic characteristic of the equipment image.

[0068] To construct a reference phase consistency template, multiple typical normal equipment images can be selected, calculate their phase consistency metrics, and take the average as the reference template. Calculate the difference between the phase consistency metric and the reference template as the local anomaly metric. The local anomaly metric reflects the deviation degree between the image to be detected and the normal reference.

[0069] Construct an adaptive threshold based on the mean and standard deviation of the local anomaly metric, and take the mean plus 2 times the standard deviation as the adaptive threshold. Compare the local anomaly metric with the adaptive threshold. If the anomaly metric of a certain area exceeds the threshold, it is identified as a mutant anomaly area.

[0070] Calculate the spatial gradient of the local anomaly metric. Use the Sobel operator to calculate the gradients in the x and y directions, and compare the gradient magnitude with a preset gradient threshold. If the gradient magnitude exceeds the threshold, it is identified as a gradual change anomaly area. The gradient threshold can be adjusted according to the actual application scenario, for example, set to 20% of the mean of the anomaly metric.

[0071] Perform temporal correction and spatial correction on the mutant anomaly area and the gradual change anomaly area. Temporal correction is to interpolate and correct the anomaly area using the normal data at adjacent times. Spatial correction is to interpolate and correct the anomaly area using the data in the surrounding normal areas. Specifically, the bilinear interpolation method can be used for interpolation calculation.

[0072] Fuse the first result of temporal correction and the second result of spatial correction based on an adaptive weight to obtain a corrected phase consistency metric; the adaptive weight can be dynamically adjusted according to the size and duration of the anomaly area. Exemplarily, for small-area short-time anomalies, the weight of temporal correction can be increased; for large-area long-time anomalies, the weight of spatial correction can be increased.

[0073] Convert the corrected phase consistency metric into local state features through a non-linear mapping. Use the sigmoid function for non-linear mapping to compress the phase consistency metric between 0 and 1. The parameters of the sigmoid function can be adjusted according to the actual application scenario to obtain an appropriate dynamic range.

[0074] Perform weighted aggregation on the local state features to obtain equipment state data. Adopt the method of spatial pyramid pooling to divide the image into grids of multiple scales, perform average pooling within each grid to obtain global features of different scales. Then perform weighted summation on the features of different scales to obtain the final equipment state data. The weights can be set according to the importance of different scales. For example, higher weights can be assigned to features of smaller scales.

[0075] Exemplarily, assume that the size of the input equipment image is 1024×1024 pixels. After performing the Gabor transform, 32 complex response maps are obtained, and each response map is also 1024×1024 in size. The calculated phase consistency metric is also a matrix of size 1024×1024. Perform a two-dimensional Fourier transform on the phase consistency metric to obtain a power spectrum of 1024×1024. Calculate the radial average value to obtain a one-dimensional curve of 512 points.

[0076] Assume that the maximum response is observed at a period of 64 pixels, then the periodic feature of the equipment image is determined to be 64 pixels. The constructed reference phase consistency template is also a matrix of size 1024×1024. The calculated local anomaly metric is also 1024×1024 in size.

[0077] Assume that the mean of the local anomaly metric is 0.1 and the standard deviation is 0.05, then the adaptive threshold is set to 0.2. It is found by comparison that an area of 100×100 pixels exceeds the threshold and is identified as a mutant anomaly area. Calculate the spatial gradient. Assume that the gradient amplitude of an area of 200×200 pixels exceeds the threshold of 0.02 and is identified as a gradual anomaly area.

[0078] Perform temporal and spatial correction on these anomaly areas to obtain a corrected phase consistency metric matrix of 1024×1024. Obtain local state features through non-linear mapping, which is also 1024×1024 in size. Finally, adopt a 3-level spatial pyramid pooling to obtain 21 global features of 1×1, 2×2, and 4×4, and perform weighted aggregation to obtain the final equipment state data.

[0079] In an alternative embodiment, a dynamic graph structure is constructed through a two-stream feature projection network, and the multi-head attention mechanism and temporal consistency constraint are used to perform feature interaction matching on the athlete's action feature vector and the equipment state data. The calculated effective score data includes: Temporally collect the athlete's action feature vector and the equipment state data based on an adaptive sampling rate, and perform dimension unification processing through a two-stream feature projection network constructed by a multi-layer perceptron and a recurrent neural network to obtain action features and equipment features; Construct the action features and the equipment features as node representations of a dynamic graph structure, calculate the edge weight values between the node representations through a sparse attention weight matrix based on locality-sensitive hashing, and construct a weighted graph structure; Generate a query matrix, a key matrix, and a value matrix for the node representations of the graph structure through a multi-scale convolutional network and position encoding, and perform multi-head parallel calculation using a feature gradient-adjusted scaled dot-product attention mechanism, and determine the fused features through residual connection and layer normalization; Calculate the temporal attention weights of the fused features through a gated recurrent unit, calculate the attention coefficients by combining the current node hidden state and the context vector, weight-aggregate the features corresponding to the neighboring nodes in the time series based on the attention coefficients, and update the node hidden state using skip connections; Calculate the similarity matrix between the action features and the equipment features, and calculate the Mahalanobis distance of the similarity matrix at adjacent temporal positions through a sliding window to determine the temporal consistency constraint; Obtain the initial score by non-linearly transforming the similarity matrix and the temporal consistency constraint, and calculate the maximum similarity mean of the initial score in the temporal dimension using the exponential moving average method to determine the score confidence; Based on the initial score and the score confidence, perform anomaly detection and smoothing filtering through an adaptive threshold to generate effective score data.

[0080] In a specific embodiment, temporally collect and feature-project the athlete's action feature vectors and equipment state data, sample the athlete's actions and equipment states using an adaptive sampling rate, and the sampling rate can be dynamically adjusted according to the action change speed. For example, when the action changes violently, the sampling rate is increased to 100 Hz, and when the action is gentle, it is decreased to 10 Hz. The collected raw data is processed by a two-stream feature projection network, which includes two parallel branches: an action feature branch and an equipment feature branch. Each branch is composed of a multi-layer perceptron and a long short-term memory network connected in series. Taking the action feature branch as an example, the original action data first passes through 3 fully connected layers, and the number of hidden layer units is 512, 256, and 128 respectively, and the activation function is ReLU. Then it is input into an LSTM layer with 64 hidden units. Finally, a 128-dimensional action feature vector is output. The equipment feature branch adopts a similar structure and finally outputs a 128-dimensional equipment feature vector.

[0081] Take the action features and equipment features at each time step as nodes in the graph, calculate the similarity between nodes using the locality-sensitive hashing algorithm, and establish edge connections between node pairs with a similarity higher than the threshold of 0.8. The weight value of the edge is determined by the node similarity. For example, assume that at a certain time step, the similarity between the action feature node v1 and the equipment feature node v2 is 0.9, then an edge with a weight of 0.9 is established between v1 and v2.

[0082] Perform multi-head attention calculation on the node representation in the graph structure, extract multi-scale features of nodes through a 3-layer graph convolutional network, and the kernel sizes of the convolutional layers are 3, 2, and 1 respectively. At the same time, encode the positions of the nodes in the graph to obtain position embedding vectors. Concatenate the multi-scale features with the position embedding vectors, and obtain the query matrix Q, key matrix K, and value matrix V through linear transformation, and the dimensions of the matrices are all 128×64. Adopt an 8-head parallel attention mechanism, and the dimension of each attention head is 8. When calculating the attention weights, introduce feature gradient information for adjustment. Specifically, the L2 norm of the feature gradient is used as the scaling factor. Finally, concatenate the outputs of the 8 attention heads, and obtain a 128-dimensional fused feature vector through residual connection and layer normalization.

[0083] To capture temporal dependencies, a gated recurrent unit (GRU) is used to further process the fused features, and the hidden state dimension of the GRU is 128; at each time step, calculate the attention coefficients between the current node's hidden state and the features of the nodes in the adjacent 5 time steps. The attention coefficients are determined by the dot product of the current hidden state and the target node's features, and are normalized by softmax to obtain weights. Weightedly sum the features of the adjacent nodes according to the attention weights, concatenate them with the current hidden state, and pass through a fully connected layer to obtain the updated hidden state.

[0084] In the feature interaction and matching stage, first calculate the similarity matrix between the action features and the equipment features. Using cosine similarity measurement, a T×T-dimensional similarity matrix is obtained, where T is the sequence length. Then apply a sliding window of size 5 in the time dimension to calculate the Mahalanobis distance between adjacent time step similarity vectors as the temporal consistency constraint. The covariance matrix of the Mahalanobis distance is estimated from the statistical data of the past 100 time steps.

[0085] Combine the similarity matrix with the temporal consistency constraint, and perform a non-linear transformation through a two-layer perceptron to obtain the initial score. The number of hidden layer units of the perceptron is 64, the number of output layer units is 1, and the activation function is ReLU. Apply exponential moving average to the initial score sequence with a decay factor of 0.9 to obtain a smoothed score sequence. Calculate the maximum value within each sliding window of length 10 of the smoothed score sequence as the score confidence of the window.

[0086] Generate valid score data based on the initial score and the score confidence. Calculate the mean μ and standard deviation σ of the score sequence, and use μ + 3σ as the adaptive threshold for anomaly detection. Consider the scores exceeding the threshold as outliers and correct them using linear interpolation. Then apply Gaussian filtering for smoothing, and set the standard deviation of the filter to 1.5. The finally output valid score data not only retains the main features of the original scores but also has good temporal consistency.

[0087] In this embodiment, effective feature extraction, dynamic association modeling, and temporal consistency constraints of the athlete's movements and equipment states are achieved, thereby obtaining an accurate and reliable sports performance score; this method can be widely applied to the intelligent scoring systems of various sports events, providing an objective reference basis for referees and coaches.

[0088] In an alternative embodiment, the node representation of the graph structure generates query matrix, key matrix, and value matrix through a multi-scale convolutional network and positional encoding, and uses a scaled dot-product attention mechanism with feature gradient adjustment for multi-head parallel computing, and determines the fused feature through residual connection and layer normalization, including: Extract features of the node representation of the graph structure through a multi-scale convolutional network, and perform feature concatenation to obtain a multi-scale feature representation; Introduce positional encoding, dynamically adjust the positional encoding through an adaptive weight matrix to obtain enhanced positional encoding, and perform feature enhancement operation on the enhanced positional encoding and the multi-scale feature representation to obtain enhanced features; Based on the enhanced features, generate query matrix, key matrix, and value matrix through query weight matrix, key weight matrix, and value weight matrix respectively, calculate the feature gradient of the enhanced features, and input the feature gradient into a feature adjustment network to obtain an attention adjustment factor; Perform multi-head parallel computing on the query matrix, key matrix, and value matrix through the scaled dot-product attention mechanism to obtain multiple attention head features, dynamically adjust the multiple attention head features through the attention adjustment factor to obtain an attention output feature, and perform residual connection on the attention output feature and the multi-scale feature representation to obtain a feature fusion result; Input the feature fusion result into a layer normalization network to generate a normalization coefficient and a bias parameter, and process to obtain the final fused feature.

[0089] In this embodiment, convolutional kernels of different scales are used to perform convolutional operations on node features. For example, convolutional kernels of 1×1, 3×3, and 5×5 scales can be used. Each scale of convolutional kernel outputs a feature map, and then these feature maps are concatenated in the channel dimension to obtain a multi-scale feature representation. Assume that the input node feature dimension is 64, and 3 scales of convolutional kernels are used, and each convolutional kernel outputs 32 channels, then the final multi-scale feature representation dimension is 96.

[0090] Introduce positional encoding and perform dynamic adjustment to generate initial positional encoding. The sine positional encoding method can be used to generate a positional encoding vector with the same number of nodes as the number of nodes. Then, perform a linear transformation on the positional encoding through a learnable weight matrix to obtain enhanced positional encoding. Exemplarily, the dimension of the initial positional encoding is 64, and the 32-dimensional enhanced positional encoding is obtained through transformation by a 32×64 weight matrix; perform an element-wise addition operation on the enhanced positional encoding and the multi-scale feature representation to obtain enhanced features.

[0091] Use three independent linear transformation layers (i.e., fully connected layers) as the query weight matrix, key weight matrix, and value weight matrix respectively. Input the enhanced features into these three linear transformation layers to obtain the query matrix, key matrix, and value matrix. Assume that the dimension of the enhanced features is 96, and it is desired to obtain 8 attention heads, each with a dimension of 12. Then, the output dimensions of these three linear transformation layers are all 96.

[0092] Obtain the feature gradient by taking the derivative of the enhanced features. Input the feature gradient into a feature adjustment network, which can be a two-layer multi-layer perceptron, and output an attention adjustment factor with the same number of dimensions as the number of attention heads. For example, the dimension of the feature gradient is 96, and an 8-dimensional attention adjustment factor is obtained through a two-layer fully connected network of 96→32→8.

[0093] Divide the query matrix, key matrix, and value matrix into 8 heads respectively, each with a dimension of 12. For each head, calculate the dot product of the query matrix and the key matrix, divide the result by the square root of 12 for scaling, and then obtain the attention weights through the softmax function. Multiply the attention weights by the value matrix to obtain the attention output of each head. Finally, concatenate the outputs of the 8 heads to obtain a 96-dimensional multi-head attention feature.

[0094] Dynamically adjust the multi-head attention feature using the obtained attention adjustment factor. Multiply the 8-dimensional attention adjustment factor element-wise with the outputs of the 8 attention heads to achieve dynamic adjustment of each attention head. The adjusted multi-head attention feature is the attention output feature.

[0095] Perform a residual connection between the attention output feature and the multi-scale feature representation. Specifically, perform an element-wise addition of the 96-dimensional attention output feature and the 96-dimensional multi-scale feature representation to obtain the feature fusion result.

[0096] Input the feature fusion result into a layer normalization network. The layer normalization network contains two learnable parameters: a scaling parameter and a bias parameter, both of which are 96-dimensional vectors. First, perform normalization with a mean of 0 and a variance of 1 on the feature fusion result, then multiply it element-wise with the scaling parameter, and then add the bias parameter to obtain the final fused feature.

[0097] Exemplarily, assume that the input graph structure has 100 nodes, and the initial feature dimension of each node is 64. First, through a multi-scale convolutional network, using 1×1, 3×3, and 5×5 convolutional kernels, each convolutional kernel outputs 32 channels, obtaining a 96-dimensional multi-scale feature representation.

[0098] Generate an initial position encoding of 100×64, and obtain an enhanced position encoding of 100×32 through transformation by a 32×64 weight matrix. Add the enhanced position encoding to the multi-scale feature representation to obtain an enhanced feature of 100×96.

[0099] Use three linear transformation layers of 96×96 to transform the enhanced feature into a query matrix, a key matrix, and a value matrix, and the dimension of each matrix is 100×96. At the same time, calculate the gradient of the enhanced feature, and through a two-layer fully connected network of 96→32→8, obtain an 8-dimensional attention adjustment factor.

[0100] Divide the query matrix, the key matrix, and the value matrix into 8 heads, and the dimension of each head is 12. Perform scaled dot-product attention calculation on each head to obtain 8 attention head features of 100×12. Concatenate the features of these 8 heads to obtain a multi-head attention feature of 100×96.

[0101] Use the 8-dimensional attention adjustment factor to adjust the 8 attention heads to obtain the adjusted attention output feature. Perform a residual connection on this feature and the multi-scale feature representation to obtain a feature fusion result of 100×96.

[0102] Use two learnable parameters of 96 dimensions to perform layer normalization on the feature fusion result to obtain the final fused feature of 100×96 dimensions.

[0103] As Figure 4 shown, the complete processing process from the graph structure node representation to the final fused feature is described in detail. The input graph structure node representations enter the multi-scale convolutional network and the position encoding processing module respectively. The multi-scale convolutional network extracts features in parallel through convolutional layers of small, medium, and large scales, and generates a multi-scale feature representation through a feature concatenation operation; at the same time, the position encoding processing module generates a position encoding and dynamically adjusts it through an adaptive weight matrix to form an enhanced position encoding. These two paths of features are combined through a feature enhancement operation to form an enhanced feature containing multi-scale spatial information and position information.

[0104] The enhanced features are then fed into three parallel modules: the query-key-value matrix generation module generates the query matrix, key matrix, and value matrix required by the attention mechanism; the feature gradient adjustment module calculates the feature gradient and generates an attention adjustment factor through the feature adjustment network; the multi-head attention mechanism module performs scaled dot-product attention calculation. Among them, the query-key-value matrix is directly input into the multi-head attention mechanism to maintain the direct transmission of the information flow. The attention head features obtained by the multi-head attention calculation are combined with the attention adjustment factor for dynamic feature adjustment to generate the attention output features.

[0105] The attention output features are subjected to residual connection with the original multi-scale feature representation and processed by layer normalization, and finally the fused features are output. This design makes full use of multi-scale spatial information, position information, and self-attention mechanism, and at the same time enhances the model's perception ability of key information through the dynamic adjustment of feature gradients, which is an efficient graph structure feature processing framework.

[0106] Traditional graph structure feature extraction mainly relies on models such as graph convolutional networks (GCNs) and graph attention networks (GATs). At the same time, the powerful representation ability demonstrated by the self-attention mechanism in the Transformer architecture provides a new idea for graph structure data processing.

[0107] Existing graph structure feature processing technologies have some limitations. Traditional GCNs mainly adopt convolutional operations with a single receptive field and cannot capture structural information at different scales simultaneously, resulting in limited expressive ability for complex graph structures. Existing methods usually adopt fixed position encoding methods, such as Laplacian eigenvectors or random walk encodings, lacking the ability to adaptively adjust according to specific tasks. Traditional graph attention usually adopts simple weighted summation or attention calculation with fixed parameters, lacking dynamic adjustment of feature importance and having limited ability to capture key information. In addition, existing methods often adopt simple feature concatenation or weighted fusion methods, lacking the ability of multi-path parallel processing and dynamic feature adjustment.

[0108] Aiming at the limitations of the existing technologies, the method of this embodiment introduces convolutional layers of three different scales, small, medium, and large, to extract features in parallel, generates a multi-scale feature representation through feature concatenation operations, and effectively captures structural information in different ranges. The position encoding is dynamically adjusted through an adaptive weight matrix, enabling the position information to be adaptively adjusted according to specific tasks and data characteristics, enhancing the model's understanding ability of the topological structure. The feature gradient calculation and feature adjustment network are introduced to generate an attention adjustment factor to dynamically adjust the attention head features, enhancing the model's perception ability of important features. A multi-module parallel processing architecture is adopted, and the original feature information is retained through residual connection, reducing the difficulty of deep network training and effectively preventing information loss.

[0109] There are correlations at different scales in the graph-structured data, and it is necessary to capture both local and global features simultaneously. The importance of node position information varies in different tasks, and it is necessary to dynamically adjust the influence of position encoding according to specific tasks. Different features in the graph structure contribute differently to the prediction results, and it is necessary to dynamically identify important features through gradient information. Deep neural networks are prone to information attenuation, and it is necessary to ensure the effective transmission of key information through direct connections and parallel processing.

[0110] This solution has achieved a significant improvement in performance compared with the prior art. The combination of multi-scale convolution and adaptive position encoding has greatly enhanced the expressive ability of complex graph structures and can capture structural information at different scales simultaneously. The attention mechanism adjusted by feature gradients can accurately identify and strengthen important features, filter redundant information, and improve the discriminative ability of the model. The design of the parallel processing architecture and direct connections reduces computational bottlenecks and improves the efficiency of feature extraction and fusion. Adaptive position encoding and dynamic feature adjustment enable the model to better adapt to different types of graph-structured data and enhance the generalization ability. The application of residual connections and layer normalization effectively alleviates the problem of gradient disappearance in deep neural networks, making model training more stable and efficient.

[0111] The real-time score statistics system for sports events based on AI visual recognition in the embodiments of the present invention includes: The first unit is used to collect image data of the sports event venue through multiple cameras; perform object detection on the image data to obtain athlete image data, equipment image data, and scoring area image data; The second unit is used to generate athlete joint heat maps from the athlete image data by using a feature pyramid network, construct a dynamic connection skeleton graph structure through temporal smoothing constraints, and extract athlete action feature vectors through a graph convolutional network and a recurrent neural ordinary differential equation; The third unit is used to perform periodic detection and anomaly correction on the equipment image data by using a phase consistency metric network to obtain equipment status data; The fourth unit is used to construct a dynamic graph structure through a two-stream feature projection network, perform feature interaction matching on the athlete action feature vectors and the equipment status data by using a multi-head attention mechanism and temporal consistency constraints, and calculate and obtain effective score data; The fifth unit is used to generate a score data packet containing the scoring time, the scoring athlete number, the scoring type, and the scoring video segment according to the effective score data; The sixth unit is used to send the score data packet to the game scoring system and store the scoring video segment in the video storage server.

[0112] In the third aspect of the embodiments of the present invention, There is provided an electronic device, including: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0113] In a fourth aspect of the embodiments of the present invention, A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0114] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present invention.

[0115] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A real-time score statistics method for sports competitions based on AI visual recognition, characterized in that, Including: Collecting image data of a sports event venue through multiple cameras; Performing object detection on the image data to obtain athlete image data, equipment image data, and scoring area image data; Using a feature pyramid network to generate athlete joint heatmaps from the athlete image data, constructing a dynamic connection skeleton graph structure through temporal smoothing constraints, and extracting athlete action feature vectors through a graph convolutional network and a recurrent neural ordinary differential equation; Using a phase consistency metric network to perform periodic detection and anomaly correction on the equipment image data to obtain equipment status data; Constructing a dynamic graph structure through a two-stream feature projection network, and performing feature interaction matching on the athlete action feature vectors and the equipment status data using a multi-head attention mechanism and temporal consistency constraints to calculate effective score data; Generating a score data packet including the scoring time, the athlete number of the scorer, the scoring type, and the scoring video segment based on the effective score data; Sending the score data packet to the game scoring system and storing the scoring video segment in a video storage server.

2. The method according to claim 1, wherein Using a feature pyramid network to generate athlete joint heatmaps from the athlete image data, constructing a dynamic connection skeleton graph structure through temporal smoothing constraints, and extracting athlete action feature vectors through a graph convolutional network and a recurrent neural ordinary differential equation includes: Performing downsampling and feature mapping on the athlete image data through the feature pyramid network to obtain feature branches at different scales; mapping the feature branches through convolutional operations and activation functions to determine the spatial position distribution of the joints and generate athlete joint heatmaps; Performing weighted integral calculation on the response values of each joint point in the athlete joint heatmap to obtain joint spatial coordinates, and performing smoothing constraint calculation in combination with the coordinate information of adjacent temporal frames to obtain a joint smooth coordinate sequence; Establishing a node set based on the joint smooth coordinate sequence, calculating the node connection probability through the temporal displacement change between nodes, and forming a dynamic connection skeleton graph structure; Inputting the dynamic connection skeleton graph structure into a multi-layer stacked graph convolutional network, normalizing the node features and updating the weights in each layer of the network, calculating the difference in velocity vectors and position vectors between nodes, and obtaining node motion correlation features; Decomposing and mapping the node motion correlation features to obtain multiple basic motion representation vectors, and calculating the corresponding combined weight coefficients to fuse and form a global action feature vector; Inputting the global action feature vector into a recurrent neural ordinary differential equation for temporal modeling, extracting multi-scale temporal features through multi-layer dilated convolutions, and outputting athlete action feature vectors.

3. The method according to claim 2, wherein Establishing a node set based on the joint smooth coordinate sequence, calculating the node connection probability through the temporal displacement change between nodes, and forming a dynamic connection skeleton graph structure includes: Constructing the joint smooth coordinate sequence into an initial node set, calculating the temporal displacement difference between adjacent nodes in the initial node set, and constructing a node state vector; Establishing a Bayesian network model based on the node state vector, and encoding the spatial position relationship and temporal displacement difference between nodes as a conditional probability distribution; Construct a Markov random field to describe the spatial constraint relationships of nodes. The spatial constraint relationships include node distance constraint, angle constraint, and topological connection constraint, and generate a spatial potential function of the nodes; Combine the conditional probability distribution and the spatial potential function, and use the variational expectation maximization algorithm to iteratively calculate the node connection probability; Construct a probability graph network based on the node connection probability, calculate the mutual information and conditional entropy between each pair of nodes, use the weighted sum of the mutual information and the conditional entropy as the connection uncertainty, and set the node connection threshold according to the connection uncertainty; Determine the connection relationship between nodes according to the node connection threshold, and perform online update on the connection relationship based on the temporal displacement difference between the nodes to form a dynamic connection skeleton graph structure.

4. The method according to claim 3, wherein Combining the conditional probability distribution and the spatial potential function, using the variational expectation maximization algorithm to iteratively calculate the node connection probability includes: Multiply the conditional probability distribution by the spatial potential function and normalize it to construct a joint probability distribution model of the node state and the connection relationship; Construct an auxiliary probability distribution to represent the connection relationship between nodes, and use the auxiliary probability distribution as a variational distribution to approximate the posterior distribution of the joint probability distribution model; Construct a variational objective function based on the joint probability distribution model and the variational distribution. The variational objective function includes an observation data log-likelihood term, a KL divergence term of the posterior distribution, and a spatial potential function term; Perform expectation maximization iterative optimization on the variational objective function. In the expectation step of each iteration, fix the model parameters to optimize the variational distribution. In the maximization step, fix the variational distribution, obtain an adaptive learning rate by calculating the first moment and the second moment of the gradient, use the product of the adaptive learning rate and the gradient as the parameter update amount and perform truncation processing, optimize the model parameters, and determine the optimized variational distribution; Calculate the initial connection probability between node pairs based on the optimized variational distribution, and perform softmax normalization in combination with the temperature parameter to obtain the normalized connection probability; Monitor the difference of the variational objective function between adjacent iterations. When the difference is less than a preset difference threshold, determine that the iteration converges, and output the final node connection probability.

5. The method according to claim 1, wherein Use a phase consistency metric network to perform periodic detection and anomaly correction on the equipment image data to obtain equipment status data including: Perform multi-directional and multi-scale Gabor transforms on the equipment image data to obtain the complex response of the equipment image data, calculate the local phase value based on the complex response, and calculate the phase consistency metric in combination with the amplitude of the complex response and the local phase value; Perform a two-dimensional Fourier transform on the phase consistency metric to obtain a power spectrum, calculate the radial average value of the power spectrum, and determine the periodic characteristics of the equipment image data based on the maximum response of the radial average value; Construct a reference phase consistency template, calculate the difference between the phase consistency metric and the reference phase consistency template as the local anomaly metric, and construct an adaptive threshold based on the mean and standard deviation of the local anomaly metric; Compare the local anomaly metric with the adaptive threshold to identify mutant anomaly regions, calculate the spatial gradient of the local anomaly metric and compare it with a preset gradient threshold to identify gradual anomaly regions; Perform temporal correction and spatial correction on the mutant anomaly regions and the gradual anomaly regions, and fuse the first result of temporal correction and the second result of spatial correction based on adaptive weights to obtain a corrected phase consistency metric; Convert the corrected phase consistency metric into local state features through non-linear mapping, and perform weighted aggregation on the local state features to obtain equipment state data.

6. The method according to claim 1, characterized in that Construct a dynamic graph structure through a two-stream feature projection network, and use the multi-head attention mechanism and temporal consistency constraints to perform feature interaction matching on the athlete's action feature vector and the equipment state data, and calculate the effective score data including: Temporally collect the athlete's action feature vector and the equipment state data based on an adaptive sampling rate, and perform dimension unification processing through a two-stream feature projection network constructed by a multi-layer perceptron and a recurrent neural network to obtain action features and equipment features; Construct the action features and the equipment features as node representations of a dynamic graph structure, calculate the edge weight values between the node representations through a sparse attention weight matrix based on locality-sensitive hashing, and construct a weighted graph structure; Generate a query matrix, a key matrix, and a value matrix for the node representations of the graph structure through a multi-scale convolutional network and position encoding, perform multi-head parallel calculations using a scaled dot-product attention mechanism adjusted by feature gradients, and determine the fused features through residual connections and layer normalization; Calculate the temporal attention weights of the fused features through a gated recurrent unit, calculate the attention coefficients by combining the current node hidden state and the context vector, perform weighted aggregation on the features corresponding to the temporal neighborhood nodes based on the attention coefficients, and update the node hidden state using skip connections; Calculate the similarity matrix between the action features and the equipment features, and calculate the Mahalanobis distance of the similarity matrix at adjacent temporal positions through a sliding window to determine the temporal consistency constraint; Non-linearly transform the similarity matrix and the temporal consistency constraint to obtain an initial score, and use the exponential moving average method to calculate the maximum similarity mean of the initial score in the temporal dimension to determine the score confidence; Based on the initial score and the score confidence, perform anomaly detection and smoothing filtering through an adaptive threshold to generate effective score data.

7. The method according to claim 6, wherein Generate a query matrix, a key matrix, and a value matrix for the node representations of the graph structure through a multi-scale convolutional network and position encoding, perform multi-head parallel calculations using a scaled dot-product attention mechanism adjusted by feature gradients, and determine the fused features including: Extract features from the node representations of the graph structure through a multi-scale convolutional network, and perform feature concatenation to obtain a multi-scale feature representation; Introduce position encoding, dynamically adjust the position encoding through an adaptive weight matrix to obtain enhanced position encoding, and perform feature enhancement operations on the enhanced position encoding and the multi-scale feature representation to obtain enhanced features; Based on the enhanced features, generate a query matrix, a key matrix, and a value matrix by querying the weight matrix, the key weight matrix, and the value weight matrix respectively, calculate the feature gradient of the enhanced features, and input the feature gradient into a feature adjustment network to obtain an attention adjustment factor; Perform multi-head parallel calculations on the query matrix, the key matrix, and the value matrix through the scaled dot-product attention mechanism to obtain multiple attention head features, dynamically adjust the multiple attention head features through the attention adjustment factor to obtain an attention output feature, and perform a residual connection on the attention output feature and the multi-scale feature representation to obtain a feature fusion result; Input the feature fusion result into a layer normalization network to generate a normalization coefficient and a bias parameter, and process to obtain a final fusion feature.

8. A real-time score statistics system for sports competitions based on AI vision recognition, which is used to implement the method described in any one of the foregoing claims 1-7, characterized in that, Comprising: A first unit for collecting image data of a sports event venue through multiple cameras; Perform object detection on the image data to obtain athlete image data, equipment image data, and scoring area image data; A second unit for generating an athlete joint heat map for the athlete image data using a feature pyramid network, constructing a dynamic connection skeleton graph structure through temporal smoothing constraints, and extracting an athlete action feature vector through a graph convolutional network and a recurrent neural ordinary differential equation; A third unit for performing periodic detection and anomaly correction on the equipment image data using a phase consistency metric network to obtain equipment status data; A fourth unit for constructing a dynamic graph structure through a two-stream feature projection network, and performing feature interaction matching on the athlete action feature vector and the equipment status data using a multi-head attention mechanism and temporal consistency constraints to calculate effective score data; A fifth unit for generating a score data packet containing the scoring time, the athlete number, the scoring type, and the scoring video segment according to the effective score data; A sixth unit for sending the score data packet to the game scoring system and storing the scoring video segment in a video storage server.

9. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Action prediction method and system based on lattice point optical flow

    CN115100559A

  • Behavior posture recognition analysis method and system

    CN119181137A

  • Basketball goal number calculation method based on computer vision

    CN119228850A

  • System and methods for providing a user key performance indicators for basketball

    US20200009443A1

  • Methods and apparatus for team classification in sports analysis

    US20240005701A1

Cited By

  • Indoor component intelligent configuration method and device, equipment and storage medium

    CN122133239A

  • Indoor component intelligent configuration method and device, equipment and storage medium

    CN122133239B

  • Automatic scoring system of racket sports competitions

    TWI919997B