Curling sports acquisition perception and strategy prediction method applied to smart game watching

By combining a multi-camera fusion visual large model and improved self-game reinforcement learning, with a global Q network and expert strategy, the accuracy issues of positioning tracking and strategy prediction in curling competitions were solved, achieving efficient, accurate and stable curling sports acquisition, perception and strategy prediction.

CN120726541AActive Publication Date: 2025-09-30SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511187215.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-09-30
Estimated Expiration
2045-08-25

AI Technical Summary

Technical Problem

Existing curling game positioning and tracking solutions lack accuracy when changing perspectives and venues, and existing strategy methods lack a global awareness of the entire game, resulting in a large difference between predicted strategies and actual strategies, and cannot effectively support smart viewing.

Method used

A multi-camera fusion vision large model is used for curling target segmentation and coordinate positioning. Combined with improved self-game reinforcement learning and Monte Carlo tree search, a strategy prediction model that integrates local and global information is developed. A global Q network is introduced to optimize strategy generation, and expert strategies are combined to improve prediction accuracy.

Benefits of technology

The accuracy of curling image recognition and game prediction has been improved, especially in the strategic predictions of the last few rounds of the game, which are more consistent and stable, with strong adaptability and support for smart viewing needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726541A_ABST
    Figure CN120726541A_ABST
Patent Text Reader

Abstract

The invention relates to a curling movement acquisition perception and strategy prediction method applied to smart game watching, and the method comprises the following steps: fusing the same time frame image in a curling throwing video obtained by cameras at a plurality of positions, and obtaining the current frame curling observation coordinate of each curling target; performing track matching on each curling target based on a video frame sequence of the curling video to obtain motion tracks of all curling in a current curling throwing process, and taking the motion tracks as curling motion acquisition perception results; and obtaining a current actual curling state, and obtaining a maximum possible action corresponding to the current actual curling state as a strategy prediction result in combination with an expert strategy method and the improved curling strategy prediction model. Compared with the prior art, the method has the advantages that the prediction accuracy of the last few curling matches of the whole game is improved, the curling image recognition accuracy is improved, and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of curling prediction and analysis, and in particular to a curling sports collection, perception, and strategy prediction method for smart game viewing. Background Art

[0002] Among current curling prediction schemes, CN202111107286.X (A Curling On-Site Analysis and Decision Support Method Based on Reinforcement Learning) discloses a curling on-site analysis and decision support method based on reinforcement learning. This method constructs a digital model of the curling game scenario and situation and establishes a corresponding decision support system. The system primarily consists of three modules: a curling game situation awareness module that acquires real-time information such as the position, speed, and resting state of the curling stones; a curling field digital extraction module that accurately determines the position and category of the curling stones at critical moments by mapping the field position to the captured data; and a curling game decision analysis module that uses this information to simulate and calculate using a reinforcement learning algorithm, outputting a recommended next shot position to assist in tactical decision-making.

[0003] CN202010435770.4 (A Curling Ball Motion State Estimation Method) discloses a curling ball motion state estimation method based on image processing, focusing on improving the accuracy of curling ball state and handle angle perception. The curling ball is detected and tracked using a curling ball target detection network, a rotation angle detection network, and a target tracking network to obtain the center coordinates of the curling ball. Finally, through coordinate transformation, the center coordinates and rotation angle in the image coordinate system are converted to the curling ball coordinates and rotation angle in the field coordinate system.

[0004] CN202110774457.8 (A Curling Strategy Generation Method Based on Monte Carlo Reinforcement Learning) discloses a curling strategy generation method based on Monte Carlo reinforcement learning, addressing the difficulty of obtaining valid datasets for strategy decision-making. This method utilizes a policy value network model and a value network to obtain initial actions, combined with an improved Monte Carlo tree search algorithm. Data from the self-play phase is then used to further update the policy value network.

[0005] CN202410066718.4 (A curling shot automated learning and decision search method) discloses a curling shot decision search method based on latent space. This method maps the actual game state into a low-dimensional latent state representation, uses Monte Carlo tree search to search for strategies in the latent space, and finally calculates the optimal strategy based on node value and visit count.

[0006] However, there are two main problems in the existing technology: (1) Existing vision-based solutions for curling positioning and tracking have the following main drawbacks: a. Existing technologies all use relatively traditional object detection algorithms, which have limited accuracy and generalization. When applied to new curling venues, due to changes in camera position and angle, a large amount of image data needs to be re-collected and annotated.

[0007] b. Existing methods typically use object detection networks like YOLO, which can only locate the target's bounding box. However, due to the large size of the curling stone (approximately 30 cm in diameter), the position of the bounding box varies significantly in images from different perspectives. For example, when viewing the curling stone from an angle, the center of the bounding box is not exactly at the center of the stone. Therefore, directly inferring the center coordinates of the curling stone from the detected bounding box will result in significant error. These technologies only achieve high accuracy when the camera is viewing the curling stone vertically from above, which places high demands on the number of cameras and their installation locations.

[0008] When it comes to curling strategy generation, existing technologies have the following shortcomings: a. Inadequate long-term strategy modeling. Existing strategy methods, during training and inference, only consider the outcome of the current round, not the overall match. Consequently, the predicted strategies for the final few bowls of the last two rounds often differ significantly from the strategies actually employed by the players. For example, in the final bowl of the last round, with the backhand trailing by 3 points, the players will typically adopt an aggressive strategy to secure a throw of ≥3 points to win the match. However, existing strategies lack a holistic view of the entire match, resulting in strategies that often win the current round by 1 or 2 points but lose the entire match.

[0009] b. Existing strategic approaches rely solely on data-driven algorithms, which may be limited to local optimal solutions. The reasoning process does not support the integration of commentators and coaches' judgments on the current situation combined with their professional knowledge, and cannot effectively support the actual needs of commentators in selecting and predicting tactical types during smart viewing. Summary of the Invention

[0010] The purpose of the present invention is to provide a curling motion acquisition, perception and strategy prediction method for smart viewing in order to improve the accuracy of curling game predictions for the last few rounds of the entire game and to improve the accuracy of curling image recognition.

[0011] The purpose of the present invention can be achieved by the following technical solutions: A curling sports collection, perception, and strategy prediction method for smart game viewing includes the following steps: S1. Fusing the same time frame images of a curling throw video obtained by cameras at multiple locations to obtain the current frame curling observation coordinates of each curling target, performing trajectory matching on each curling target based on the video frame sequence of the curling video, and obtaining the motion trajectories of all curling stones in the current curling throw as the curling motion acquisition perception result; S2. Obtain the current actual curling state, combine the expert strategy method and the improved curling strategy prediction model, obtain the maximum possible action corresponding to the current actual curling state as the strategy prediction result, and render the strategy prediction result; The training process of the improved curling strategy prediction model is as follows: Construct an initial policy network and an initial value network, obtain training data, and perform supervised training on the policy network and the value network based on the training data to obtain the first policy network and the first value network; Based on the improved self-game reinforcement learning and Monte Carlo tree search method, the first strategy network and the first value network were trained again to obtain an improved curling strategy prediction model.

[0012] Furthermore, the training data includes training label data {(TS t , Ta t , Tz t )}, the training label data includes multiple game label data, each game label data includes multiple round label data, each round label data includes all the states of a round of curling game and the final score Tz of the round of curling game t , all states of a curling game include the training state of each curling throw, and the training state TS of the t-th curling throw in all states of a curling game t Transition to the training state TS for the t+1th curling throw t+1 The corresponding training action is Ta t , supervised training is performed on the policy network and the value network based on the training label data to obtain the first policy network and the first value network.

[0013] Furthermore, the specific steps of fusing the same time frame images of a curling throw video obtained by cameras at multiple positions to obtain the current frame curling observation coordinates of each curling target are as follows: Multiple curling targets are identified based on the same time frame images in a curling throwing video obtained by cameras at multiple positions, and the same time frame segmentation maps of each curling target from multiple perspectives are obtained. The current frame actual coordinates of the curling on the track from each perspective are determined based on the same time frame segmentation maps of each curling target from each perspective, and the current frame actual coordinates of the same curling target from various perspectives are fused to obtain the current frame curling observation coordinates of each curling target.

[0014] Furthermore, the specific steps of performing trajectory matching on each curling target based on the frame sequence of the curling video and obtaining the motion trajectories of all curling stones are as follows: For the same video frame t, assume that a curling targets are identified, and the current frame curling observation coordinates of the a curling targets and the state vectors of the a curling targets in video frame t are obtained. The state vector of video frame t+1 is predicted by using the state vector of video frame t, the current frame curling observation coordinates of video frame t, and the Kalman filter state transition model to obtain the predicted coordinates of the a curling targets in video frame t+1. The Euclidean distance between the current frame curling observation coordinates of the curling target in video frame t+1 and the predicted coordinates of the curling target in video frame t+1 is calculated. The Hungarian algorithm is used to perform optimal matching based on the Euclidean distance. The current frame curling observation coordinates of the curling target in video frame t+1 that is successfully matched are updated to the motion trajectory sequence of the corresponding curling target, and the state vector of the curling target in video frame t+1 is obtained. The current frame curling observation coordinates of the curling target in video frame t+1 that fails to match are used as the newly appeared curling target. The state vector of the newly appeared curling target in video frame t+1 is initialized, and the video frame is updated to t=t+1. The above steps are repeated to obtain the motion trajectories of all curling stones in the current curling throw.

[0015] Furthermore, the first strategy network and the first value network are retrained based on the improved self-game reinforcement learning and Monte Carlo tree search method to obtain the improved curling strategy prediction model. The specific steps are as follows: A1. The initial curling stone state is used as the current curling stone state. A2. Input the current curling state into the first policy network and the first value network respectively to obtain the probability distribution of the action corresponding to the current curling state. And the probability distribution of scoring in one game v t , the probability distribution of actions Sample a set of discrete actions As the base action set for the current curling state ={ }, Represents the 1st, 2nd…mth discrete action of the base action set, m represents the total number of discrete actions, and the discrete action is obtained The corresponding probabilities P t ; A3. Set the current curling state and basic actions , probability P t And the probability distribution of scoring in one game v t As the root node of the Monte Carlo tree, after selecting, expanding, simulating and backtracking in sequence, the cumulative reward and number of visits of the node of the Monte Carlo tree are updated; A4. Use the normalized number of visits to the child nodes of the root node as the first action probability distribution; Sampling the first action probability distribution to obtain a first action corresponding to the current curling state, executing the first action based on the current curling state, using the obtained state as the new current curling state, and returning to A2 until the current round ends, where the end of the current round indicates that the current round has reached the throw count threshold. A5. Integrate the current state of each curling stone, the first action corresponding to each current state of each curling stone, the cumulative score from the first round to the current round, the number of rounds remaining in the current game up to the current round, and the win / loss status of the second player in the first round in the current game as the global Q network training set; A6. Use the first action corresponding to each current curling state as a label for the first policy network, and the current round result at the end of the current round as a label for the first value network for training, to obtain a second policy network and a second value network. A7. Build a global Q network and train the global Q network based on the global Q network training set to obtain a trained Q network. A8. Using the initial curling state as the current curling state; A9. Input the current curling state into the second policy network and the second value network respectively to obtain the probability distribution of the action corresponding to the current curling state. , and the probability distribution of scoring in one round v t ', the probability distribution of the action Sample a set of discrete actions , ,……, As the base action set for the current curling state , ,……, }, , ,……, Indicates the 1st, 2nd, ... A discrete action, Represents the total number of discrete actions and obtains discrete actions , ,……, The corresponding probabilities P t '; A10. Set the current curling state and basic actions , probability P t ' and the probability distribution of one-game score v t ' As the root node of the Monte Carlo tree, after selecting, expanding, simulating and backtracking in sequence, the cumulative reward and number of visits of the node of the Monte Carlo tree are updated; A11. Take the normalized access count of the child nodes of the root node as the second probability distribution , and sample a set of sampling actions corresponding to the current curling state without replacement from the second probability distribution Denote the 1st, 2nd,... k sampling actions, and the probability corresponding to each sampling action , where i = 1, 2,... k , k denotes the total number of sampling actions; A12. Input the current curling state, k sampling actions, the cumulative score from the first inning to the current inning corresponding to the current curling state, and the remaining innings in the current game corresponding to the current curling state into the trained Q-network, and obtain the probability of winning or losing for the second player in the first inning of the current game corresponding to each sampling action , and calculate the probability of integrating local and global information based on the probability of winning or losing and the probability corresponding to each sampling action , and on the basis of the current curling state, execute the sampling action with the maximum probability of integrating local and global information , and use the obtained state as the new current curling state, and return to A9 until the end of the current inning; A13. Integrate the various current curling states, the probabilities of integrating local and global information respectively corresponding to the various current curling states the maximum actions, the cumulative scores from the first inning to the current inning corresponding to the various current curling states, the remaining innings in the current game corresponding to the various current curling states, and the winning or losing situation of the second player in the first inning of the current game as the global Q-network fine-tuning dataset, and fine-tune the trained global Q-network based on the fine-tuning dataset to obtain the fine-tuned global Q-network<ooo0159>

[0016] Furthermore, the specific steps to obtain the current actual curling state and combine the expert strategy method and the improved curling strategy prediction model to obtain the most likely action corresponding to the current actual curling state are as follows: )]] B1. Input the current actual curling state into the second policy network and the second value network respectively. The second policy network and the second value network respectively output the probability distribution of the actions corresponding to the current actual curling state and the probability distribution of the score of one inning corresponding to the current actual curling state B2. Sample a set of actual actions and the probabilities corresponding to the actual actions from the probability distribution of the actions corresponding to the current actual curling state, and select f (f < e) expert strategies from the predefined e expert strategies , fIndicates the 1st, ..., f expert strategies, one expert strategy represents one action, and the normalized credibility of each expert strategy is obtained at the same time; B3. The expert strategy and a set of actual actions constitute the actual discrete action set. The probability corresponding to the actual action and the normalized credibility of each expert strategy constitute the probability corresponding to the actual discrete action set. B4. Use the current actual curling state, the actual discrete action set, the probability corresponding to the actual discrete action set, and the probability distribution of the score in one round corresponding to the current actual curling state as the root node of the Monte Carlo tree. After selecting, expanding, simulating, and backtracking in sequence, update the cumulative reward and visit count of the node in the Monte Carlo tree; B5. Take the normalized number of visits to the child nodes of the root node as the first probability distribution ', sample a set of sample actions corresponding to the current actual curling state from the first probability distribution without replacement Indicates the 1st, 2nd, ... k The actual actions sampled, and the probability corresponding to each actual action sampled ,in i =1, 2…… k , k Indicates the total number of actual actions sampled; B6. Input the current actual curling state, k sampled actual actions, the cumulative scores from the first game to the current game corresponding to the current actual curling state, and the number of games remaining in the current game corresponding to the current actual curling state up to the current game into the fine-tuned global Q network. The fine-tuned global Q network outputs the probability of the second player in the first game corresponding to each sampled actual action winning or losing in the current game, calculates the actual probability of fusing local and global information, and takes the sampled actual action with the largest actual probability as the most likely action corresponding to the current actual curling state.

[0017] Furthermore, the probability of fusing local and global information is:

[0018] in, represents the probability of fusing local and global information, Indicates the i Sampling action, is the dynamic coefficient.

[0019] Furthermore, the dynamic coefficient is:

[0020] in, is a hyperparameter, c is half of the total number of rounds in a curling game, and n tThe number of remaining rounds in the current game corresponding to the current curling state.

[0021] Furthermore, the state space corresponding to the state is the position of all curling stones in sequence and the position of the base camp, whether the current curling stone is thrown by the second team in the first round, which throw the current curling stone is in this round, and the current round number.

[0022] Furthermore, the action space corresponding to the action is the target position of the curling stone and whether to rotate clockwise or counterclockwise to reach the target position.

[0023] Compared with the prior art, the present invention has the following beneficial effects: The improved self-playing reinforcement learning algorithm designed by the present invention incorporates a global Q-network to enhance the algorithm's long-term strategy modeling. Based on dynamic coefficients, it effectively balances the generated strategy's individual round wins and losses with the overall match score at different stages of the game. This integrates local and global information. Through the dynamic coefficients, the generated strategy focuses more on individual round scores in the early stages of the game (when many rounds remain) and more on the overall match score in the later stages (when fewer rounds remain). This dynamic adjustment mechanism, integrating local and global information, enables the present invention to combine the flexibility and targeted nature of single-round decision-making at the micro level with the stability and goal-oriented nature of overall match strategies at the macro level. This ensures the coherence and consistency of strategies at each stage, ultimately achieving an optimal balance between individual round benefits and overall match score. This significantly improves the adaptability and accuracy of the present curling strategy prediction method in complex game scenarios. In addition, the present invention is based on a curling target detection and segmentation scheme based on a large visual model. Benefiting from the powerful generalization ability of the large visual model, this scheme can achieve relatively accurate curling target segmentation without the need for a large amount of labeled data of curling scenes. The present invention adopts a curling coordinate positioning method based on differentiable rendering optimization, and utilizes the consistency constraints of the standard 3D model of the curling target and the actual picture to accurately solve the true coordinates of the curling. Compared with other schemes that use rough target detection and positioning methods, the method of this scheme can not only obtain more accurate curling position information, but also does not require the camera to be installed directly above the venue, and has a wider range of application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 A structural diagram of a system corresponding to the present invention; Figure 2 This is the global Q network architecture diagram. DETAILED DESCRIPTION

[0025] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0026] The present invention proposes a curling motion collection, perception, and strategy prediction method for smart game viewing, which includes the following steps: S1. Fusing the same time frame images of a curling throw video obtained by cameras at multiple locations to obtain the current frame curling observation coordinates of each curling target, performing trajectory matching on each curling target based on the video frame sequence of the curling video, and obtaining the motion trajectories of all curling stones in the current curling throw as the curling motion acquisition perception result; S2. Obtain the current actual curling state, combine the expert strategy method and the improved curling strategy prediction model, obtain the maximum possible action corresponding to the current actual curling state as the strategy prediction result, and render the strategy prediction result; The training process of the improved curling strategy prediction model is as follows: Construct an initial policy network and an initial value network, obtain training data, and perform supervised training on the policy network and the value network based on the training data to obtain the first policy network and the first value network; Based on the improved self-game reinforcement learning and Monte Carlo tree search method, the first strategy network and the first value network were trained again to obtain an improved curling strategy prediction model.

[0027] The basic content and core invention points of the present invention are as follows: (1) In the acquisition and perception part, the present invention uses multiple cameras set up around and above the curling competition venue, aiming at the same curling track to shoot and simultaneously collect images of the curling venue. The images from each perspective are processed by the Segment Anything Model (SAM) visual model to segment and locate the curling target in each frame of the image. Then, for each segmented curling stone, a coordinate rotation angle optimization algorithm based on differentiable rendering is used to obtain the absolute coordinates and angle of the curling stone on the field. Finally, based on the continuity of the curling motion, the position of each curling stone is tracked by the Kalman filter algorithm, and the motion trajectory of each curling stone can be obtained. At the same time, according to the curling throwing rules, it can also be confirmed which team each curling stone comes from and which throw it is.

[0028] (2) Based on the KR-DL-UCT (Kernel regression-Deep Learning-Upper ConfidenceBound Apply to Tree Monte Carlo tree search) algorithm, the present invention introduces a global Q network and proposes a mechanism for the strategy model to integrate global wins and losses with single-game scores, optimizes the sampling process of the Monte Carlo tree search algorithm, and guides the strategy model to optimize the global nature of the strategy in long-term games.

[0029] (3) Data-driven strategy models may fail in extreme situations (largely behind), while expert strategies are stable in certain situations. This paper proposes an expert-learning strategy hybrid control system that effectively integrates expert strategies with strategies based on deep reinforcement learning to improve strategy robustness and generalization.

[0030] The present invention consists of three parts, including acquisition and perception module, simulation and strategy prediction module, and VR animation rendering module. Figure 1 shown.

[0031] 1. Collection and perception module: a. Multi-camera system: The present invention sets up more than 5 cameras around the curling venue. Each camera needs to be dispersed to a different viewing angle and spread out over an area of ​​more than 270 degrees around the venue. The specific position of the camera can be adjusted according to the competition venue so that the target venue area can be clearly seen. The video signals of multiple cameras are transmitted to the switch via optical fiber using the Network Device Interface (NDI) transmission equipment and protocol. The server connected to the switch via a network cable can synchronously read the video signals of each channel. Each camera needs to be calibrated in advance with the internal parameters of the resolution and focal length settings used, and a checkerboard calibration method can be used. After the camera position is determined, the external parameters of the camera are calibrated using the marking lines and marking points in the curling venue.

[0032] b. Curling Object Detection: For each viewpoint, a large-scale visual model is used to perform segmentation processing, generating segmentation maps for each curling stone on the track. The large-scale visual model used in this solution is the SAM model. This model has been extensively trained on large-scale segmentation datasets and can automatically detect and segment various objects without requiring additional training data, demonstrating strong generalization. The pre-trained SAM segmentation model is a general-purpose model that requires a target prompt to segment specific objects. To apply it to the specific curling scenario, this solution fine-tuned the model. First, images containing curling stones were collected from the internet or at tournaments. These images were then annotated using the pre-trained SAM model. This involves feeding the images into the SAM model and providing manual prompts indicating the location of the curling stones, either as points or boxes. The SAM model then outputs segmentation maps for all curling stones. Approximately 1,000 annotated curling images were used to fine-tune the pre-trained SAM model, enabling it to segment curling stones without any prompts.

[0033] c. Curling Stone Coordinate Prediction: Since curling stones have a relatively standard shape and size and slide only on the track surface, these constraints and the curling stone segmentation map obtained in step (b) allow the actual coordinates of the curling stone on the track to be calculated through optimization. First, a 3D mesh model of a standard curling stone is prepared, which is consistent with the curling stone in the actual game (the 3D model color is normalized to binary grayscale). For each curling stone segmentation map, the average value of the segmented pixels is calculated. Using the intrinsic and extrinsic parameters of the corresponding camera, these coordinates are back-projected back to the ground coordinates of the playing field to serve as the optimized initial coordinates for the curling stone. In virtual 3D space, the 3D curling stone model is initialized to the corresponding position. Then, using the intrinsic and extrinsic parameters of the corresponding camera, an approximate differentiable renderer (OpenDR) is used to project the resulting image. When the curling stone coordinates and rotation angle are accurate, this 3D projected segmentation map should overlap as closely as possible with the segmentation map obtained in step (b). Therefore, the plane coordinates and rotation angles of the curling stone can be optimized using an optimization method, as shown in the following formula:

[0034] Among them, OpenDR is the renderer, Obj is a fixed curling 3D model, x, y, and θ are the horizontal and vertical coordinates and rotation angles of the curling in the world coordinate system (track plane coordinate system), respectively, and M is the predicted curling segmentation map in b. The absolute value error between the rendered image and the predicted segmentation map is used as the loss function of both. Since the entire calculation process is differentiable, the gradient of the loss function will be able to backpropagate to the three variables x, y, and θ, so that the coordinate and rotation variables can be optimized using gradient optimization methods such as Newton's method. In this way, the actual coordinates and rotation of the curling can be obtained, and the inverse of the optimized nearest distance is used as the confidence level of the coordinates.

[0035] d. Multi-camera data fusion: Because multiple cameras simultaneously capture the playing field, excluding obstructed viewpoints, a single curling stone may be captured by multiple cameras. Therefore, the coordinates of the curling stone obtained from each viewpoint must be fused. First, the Hungarian matching algorithm is used to match the same curling stone from each viewpoint. The coordinates from each viewpoint are then weighted and averaged based on confidence to obtain the final coordinates.

[0036] e. Curling Tracking: Using a Kalman filter tracking algorithm, all curling stones are tracked sequentially in the video frame sequence to obtain each stone's trajectory. Specifically, a state vector is constructed using the stone's x, y coordinates, rotation angle θ, and velocity. When a new stone appears, the state vector is initialized, generating a new trajectory. At the next moment (a new frame), the current state is predicted based on the previous state and the state transition model. The Euclidean distance between all observed stone coordinates and predicted trajectory coordinates in the next frame, obtained through d, is calculated. The Hungarian algorithm is used to optimally match each observation to the trajectory. Successfully matched observations are updated to the corresponding trajectory sequence; unmatched observations are considered new. This method ensures that each stone has a continuous trajectory. Based on the curling rules and the order of the teams, the color and sequence number of each newly appeared stone's trajectory can be determined.

[0037] During the t+1 frame matching process, the predicted state vector for frame t+1 is derived using the observed state vectors of the curling stone's historical trajectory (frames t, t-1, ...) using the Kalman filter state prediction formula. The state vector for frame t+1 is obtained when the trajectory of the corresponding curling stone is updated after the t+1 frame matching is completed, using data such as the velocity calculated from the current trajectory (frames t+1, t...). The sequence of curling stone motion trajectories is the sequence of observed state vectors.

[0038] For the same video frame t, assume that the current frame curling observation coordinates of a curling target are identified and the state vectors of a curling target in video frame t are respectively assumed. The state vector of video frame t+1 is predicted by using the state vector of video frame t, the current frame curling observation coordinates of video frame t, and the Kalman filter state transition model to obtain the predicted coordinates of the a curling target in video frame t+1. The Euclidean distance between the current frame curling observation coordinates of the curling target in video frame t+1 and the predicted coordinates of the curling target in video frame t+1 is calculated. The Hungarian algorithm is used to perform optimal matching based on the Euclidean distance. The current frame curling observation coordinates of the curling target in video frame t+1 that is successfully matched are updated to the motion trajectory sequence of the corresponding curling target, and the state vector of the curling target in video frame t+1 is calculated based on the trajectory sequence. The current frame curling observation coordinates of the curling target in video frame t+1 that fails to match are used as the state vector of the newly appeared curling target. The state vector of the newly appeared curling target in video frame t+1 is initialized. The video frame is updated to t+2. The above steps are repeated to obtain the motion trajectories of all curling targets. The observation coordinates of the curling stone in the current frame are obtained from the video frame using a target recognition algorithm.

[0039] In the above steps, after identifying the observed coordinates of the a curling stone in the current frame and the state vectors of each curling stone in video frame t, for the first curling stone, its state vector in video frame t is used as the optimal estimate at the previous moment. Based on this optimal estimate, a priori estimate for frame t+1 (i.e., the predicted value of the state vector in frame t+1) is calculated. The priori error covariance and Kalman gain are also calculated. The observed coordinates of the first curling stone in the current frame are used as the measurement vector to obtain the posterior state estimate and posterior error covariance for frame t+1. These posterior state estimate and posterior error covariance are used as input for the next iteration. The predicted value of the state vector in frame t+1 is combined with the observed coordinates in video frame t to obtain the predicted coordinates of the curling stone in video frame t+1. The above steps are repeated for all curling stones to obtain the predicted coordinates of each of the a curling stone in frame t+1. The current frame curling observation coordinates of each curling target in the known video frame t+1 are matched with the predicted coordinates of a curling targets respectively. For a curling target, the current frame curling observation coordinates that successfully match its predicted coordinates are used as the next coordinate in the motion trajectory sequence of the curling target.

[0040] 2. Simulation and strategy prediction: a. Curling physics simulation platform: Based on a digital curling physics simulation platform, the present invention utilizes the aforementioned acquisition and perception module to collect multiple sets of sweep-free curling motion trajectories on-site. Based on these collected trajectories, parameters such as the friction coefficient in the physics simulation platform are configured through system identification methods, making the physics simulation environment more realistic than the actual playing field.

[0041] b. Overview of the policy module framework: The strategy prediction module in this invention is based on the KR-DL-UCT algorithm. Its network architecture comprises a strategy-value network. First, information such as the sequential positions of all stones in the current round and the home base position is converted into a 32x32 pixel image and input into the network. Multiple convolutional layers are used to generate latent features. These latent features are then fed into the strategy tap and the value tap, respectively. The strategy tap returns the probability distribution (action) of whether the stone will stop at the aforementioned 32x32 pixel location and whether it will reach that location by rotating clockwise or counterclockwise, i.e., the probability distribution of 32x32x2 candidate actions. The value network outputs the probability distribution of the score for the current situation. The score ranges from [-8 to 8] (value), with a total of 17 values ​​(each player throws 8 stones per round).

[0042] In order to solve the limitations of the above-mentioned discrete state and action space in high-dimensional continuous action scenarios, the Monte Carlo Tree Search (MCTS) method is further adopted. First, the state S of the adaptive strategy network at that moment is constructed based on the current field information obtained by perception. t,Input into the strategy-value network to obtain the probability distribution of the action , score probability distribution v t According to the action probability distribution output by the policy network , sampling a set of discrete actions As a set of base actions in the current state And the probability P corresponding to these actions t . KR-DL-UCT algorithm is based on S t 、 、P t 、v t The root node of the Monte Carlo tree is constructed for the input. Combining the physical simulation platform and the policy-value network in 5.2a, the four-step process (selection-expansion-simulation-backtracking) is called multiple times in the continuous state and action space. Finally, the normalized number of visits to the child nodes of the root node is used as a probability distribution, and the sampling is obtained based on the current state S t Action strategy .

[0043] Finally, the above action strategies are transformed into Map to the corresponding curling throwing speed and rotation direction. t The predicted curling throwing speed, rotation direction and other information are input into the simulation platform, and the physical simulation platform simulates the motion trajectory of the thrown curling stone and its collision with the curling stone on the field.

[0044] The training of the above strategy-value network is divided into two stages: supervised learning and self-game reinforcement learning. In the supervised learning stage, the real collected data {(S t 、a t 、z t )} for the label (z t is the final score of the current game), the strategy-value network parameters are optimized through the cross entropy loss function to make it correspond to the output {(S t 、 、v t )} is consistent with the label.

[0045] In the self-game reinforcement learning stage, the action strategy returned by Monte Carlo tree search during the self-game is collected. and the game result z t For labels, the same loss function optimization strategy as in the supervised learning stage is adopted - the value network parameters.

[0046] c. Improved self-game reinforcement learning algorithm Based on the massive results of the first stage of self-game training (the original self-game training process of KR-DL-UCT), we constructed {(St 、a t 、n t d t , r)} dataset, where n t Indicates the number of remaining games up to the current game, d t represents the cumulative score of all games (a win for the second player in the first game is a positive score, a loss for a negative score), and r = {1, 0, -1} represents whether the second player in the first game wins the entire game (a win is 1, a draw is 0, and a loss is -1).

[0047] A global Q network is designed, which is based on (S t 、a t 、n t d t ) is input and outputs the probability distribution of r in three categories {1, 0, -1}. The network architecture of the global Q network is as follows Figure 2 As shown in Figure 2, the global Q network is trained using the above-constructed dataset, with cross entropy as the loss function.

[0048] In the second stage of self-game training, the final sampling process of Monte Carlo tree search is modified as follows: the normalized number of visits to the child nodes of the root node is used as the probability of the corresponding action of the child node, which is called the local strategy probability distribution .according to Sample a set of scalars based on the current state S without replacement t Action And the corresponding probability , i = 1, 2, ..., k . Further, we can calculate n t d t , will (S t 、 、n t d t ) is input into the global Q network to obtain the corresponding global score probability distribution, and from it the probability of the party winning the global game is obtained (the probability of the second hand in the first game is r = 1, and the probability of the second hand is r = -1), which is recorded as . The indicator of fusing local and global information is calculated according to the formula.

[0049]

[0050]

[0051] Pick The largest corresponding action is returned. is a hyperparameter, c is generally taken as half of the total number of games. , the generated strategy pays more attention to the score of a single game (β is small) in the early stage of the game (when there are many remaining innings), and focuses more on the cumulative score of the whole game (β is large) in the later stage (when there are few remaining innings). In the second stage of self-play training, based on the above improved method, the algorithm performs self-play again, and accordingly updates the dataset {(S t , a t , n t , d t , r)}, and trains and fine-tunes the global Q-network based on the updated dataset, still using cross-entropy as the loss function. The a t in the updated dataset is the corresponding action with the largest value.

[0052] d. Expert-Learning Strategy Hybrid Control System Fuse the rule-based expert strategy (such as defensive / offensive tactics) with the above data-driven strategy to form a hybrid action probability distribution, so as to integrate the professional knowledge of coaches, athletes, commentators, etc. during the game, improve the diversity and interpretability of the strategy, and better support intelligent viewing of the game.

[0053] Specifically, we pre-construct an expert strategy library and define e expert strategies {π exp1 ,…,π expe}, each strategy corresponds to a specific tactic (such as "protect the center pot", "clear the opponent's pot"), and each expert strategy is implemented using a rule-based algorithm, which returns the action of the corresponding strategy according to the current state . (The above strategies do not require training.) In the inference stage, the system accepts f (f < e) expert strategies selected from the predefined e expert strategies and the corresponding credibility input manually (the value range is [0,1], and the higher the credibility, the more the expert tends to this strategy), and normalizes the credibility of these strategies so that their sum is. Further, we expand the input of Monte Carlo tree search from the discrete action set sampled from the output distribution of the policy-value network to , and at the same time P t is expanded to P t and the normalized credibility of the expert strategy actions, and construct the root node of the Monte Carlo tree from the above expanded input. The number of expert strategies f and the number of data-driven algorithm strategies m can be set and adjusted to meet the system's different tendencies towards the two. The subsequent steps remain unchanged. <In the above process, the self-game training process of the model, that is, retraining the first policy network and the first value network based on the improved self-game reinforcement learning and Monte Carlo tree search method to obtain the improved curling strategy prediction model, is as follows: A1. The initial curling stone state is used as the current curling stone state. A2. Input the current curling state into the first policy network and the first value network respectively to obtain the probability distribution of the action corresponding to the current curling state. And the probability distribution of scoring in one game v t , the probability distribution of actions Sample a set of discrete actions As the base action set for the current curling state ={ }, Represents the 1st, 2nd…mth discrete action of the base action set, m represents the total number of discrete actions, and the discrete action is obtained The corresponding probabilities P t ; A3. Set the current curling state and basic actions , probability P t And the probability distribution of scoring in one game v t As the root node of the Monte Carlo tree, after selecting, expanding, simulating and backtracking in sequence, the cumulative reward and number of visits of the node of the Monte Carlo tree are updated; A4. Use the normalized number of visits to the child nodes of the root node as the first action probability distribution; Sampling the first action probability distribution to obtain a first action corresponding to the current curling state, executing the first action based on the current curling state, using the obtained state as the new current curling state, and returning to A2 until the current round ends, where the end of the current round indicates that the current round has reached the throw count threshold. A5. Integrate the current state of each curling stone, the first action corresponding to each current state of each curling stone, the cumulative score from the first round to the current round, the number of rounds remaining in the current game up to the current round, and the win / loss status of the second player in the first round in the current game as the global Q network training set; A6. Use the first action corresponding to each current curling state as a label for the first policy network, and the current round result at the end of the current round as a label for the first value network for training, to obtain a second policy network and a second value network. A7. Build a global Q network and train the global Q network based on the global Q network training set to obtain a trained Q network. A8. Using the initial curling state as the current curling state; A9. Input the current curling state into the second policy network and the second value network respectively to obtain the probability distribution of the action corresponding to the current curling state. , and the probability distribution of scoring in one round v t ', the probability distribution of the action Sample a set of discrete actions , ,……, As the base action set for the current curling state , ,……, }, , ,……, Indicates the 1st, 2nd, ... A discrete action, Represents the total number of discrete actions and obtains discrete actions , ,……, The corresponding probabilities P t '; A10. Set the current curling state and basic actions , probability P t ' and the probability distribution of one-game score v t ' As the root node of the Monte Carlo tree, after selecting, expanding, simulating and backtracking in sequence, the cumulative reward and number of visits of the node of the Monte Carlo tree are updated; A11. Use the normalized number of visits to the child nodes of the root node as the second probability distribution , sample a set of sampled actions corresponding to the current curling state from the second probability distribution without replacement Indicates the 1st, 2nd, ... k sampling actions, and the probability corresponding to each sampling action ,in i =1, 2…… k , k Indicates the total number of sampled actions; A12. Input the current curling state, k sampled actions, the cumulative scores from the first round to the current round corresponding to the current curling state, and the number of rounds remaining in the current game corresponding to the current curling state into the trained Q network, and obtain the probability of the second player in the first round winning or losing in the current game corresponding to each sampled action. , based on the probability of winning or losing and the probability corresponding to each sampled action Calculate the probability of fusing local and global information , based on the current state of the curling stone, the probability of fusing local and global information is performed The maximum sampling action, and the obtained state is used as the new current state of throwing the curling stone, and return to A9 until the end of the current game; A13. Integrate the states of each current curling stone throw, the probabilities of fusing local and global information corresponding to the states of each current curling stone throw respectively The maximum action, the cumulative scores from the first game to the current game corresponding to the states of each current curling stone throw, the remaining number of games in the current match corresponding to the states of each current curling stone throw, and the win-loss situation of the second hand in the first game in the current match are integrated as the global Q-network fine-tuning dataset, and the trained global Q-network is fine-tuned based on the fine-tuning dataset to obtain the fine-tuned global Q-network.

[0055] When actually making predictions, the specific steps to obtain the most likely action corresponding to the current actual curling stone state by combining the expert strategy method and the improved curling stone strategy prediction model are as follows: B1. Input the current actual curling stone state into the second policy network and the second value network respectively, and the second policy network and the second value network output the probability distribution of the action corresponding to the current actual curling stone state and the probability distribution of the score of one game corresponding to the current actual curling stone state respectively; B2. Sample a set of actual actions and the probabilities corresponding to the actual actions from the probability distribution of the actions corresponding to the current actual curling stone state, and select f (f < e) expert strategies from the predefined e expert strategies , representing the 1st, ……, the f th expert strategy, and 1 expert strategy represents 1 action, and at the same time obtain the normalized credibility of each expert strategy; B3. The expert strategies and a set of actual actions form an actual discrete action set, and the probabilities corresponding to the actual actions and the normalized credibility of each expert strategy form the probabilities corresponding to the actual discrete action set; B4. Use the current actual curling stone state, the actual discrete action set, the probabilities corresponding to the actual discrete action set, and the probability distribution of the score of one game corresponding to the current actual curling stone state as the root node of the Monte Carlo tree, and then perform selection, expansion, simulation, and backtracking in sequence to update the cumulative rewards and visit counts of the nodes of the Monte Carlo tree; B5. Use the normalized visit counts of the child nodes of the root node as the first probability distribution ’, and sample a set of sampling actions corresponding to the current actual curling stone state without replacement from the first probability distribution representing the 1st, 2nd …… k th sampling actual action, and the probability corresponding to each sampling actual action , where i = 1, 2 ……k , k Indicates the total number of actual actions sampled; B6: Input the current actual curling state, k sampled actual actions, the cumulative score from the first round to the current round corresponding to the current actual curling state, and the number of rounds remaining in the current game corresponding to the current actual curling state into the fine-tuned global Q network. The fine-tuned global Q network outputs the probability of the second player in the first round corresponding to each sampled actual action winning or losing in the current game. Calculate the actual probability of the fusion of local and global information, and take the sampled actual action with the largest actual probability as the most likely action corresponding to the current actual curling state. The formula for the actual probability of fusion of local and global information and the probability of fusion of local and global information in A12 are the same. The formula is the same.

[0056] 3. Animation Rendering The VR event screen is reconstructed in 3D based on the information obtained by the acquisition and perception module, and the predicted strategy effect is rendered according to the information obtained by the physical simulation platform based on the strategy simulation predicted by the strategy module.

[0057] The present invention adopts a curling target detection and segmentation scheme based on a large visual model. Benefiting from the powerful generalization ability of the large visual model, this scheme can achieve relatively accurate curling target segmentation without the need for a large amount of labeled data of curling scenes.

[0058] The present invention adopts a curling coordinate positioning method based on differentiable rendering optimization, and uses the standard 3D model of the curling target and the consistency constraint of the actual picture to accurately solve the true coordinates of the curling. Compared with other schemes that use rough target detection and positioning methods, the method of this scheme can not only obtain more accurate curling position information, but also does not require the camera to be installed directly above the venue, and has a wider range of application scenarios.

[0059] The improved self-game reinforcement learning algorithm designed in the present invention improves the long-term strategy modeling of the algorithm by introducing a global Q network, and effectively balances the single-game win or loss of the generated strategy with the cumulative score of the entire game based on dynamic coefficients at different stages of the game.

[0060] The expert-learning strategy hybrid control system method adopted in this invention integrates expert domain knowledge through an expert strategy library, and combines data-driven methods with Monte Carlo tree search to further improve tactical diversity and strategy quality, avoiding deep reinforcement learning strategies from falling into local optimality. It can also effectively support the actual needs of commentators in smart viewing to select and predict tactical types.

[0061] The above describes in detail the preferred embodiments of the present invention. It should be understood that those skilled in the art can make numerous modifications and variations based on the concepts of the present invention without inventive effort. Therefore, any technical solutions that can be derived by those skilled in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.

Claims

1. A curling sports data collection, perception and strategy prediction method for smart game viewing, characterized in that: The method comprises the following steps: S1. Fusing the same time frame images of a curling throw video obtained by cameras at multiple locations to obtain the current frame curling observation coordinates of each curling target, performing trajectory matching on each curling target based on the video frame sequence of the curling video, and obtaining the motion trajectories of all curling stones in the current curling throw as the curling motion acquisition perception result; S2. Obtain the current actual curling state, combine the expert strategy method and the improved curling strategy prediction model, obtain the maximum possible action corresponding to the current actual curling state as the strategy prediction result, and render the strategy prediction result; The training process of the improved curling strategy prediction model is as follows: Construct an initial policy network and an initial value network, obtain training data, and perform supervised training on the policy network and the value network based on the training data to obtain the first policy network and the first value network; Based on the improved self-game reinforcement learning and Monte Carlo tree search method, the first strategy network and the first value network were trained again to obtain an improved curling strategy prediction model.

2. The curling sports data collection, perception and strategy prediction method for smart game viewing according to claim 1, characterized in that: The training data includes training label data {(TS t , Ta t , Tz t )}, the training label data includes multiple game label data, each game label data includes multiple round label data, each round label data includes all the states of a round of curling game and the final score Tz of the round of curling game t , all states of a curling game include the training state of each curling throw, and the training state TS of the t-th curling throw in all states of a curling game t Transition to the training state TS for the t+1th curling throw t+1 The corresponding training action is Ta t , supervised training is performed on the policy network and the value network based on the training label data to obtain the first policy network and the first value network.

3. The curling sports data collection, perception and strategy prediction method for smart game viewing according to claim 1, characterized in that: The specific steps for fusing the same time frame images of a curling throw video obtained by cameras at multiple locations to obtain the current frame curling observation coordinates of each curling target are as follows: Multiple curling targets are identified based on the same time frame images in a curling throwing video obtained by cameras at multiple positions, and the same time frame segmentation maps of each curling target from multiple perspectives are obtained. The current frame actual coordinates of the curling on the track from each perspective are determined based on the same time frame segmentation maps of each curling target from each perspective, and the current frame actual coordinates of the same curling target from various perspectives are fused to obtain the current frame curling observation coordinates of each curling target.

4. The curling sports data collection, perception and strategy prediction method for smart game viewing according to claim 1, characterized in that: The specific steps for matching the trajectory of each curling target based on the frame sequence of the curling video and obtaining the motion trajectory of all curling stones are as follows: For the same video frame t, assume that a curling targets are identified, and the current frame curling observation coordinates of the a curling targets and the state vectors of the a curling targets in video frame t are obtained. The state vector of video frame t+1 is predicted by using the state vector of video frame t, the current frame curling observation coordinates of video frame t, and the Kalman filter state transition model to obtain the predicted coordinates of the a curling targets in video frame t+1. The Euclidean distance between the current frame curling observation coordinates of the curling target in video frame t+1 and the predicted coordinates of the curling target in video frame t+1 is calculated. The Hungarian algorithm is used to perform optimal matching based on the Euclidean distance. The current frame curling observation coordinates of the curling target in video frame t+1 that is successfully matched are updated to the motion trajectory sequence of the corresponding curling target, and the state vector of the curling target in video frame t+1 is obtained. The current frame curling observation coordinates of the curling target in video frame t+1 that fails to match are used as the newly appeared curling target. The state vector of the newly appeared curling target in video frame t+1 is initialized, and the video frame is updated to t=t+1. The above steps are repeated to obtain the motion trajectories of all curling stones in the current curling throw.

5. The curling sports data collection, perception and strategy prediction method for smart viewing according to claim 1, characterized in that: The specific steps of retraining the first policy network and the first value network based on the improved self-game reinforcement learning and Monte Carlo tree search method to obtain the improved curling strategy prediction model are as follows: A1. The initial curling stone state is used as the current curling stone state. A2. Input the current curling state into the first policy network and the first value network respectively to obtain the probability distribution of the action corresponding to the current curling state. And the probability distribution of scoring in one game v t , the probability distribution of actions Sample a set of discrete actions As the base action set for the current curling state ={ }, Represents the 1st, 2nd…mth discrete action of the base action set, m represents the total number of discrete actions, and the discrete action is obtained The corresponding probabilities P t ; A3. Set the current curling state and basic actions , probability P t And the probability distribution of scoring in one game v t As the root node of the Monte Carlo tree, after selecting, expanding, simulating and backtracking in sequence, the cumulative reward and number of visits of the node of the Monte Carlo tree are updated; A4. Use the normalized number of visits to the child nodes of the root node as the first action probability distribution; Sampling the first action probability distribution to obtain a first action corresponding to the current curling state, executing the first action based on the current curling state, using the obtained state as the new current curling state, and returning to A2 until the current round ends, where the end of the current round indicates that the current round has reached the throw count threshold. A5. Integrate the current state of each curling stone, the first action corresponding to each current state of each curling stone, the cumulative score from the first round to the current round, the number of rounds remaining in the current game up to the current round, and the win / loss status of the second player in the first round in the current game as the global Q network training set; A6. Use the first action corresponding to each current curling state as a label for the first policy network, and the current round result at the end of the current round as a label for the first value network for training, to obtain a second policy network and a second value network. A7. Build a global Q network and train the global Q network based on the global Q network training set to obtain a trained Q network. A8. Using the initial curling state as the current curling state; A9. Input the current curling state into the second policy network and the second value network respectively to obtain the probability distribution of the action corresponding to the current curling state. , and the probability distribution of scoring in one round v t ', the probability distribution of the action Sample a set of discrete actions , ,……, As the base action set for the current curling state , ,……, }, , ,……, Indicates the 1st, 2nd, ... A discrete action, Represents the total number of discrete actions and obtains discrete actions , ,……, The corresponding probabilities P t '; A10. Set the current curling state and basic actions , probability P t ' and the probability distribution of one-game score v t ' As the root node of the Monte Carlo tree, after selecting, expanding, simulating and backtracking in sequence, the cumulative reward and number of visits of the node of the Monte Carlo tree are updated; A11. Use the normalized number of visits to the child nodes of the root node as the second probability distribution , sample a set of sampled actions corresponding to the current curling state from the second probability distribution without replacement Indicates the 1st, 2nd, ... k sampling actions, and the probability corresponding to each sampling action ,in i =1, 2…… k , k Indicates the total number of sampled actions; A12. Input the current curling state, k sampled actions, the cumulative scores from the first round to the current round corresponding to the current curling state, and the number of rounds remaining in the current game corresponding to the current curling state into the trained Q network, and obtain the probability of the second player in the first round winning or losing in the current game corresponding to each sampled action. , based on the probability of winning or losing and the probability corresponding to each sampled action Calculate the probability of fusing local and global information , based on the current state of the curling stone, the probability of fusing local and global information is performed The maximum sampling action is performed, and the resulting state is used as the new current curling state. Return to A9 until the current game ends. A13, the probability of fusion of local and global information corresponding to the state of each current curling stone The maximum action, the cumulative score from the first round to the current round corresponding to the current curling state, the number of remaining rounds in the current game corresponding to the current curling state, and the win-loss situation of the second player in the first round in the current game are integrated as the global Q network fine-tuning dataset. The trained global Q network is fine-tuned based on the fine-tuning dataset to obtain a fine-tuned global Q network.

6. The curling sports data collection, perception and strategy prediction method for smart viewing according to claim 5, characterized in that: The specific steps for obtaining the current actual curling state and combining the expert strategy method with the improved curling strategy prediction model to obtain the maximum possible action corresponding to the current actual curling state are as follows: B1. The current actual curling state is input into the second policy network and the second value network respectively. The second policy network and the second value network respectively output the probability distribution of the action corresponding to the current actual curling state and the probability distribution of the score of the round corresponding to the current actual curling state; B2. Sample a set of actual actions and the probabilities corresponding to the actual actions from the probability distribution of the actions corresponding to the current actual curling state, and select f (f < e) expert policies from the predefined e expert policies , denote the 1st, ……, the f th expert policies. One expert policy represents one action, and at the same time, obtain the normalized credibility of each expert policy; B3. The expert strategy and a set of actual actions constitute the actual discrete action set. The probability corresponding to the actual action and the normalized credibility of each expert strategy constitute the probability corresponding to the actual discrete action set. B4. Use the current actual curling state, the actual discrete action set, the probability corresponding to the actual discrete action set, and the probability distribution of the score in one round corresponding to the current actual curling state as the root node of the Monte Carlo tree. After selecting, expanding, simulating, and backtracking in sequence, update the cumulative reward and visit count of the node in the Monte Carlo tree; B5. Take the normalized number of visits to the child nodes of the root node as the first probability distribution ', sample a set of sample actions corresponding to the current actual curling state from the first probability distribution without replacement Indicates the 1st, 2nd, ... k The actual actions sampled, and the probability corresponding to each actual action sampled ,in i =1, 2…… k , k Indicates the total number of actual actions sampled; B6. Input the current actual curling state, k sampled actual actions, the cumulative scores from the first game to the current game corresponding to the current actual curling state, and the number of games remaining in the current game corresponding to the current actual curling state up to the current game into the fine-tuned global Q network. The fine-tuned global Q network outputs the probability of the second player in the first game corresponding to each sampled actual action winning or losing in the current game, calculates the actual probability of fusing local and global information, and takes the sampled actual action with the largest actual probability as the most likely action corresponding to the current actual curling state.

7. The curling sports data collection, perception and strategy prediction method for smart game viewing according to claim 5, characterized in that: The probability of fusing local and global information is: ; in, represents the probability of fusing local and global information, Indicates the i Sampling action, is the dynamic coefficient.

8. The curling sports data collection, perception and strategy prediction method for smart game viewing according to claim 7, characterized in that: The dynamic coefficient is: ; in, is a hyperparameter, c is half of the total number of rounds in a curling game, and n t The number of remaining rounds in the current game corresponding to the current curling state.

9. The curling sports data collection, perception and strategy prediction method for smart game viewing according to claim 1, characterized in that: The state space corresponding to the state is the position of all curling stones in sequence and the position of the base camp, whether the current curling stone is thrown by the second team in the first round, which throw the current curling stone is in this round, and the current round number.

10. The curling sports data collection, perception and strategy prediction method for smart viewing according to claim 1, characterized in that: The action space corresponding to the action is the target position of the curling stone and whether to rotate clockwise or counterclockwise to reach the target position.

Citation Information

Patent Citations

  • Curling match strategy generation method based on Monte Carlo reinforcement learning

    CN113673672A

  • Football match technique and tactical evaluation method and system

    CN118485353A

  • Automatic driving decision-making method and system based on generative world large model and multi-step reinforcement learning

    CN118790287A

  • Floating leisure article and connector

    KR1020210146176A