General training method for ball games based on deep learning technology for moving target detection and auxiliary referee system
By employing deep learning technology and a gridded motion index model, the issues of universality and camera compatibility in ball motion detection were resolved, enabling precise localization and real-time feedback of moving targets, thereby improving detection accuracy and device adaptability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING ZHIYUAN GUANGRUN SURVEY TECH CO LTD
- Filing Date
- 2022-11-10
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies lack versatility in ball sports detection, cannot be automatically extended to multiple sports, have large detection errors, require a lot of manual intervention, have poor camera compatibility, and the entire device stops working once a node malfunctions.
Deep learning technology is used to establish basic weight parameters, which are then adjusted in specific scenarios through a parameter optimization module. Combined with a gridded motion index model, this enables precise positioning and 3D coordinate detection of moving targets. Real-time analysis and feedback are then performed using multi-camera integration technology.
It achieves accurate identification and localization of moving targets, provides real-time motion index feedback, supports multi-camera integration, improves the versatility and reliability of detection, reduces human intervention, and enhances the device's adaptability.
Smart Images

Figure CN115845349B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing, and in particular to a general training method and auxiliary referee device for ball sports based on deep learning technology for moving target detection. Background Technology
[0002] By deploying a fixed number of cameras around the sports field, artificial models of the specific ball or person to be detected are created. During sports competitions or training, the landing point of the ball is captured, which can be used to assist athletes and referees in the refereeing process.
[0003] Existing technologies all involve building dedicated systems that can only serve a single sport, and there is no solution that can be automatically expanded and is universal.
[0004] Using artificial models is a difference between deep learning and conventional image detection, resulting in poor universality.
[0005] There is currently no gridded motion index model; such a model must be used in conjunction with detection methods to be effective.
[0006] In terms of ball sports detection technology, existing technologies mainly use traditional manual modeling methods for ball sports image detection, without using modern advanced detection technologies such as deep learning. This results in large detection errors, poor adaptability to different venues, and a significant amount of manual intervention required.
[0007] In terms of applicability across multiple sports, current technologies are still limited to specific events and cannot be universally applied to various ball sports.
[0008] In terms of training techniques, there is no unified grid model. Specific grid modeling can only be carried out for specific ball sports, and a single grid model cannot be universally applicable to various ball sports training programs.
[0009] In terms of multi-camera integration, there is no unified adaptive access method. All camera access requires a dedicated interface for matching, resulting in poor compatibility among multiple cameras.
[0010] In terms of device configuration, there is no unified standard. When a problem occurs in a certain node, the entire node will stop working. Summary of the Invention
[0011] In view of the above problems, the present invention proposes a general training method and auxiliary referee device for ball sports based on deep learning technology for moving target detection to overcome or at least partially solve the above problems.
[0012] A general training method for ball sports based on deep learning technology for moving target detection includes:
[0013] Step 1: Establish basic weight parameters. Deep learning is performed on the model trained by deep learning to establish the weight parameters after training.
[0014] Step 2: Parameter optimization for specific motion scenarios. When errors occur in the basic weight parameters, a parameter optimization module is designed to collect images of the actual motion scenarios, conduct optimization training, and establish optimized weight parameters for specific scenarios.
[0015] Furthermore, the weighting parameters are used to analyze motion images in real time during actual operation, identify moving targets, detect the three-dimensional coordinates of moving targets, reconstruct motion trajectories, and detect motion indicators of motion speed by combining time information.
[0016] Furthermore, the feature is that the error is due to insufficient samples for deep learning training. When this occurs, samples are created from the on-site data, incremental training is performed, and the weights are re-established.
[0017] Furthermore, the feature is that the deep learning training model includes: a gridded motion index model, which superimposes motion state indices at a certain moment on the grid to form a unified model of grid position, time, and motion state.
[0018] Furthermore, the feature is that the gridded motion index model includes establishing a grid model and establishing a motion state index model.
[0019] Furthermore, the feature is that the establishment of the grid model specifically involves establishing a nine-square grid model, which can be adjusted according to the actual sports field to adapt to the sports field; the establishment of the motion state index model involves establishing a human motion state model, a ball motion state model, and a grid motion index model.
[0020] Furthermore, the feature is that the human motion state model includes: human motion trajectory, human motion speed, human motion distance, and human motion consciousness.
[0021] Furthermore, the feature is that the ball's motion state model includes the ball's trajectory, ball's velocity, ball's starting point, ball's landing point, ball's rotation speed, and direction.
[0022] Furthermore, the feature is that the grid motion index model includes a grid index statistics table.
[0023] This aspect also provides an auxiliary referee system for a general training method for ball sports that uses deep learning technology to detect moving targets. The system includes: a human-computer interaction system host for analyzing and statistically processing various indicators and displaying them via a display controller or television interface to an operator monitor, a large screen at the sports venue, a referee's seat display screen, and a television control center.
[0024] The technical solutions provided in the embodiments of the present invention have at least the following technical effects or advantages:
[0025] Utilizing advanced computer image recognition technology (deep learning), this system accurately identifies and locates moving targets, acquiring their basic motion parameters (target recognition, 3D positioning, velocity, trajectory, and center of gravity). Combined with sports competition and training rules, it establishes a gridded set of motion index parameters. Real-time feedback on changes in these grid motion indexes during training or competition assists coaches in optimizing or improving training effectiveness. The integrated technology, which allows for arbitrary overlay and expansion of multiple cameras, can be used for Voice Assistant Referee (VAR) during competitions and for replaying high-speed, high-definition images from any angle during training to help athletes analyze their movements and postures.
[0026] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0027] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0028] Figure 1(A) shows the minimum processing network unit structure;
[0029] Figure 1(B) shows the layout of the smallest processing network unit;
[0030] Figure 2(A) shows the standard networking unit group structure;
[0031] Figure 2(B) shows the layout of the standard networking unit group;
[0032] Figure 3(A) shows the structure of a standard networking unit group that can be expanded to a maximum of 4 groups;
[0033] Figure 3(B) shows the layout that can be expanded to a maximum of 4 standard networking units;
[0034] Figure 4 The maximum supported network topology is shown;
[0035] Figure 5 The camera layout of the sports field is shown;
[0036] Figure 6The process for establishing the basic weight parameters is shown;
[0037] Figure 7 The parameter optimization process for a specific motion scenario is shown;
[0038] Figure 8 A nine-square grid model is shown;
[0039] Figure 9 A diagram illustrating how to adapt a nine-square grid to a badminton court is shown.
[0040] Figure 10(A) shows the structure of a VAR (Video Assistant Referee) system for ball sports based on deep learning technology for moving target detection;
[0041] Figure 10(B) shows the layout of a VAR (Video Assistant Referee) system for ball sports based on deep learning technology for moving target detection. Detailed Implementation
[0042] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0043] The scalable high-speed high-definition camera universal integration technology is divided into the minimum processing network unit, the standard network unit group, the maximum expandable 4 standard network unit groups, and the maximum supported network based on the number of cameras.
[0044] The basic components of the video acquisition system, also known as the minimum processing network unit, are shown in Figure 1(A), which is a structural diagram, and Figure 1(B), which is a field layout diagram. The images acquired by camera 1 and camera 2 are transmitted to the image processing terminal via gigabit network, 10-gigabit switch, and 10-gigabit network, and then processed by the field interactive and display control system.
[0045] A standard networking unit group can be equipped with up to 8 cameras depending on the coverage area. Figure 2(A) is a structural diagram and Figure 2(B) is a field layout diagram. The images collected by cameras 1 to 8 are transmitted to the image processing terminal via a 10 Gigabit switch and processed by the field interactive and display control system.
[0046] It can be expanded to a maximum of 4 standard networking units. Figure 3(A) is a structural schematic diagram and Figure 3(B) is a field layout schematic diagram. The images collected by the 4 groups of cameras 1 to 8 are transmitted to the image processing terminal through the corresponding gigabit network and 10-gigabit switch of each group, and the data is processed by the corresponding field interaction and display control system of each group.
[0047] Maximum supported network topology model such as Figure 4 As shown, data from networking units 1 to 4 are transmitted to the image processing terminal server via a 10 Gigabit switch, and then processed by the on-site interactive and display control system.
[0048] Camera layout in the sports field as follows Figure 5 As shown.
[0049] The methods for using deep learning technology to identify and accurately locate spherical and human targets are as follows:
[0050] The deep learning process involves first acquiring sample data, then selecting a suitable network structure, such as CNN or RNN, and using a deep learning framework for training. The final training result is the weight parameters. This invention designs and establishes a basic motion scene weight parameter management module. Through this module, training results for various motion scenes can be quickly managed, forming a basic weight parameter library. When the system runs in a specific motion scene, it can directly use the basic weights for testing. If the accuracy of the basic weight parameters is found to be insufficient, the weight parameter optimization module can be called to optimize the weights. After optimization, testing can be repeated until satisfactory weight parameters are obtained.
[0051] The following are examples of methods using deep learning technology:
[0052] A general training method for ball sports based on deep learning technology for moving target detection includes:
[0053] Step 1: Establish basic weight parameters. This involves using a deep learning model to perform deep learning and establish the trained weight parameters. These parameters are used in real-time analysis of motion images during actual operation to identify moving targets, detect their 3D coordinates, reconstruct trajectories, and, combined with time information, detect motion indicators such as speed. Figure 6 As shown.
[0054] Step 2: Parameter Optimization for Specific Motion Scenarios. In real-world scenarios, the base weights may exhibit significant errors. By designing a parameter optimization module, images of actual motion scenarios are collected and optimized through training to establish optimized weight parameters for specific scenarios. Errors arise because deep learning training may not have sufficient samples. When this occurs, on-site data can be used to create samples, and incremental training can be performed to re-establish the weights. Figure 7 As shown.
[0055] Taking volleyball as an example, cameras are installed around the volleyball court to collect video footage. This video footage is then analyzed using deep learning methods for image analysis, which yields better results than traditional modeling. Extracting the three-dimensional coordinates (X, Y, Z) of a player's movement and connecting these coordinates reveals the player's trajectory and speed. Similarly, obtaining the ball's three-dimensional coordinates provides data such as its flight trajectory, speed, and landing point. Analyzing this data according to volleyball rules yields volleyball performance indicators, such as average speed, maximum speed, running distance, serve speed, and ball landing point. Since each indicator has its own coordinates, dividing the court into a nine-square grid allows each indicator to be mapped to a specific grid. These grids can then form tactical zones such as attack zones, defense zones, and service zones. During player training, the system automatically identifies whether each ball is in the attack or defense zone, and the hitting effect is used to determine if the set speed, landing point, and timing requirements have been met. This training method allows for a more accurate evaluation of performance.
[0056] Deep learning training models include: gridded motion index models.
[0057] Specifically, the gridded motion index model includes:
[0058] (I) Establishing a grid model, specifically a nine-square grid model. The grid model serves as a reference system for ball sports training; for example, the target area for the training ball is represented by a grid. By setting the grid, a ball landing point model is established. Detecting and statistically analyzing this model allows us to determine the training effectiveness, such as the number of successful and unsuccessful landings within the grid area. The size and position of the grid area can be used as weights to calculate the success rate and failure rate. This model can be adjusted to suit the actual sports field. The nine-square grid model is as follows: Figure 8 As shown.
[0059] Create a 3x3 grid and number each cell sequentially from 1 to 9.
[0060] For example: adapting the above nine-square grid to a badminton court, then as follows: Figure 9 As shown.
[0061] (II) Establish a motion state index model, which includes:
[0062] (1) A model of human motion states, including:
[0063] Human motion trajectory: The trajectory of a person during movement. Deep learning is used to obtain the person's position coordinates from an image. The set of coordinates over a continuous time period constitutes the person's motion trajectory, represented by two-dimensional coordinates, P(x1,y1,x2,y2...xn,yn).
[0064] Human speed: The speed of a person's movement over a continuous time period can be calculated using V = S / T.
[0065] Human movement distance: The distance a person travels within a continuous time period; this can be obtained by summing the distances to all points.
[0066] Center of gravity: Through deep learning, the posture and movement of people in the image are detected, and the position of the person's center of gravity is identified based on the posture and movement. The center of gravity position is represented by 8 directions: front, back, left, right, left front, right front, left back, and right back. The center of gravity is obtained through deep learning training.
[0067] Human motion awareness includes four states: offense, defense, offensive and defensive, and unconscious. Motion awareness is determined by a person's positioning and the ball's landing point. When a person runs from the baseline towards the center line or the opponent's direction with a relatively high ball speed, it can be considered an offensive state; conversely, it's a defensive state. Offensive and defensive states fall somewhere in between. When a person continuously hits the ball from a certain position with the ball's speed and direction fluctuating within a small range, it's considered an unconscious state.
[0068] Reaction time: This is a parameter for human movement, calculated by subtracting the ball's return start time from the ball's arrival time.
[0069] (2) The motion state model of the ball includes:
[0070] Ball trajectory: The trajectory of a ball's flight is obtained from an image using deep learning to capture the ball's position coordinates. The set of coordinates over a continuous time period constitutes the ball's trajectory, represented by three-dimensional coordinates, B(x1,y1,z1,x2,y2,z2...xn,yn,zn).
[0071] Ball velocity: The velocity of the ball over a continuous time period can be calculated using V = S / T.
[0072] Ball origin point: The point on the ball when a person hits it or when the body touches the ball, represented by three-dimensional coordinates as B(x,y,z).
[0073] Landing point: The position of the ball when it lands, represented by three-dimensional coordinates as B(x,y,z).
[0074] The ball's rotational speed and direction: The direction and speed of the ball's rotation during flight.
[0075] Rotation, acceleration, trajectory, and landing point are determined by images captured by several cameras in this project. The images are then trained using deep learning based on neural networks to obtain a neural network model. The motion parameters of the ball can be calculated based on this neural model.
[0076] The above-mentioned gridded motion index model overlays motion state indices at a certain moment onto the grid, forming a unified model of grid position, time, and motion state.
[0077] Grid Indicator Statistics Table
[0078]
[0079] Figure 10 shows a VAR (Video Assistant Referee) system for ball sports that uses deep learning technology for moving target detection and employs human-computer interaction technology for assisted training and referee assistance.
[0080] Figure 10(A) is a schematic diagram of the system structure, and Figure 10(B) is a schematic diagram of the on-site layout. The host (control interaction server) of the human-computer interaction system can analyze and statistically analyze various indicators and perform display control, which are presented to the operator monitor (not shown in Figure 10(B)), the (sports) on-site large screen, the referee's seat display screen (not shown in Figure 10(B)), the television control center, etc., through the display controller (matrix) or the television interface.
[0081] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0082] Similarly, it should be understood that, in order to simplify this disclosure and aid in understanding one or more of the various aspects of the invention, in the above description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof. However, this method of disclosure should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, inventive aspects lie in fewer than all features of a single foregoing disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into this detailed description, wherein each claim itself is a separate embodiment of the invention.
[0083] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims.
Claims
1. A general training method for ball sports based on deep learning technology for moving target detection, comprising: Step 1: Establish basic weight parameters: Perform deep learning through the model trained by deep learning, establish the weight parameters after training, manage the weight parameters for various motion scenarios, and form a basic weight parameter library; The deep learning training model includes: a gridded motion index model; The gridded motion index model includes establishing a grid model and establishing a motion state index model; The grid model is a nine-square grid model. This model can be adjusted according to the actual sports field to adapt to the sports field. The motion state index model is a human motion state model and a ball motion state model. The motion state index at a certain moment is superimposed on the grid to form a unified model of grid position, time and motion state. Step 2, Parameter optimization for specific motion scenarios: When facing a specific motion scenario, select the basic weight parameters for that specific motion scenario from the basic weight parameter library for testing. When errors occur in the basic weight parameters, design a parameter optimization module, collect images of the actual motion scenario, perform optimization training, and establish optimized weight parameters for the specific scenario. The human motion state model includes: human motion trajectory, human motion speed, human motion distance, and human motion consciousness; The ball's motion model includes the ball's trajectory, ball's speed, ball's starting point, ball's landing point, ball's rotation speed, and direction.
2. The method according to claim 1, characterized in that, in, The weight parameters are used to analyze motion images in real time during actual operation, identify moving targets, detect the three-dimensional coordinates of moving targets, reconstruct motion trajectories, and detect motion indicators such as motion speed by combining time information.
3. The method according to claim 1, characterized in that, in, The error is because the number of samples used for deep learning training is insufficient. When this happens, samples are created from the on-site data, incremental training is performed, and the weights are re-established.
4. The method according to claim 1, characterized in that, in, The grid motion index model includes a grid index statistics table.
5. An auxiliary referee system employing the method of any one of claims 1 to 4, comprising: The host computer of the human-computer interaction system is used to analyze and statistically analyze various indicators and to display and control them. These indicators are presented to the operator monitor, the large screen at the sports venue, the referee's display screen, and the television control center via the display controller or television interface.