A method and apparatus for high-speed tennis trajectory recognition and 3D reconstruction based on monocular vision under limited computing power.
Patent Information
- Application Number
- CN202610621589.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-08
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2046-05-08
AI Technical Summary
[0008]硬件与算力成本过高:双目/多目系统需要多台昂贵的高速工业相机、严格的时钟同步硬件,以及高性能GPU工作站来处理庞大的多视角图像数据,无法下放到普通的边缘计算设备(如智能手机、轻量级计算盒子)中
[0065]①极低的硬件成本与算力门槛:相比双目/多目系统,省去了昂贵的多台相机和算力设备,只需单台普通摄像头。通过算法的轻量化设计,能在普通边缘算力设备上流畅运行。
Smart Images

Figure CN122134762B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of machine vision technology, and in particular to a method and apparatus for high-speed tennis trajectory recognition and 3D reconstruction based on monocular vision under limited computing power. Background Technology
[0002] With the development of smart sports and computer vision, 3D reconstruction of motion trajectories has been widely used in sports event refereeing assistance (such as the "Hawk-Eye" system), athlete training analysis, and mass sports entertainment.
[0003] Tennis balls are not only small in size, but also fly at extremely high speeds (serving speeds can exceed 200 km / h). When they are moving at high speeds, the images captured by cameras are often blurry pixelation with motion blur, and they are easily confused with the background of the court (such as white lines and stands). This is a typical problem of "high-speed, small target detection and tracking".
[0004] Existing technological solutions mainly include:
[0005] Multi-view / binocular vision systems (such as traditional Hawk-Eye technology): use multiple (usually 6-10) high-speed cameras to capture images simultaneously from different angles. Through extremely rigorous camera joint calibration, the 3D spatial coordinates of the tennis ball are calculated using the principles of stereo matching and triangulation.
[0006] Traditional monocular vision systems use ordinary background subtraction or basic 2D object detection algorithms (such as the YOLO version) to extract the 2D pixel coordinates of the tennis ball, and then rely solely on physical parabolic models and simple ground plane assumptions (such as assuming that the tennis ball only moves on a specific plane or must hit the ground to be measured) to roughly infer the 3D trajectory.
[0007] Disadvantages of existing technology:
[0008] Excessive hardware and computing costs: Binocular / multi-view systems require multiple expensive high-speed industrial cameras, strict clock synchronization hardware, and high-performance GPU workstations to process massive multi-view image data, making it impossible to scale up to ordinary edge computing devices (such as smartphones and lightweight computing boxes).
[0009] Deployment and calibration are extremely complicated: multi-camera systems require harsh on-site installation environments, and once a camera undergoes even a slight displacement, complex joint calibration must be performed again.
[0010] Monocular reconstruction suffers from low accuracy and poor robustness: Existing monocular technology cannot effectively solve the problem of missing depth information, and when faced with background interference and high-speed blurring, 2D detection is prone to missing detection, causing subsequent 3D fitting to completely fail. Summary of the Invention
[0011] To address the aforementioned issues, this disclosure provides a monocular vision tennis 3D reconstruction system and method operating under limited computing power. By significantly reducing hardware costs, computing power requirements, and deployment complexity, and through the combination of offline calibration and the lightweight BEVFormer algorithm framework, it solves the problems of high-speed, small target detection, tracking, and depth estimation, achieving high-precision tennis 3D trajectory reconstruction, velocity measurement, and landing point calculation comparable to multi-view systems.
[0012] The method for high-speed tennis trajectory recognition and 3D reconstruction based on monocular vision under limited computing power disclosed herein mainly includes the following steps:
[0013] S1, perform a one-time offline calibration before running;
[0014] S2, For devices with limited computing power, a lightweight BEVFormer network is constructed. The lightweight BEVFormer network takes a monocular image as input and the 3D spatial coordinates of the tennis ball in the current frame as output.
[0015] S3, through a single-channel image acquisition device, pushes the stream in real time, inputs the trained lightweight BEVFormer network, and outputs continuous 3D coordinate points;
[0016] S4. After filtering and smoothing the continuous 3D coordinate points output in step S3, calculate the flight trajectory of the tennis ball to obtain the instantaneous speed and landing point.
[0017] Furthermore, step S1 specifically includes:
[0018] The intrinsic parameter matrix K of the monocular camera, including the focal length f, was obtained in advance using Zhang Zhengyou's calibration method combined with a checkerboard pattern. x f y and principal point coordinates c x c y and distortion coefficient;
[0019] Using the standard geometric dimensions of a tennis court, the corresponding pixel corners in the image are extracted. The PnP algorithm is used to calculate the camera's extrinsic parameters relative to the tennis court, including the rotation matrix R and translation vector T, to establish a unified 3D world coordinate system.
[0020] Furthermore, step S2 specifically includes:
[0021] The BEVFormer standard architecture used for autonomous driving is pruned and modified to be lightweight and adapted to devices with limited computing power. The input of the lightweight BEVFormer network is a continuous stream of N frames of monocular video images, and the output is the 3D spatial coordinates of the tennis ball in the current frame.
[0022] The lightweight BEVFormer network employs a spatiotemporal attention mechanism, wherein:
[0023] Temporal attention: Integrating historical bird's-eye view features from the previous N-1 frames to obtain the movement trend of the tennis ball;
[0024] Spatial attention: Combining the camera intrinsic and extrinsic parameter matrices obtained in calibration step S1, 2D image features are mapped to 3D voxel space through a deformable attention mechanism, focusing on the spatial region where the tennis ball may exist.
[0025] Furthermore, in step S2, the method for lightweight pruning and modifying the BEVFormer standard architecture includes one or more of the following:
[0026] (1) Lightweight replacement of the backbone network, including: replacing the feature extraction layer Backbone of the standard BEVFormer with a lightweight convolutional neural network;
[0027] (2) For small targets, perform "spatial attention" reconstruction, including: in the spatial attention module, remove low-resolution deep features and retain high-resolution shallow features; at the same time, reduce the number of attention heads and the number of sampling points for each reference point;
[0028] (3) The BEV bird's-eye view space is reduced in size and customized, including: the perception range only covers the standard tennis court and its surroundings; the resolution of the XY plane of the BEV Grid is refined; and the voxel division of the Z axis adopts "non-equidistant refined design".
[0029] (4) For edge devices with limited computing power, special operators in the network were replaced to compress the single-frame inference delay time.
[0030] Furthermore, in step S4, the specific methods for calculating the instantaneous velocity and predicting the landing point include:
[0031] (1) Instantaneous velocity calculation: Instantaneous velocity is calculated using a time sliding window. Assume the video frame rate is FPS and the time interval is... The instantaneous spatial velocity at frame t The calculation formula is:
[0032]
[0033] Indicates the number of frames contained in the sliding window;
[0034] (2) Calculation of landing point:
[0035] The method of Z-axis extreme value detection combined with fitting is used to find the frame tt with the lowest z-axis, and 3D curve fitting is performed on the trajectory of the m frames before and n frames after tt respectively. The intersection point obtained by the two curve fittings is taken as the landing point.
[0036] Furthermore, the method also includes the following steps:
[0037] Acquire high-precision stereo data as ground truth data for network training;
[0038] The lightweight BEVFormer network is trained using monocular continuous images as input and high-precision binocular 3D data as supervision signals.
[0039] Furthermore, the method for acquiring high-precision binocular data specifically includes:
[0040] In real tennis matches or training, a binocular vision system is used to simultaneously collect binocular video data.
[0041] Using the high-precision triangulation of a binocular system, the high-precision 3D spatial coordinates P of each frame of tennis ball are calculated. t_gt =(X t_gt Y t_gt Z t_gt ), which serves as the ground truth label for network training, where t represents the t-th frame and gt represents the ground truth.
[0042] Furthermore, the specific method for training the lightweight BEVFormer network includes: using a joint loss function for end-to-end supervised training, the specific calculation formula of which is as follows:
[0043] 3D coordinate regression loss The error between the predicted coordinates and the stereo ground truth coordinates is calculated using Smooth L1 Loss.
[0044]
[0045] T represents the time span of the time sliding window used to calculate coordinate regression loss or instantaneous velocity;
[0046] Motion consistency / speed loss To prevent the predicted 3D trajectory from exhibiting jitter that does not conform to physical laws, constraints are imposed on the approximate displacement and velocity values of adjacent frames.
[0047]
[0048] Total loss function:
[0049]
[0050] in, and These are weight hyperparameters;
[0051] By backpropagating the total loss, the network weights of the lightweight BEVFormer are continuously optimized.
[0052] The high-speed tennis trajectory recognition and 3D reconstruction device based on monocular vision under limited computing power using the above method mainly includes:
[0053] The camera and site offline calibration module is used to perform a one-time offline calibration before the system is run.
[0054] The lightweight BEVFormer network module is suitable for devices with limited computing power. It takes a monocular image as input and outputs the 3D spatial coordinates of the tennis ball in the current frame.
[0055] Real-time inference and application module: Deployed on computing-limited devices, it relies solely on a single camera to push live data in real time, inputting a lightweight BEVFormer network and outputting continuous 3D coordinate points;
[0056] 3D trajectory output module: Filters and smooths the continuous 3D coordinate points output by the lightweight BEVFormer network processing module, calculates the flight trajectory of the tennis ball, and obtains the instantaneous speed and landing point.
[0057] Furthermore, the device also includes:
[0058] The high-precision ground truth training data generation module is used to acquire high-precision stereo data as ground truth data for network training.
[0059] The end-to-end network training module uses monocular continuous images as input and high-precision 3D data acquired by a binocular system as supervision signals to train the lightweight BEVFormer network end-to-end.
[0060] The main technical approaches of this solution include:
[0061] (1) System architecture: It relies solely on a monocular camera, combined with offline site / camera calibration, and uses high-precision binocular data as ground truth to train the end-to-end monocular 3D reconstruction network;
[0062] (2) Core algorithm and lightweight: Under limited computing power, a spatiotemporal attention mechanism (such as the BEVFormer architecture) is adopted to simultaneously handle the "detection" and "tracking" problems of high-speed and weak targets such as tennis balls, and directly output 3D coordinates;
[0063] (3) Temporal features to solve blur and occlusion: The temporal features of multiple consecutive frames of images are used to solve the problem of severe motion blur and occlusion of tennis balls caused by high-speed motion in a single frame.
[0064] Compared with the prior art, the beneficial effects of this disclosure are:
[0065] ① Extremely low hardware cost and computing power threshold: Compared to binocular / multi-camera systems, it eliminates the need for multiple expensive cameras and computing power equipment, requiring only a single ordinary camera. Through lightweight algorithm design, it can run smoothly on ordinary edge computing devices.
[0066] ② Extremely simple deployment method: It avoids the complexities of binocular joint calibration. Only a simple monocular camera and field line calibration is required during system initialization, which greatly improves the tolerance to site environment and installation location.
[0067] ③ Breaking through the ceiling of monocular accuracy: Compared with traditional monocular algorithms, this disclosure uses a spatiotemporal attention mechanism to solve the high-speed ghosting problem. Moreover, due to the use of high-precision binocular data for "dimensionality reduction" end-to-end training, its monocular 3D reconstruction accuracy far exceeds that of traditional methods that rely on simple physical assumptions. Attached Figure Description
[0068] The above and other objects, features and advantages of this disclosure will become more apparent from the more detailed description of exemplary embodiments of this disclosure taken in conjunction with the accompanying drawings, in which the same reference numerals generally represent the same components.
[0069] Figure 1 This displays the overall flowchart of the monocular tennis ball 3D reconstruction system under limited computing power.
[0070] Figure 2 A schematic diagram of the lightweight BEVFormer network structure;
[0071] Figure 3 This is a logic diagram for generating high-precision ground truth training data and for end-to-end training. Detailed Implementation
[0072] Preferred embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While preferred embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.
[0073] This disclosure provides a method for high-speed tennis trajectory recognition and 3D reconstruction based on monocular vision under limited computing power. It mainly solves the technical problems of existing high-precision tennis 3D trajectory reconstruction and speed measurement systems (such as binocular / multi-view vision systems) in digital scenarios of sports (tennis) events or daily training, such as high hardware costs, harsh deployment environments, and extremely high requirements for computing resources (unable to run on edge devices or devices with limited computing power).
[0074] Meanwhile, for monocular vision systems, this disclosure solves the technical problems that traditional algorithms are prone to motion blur, occlusion, and severe loss of depth information when facing high-speed, small moving targets such as tennis balls, which leads to low 3D reconstruction accuracy and easy loss of trajectory tracking.
[0075] In one exemplary implementation:
[0076] The overall architecture of a single-eye tennis ball 3D reconstruction system / method under limited computing power according to this disclosure is attached. Figure 1 As shown. Specifically, it includes the following key modules and steps:
[0077] 1. Offline camera and site positioning
[0078] Implementation: The system performs a one-time offline calibration before operation.
[0079] First, the intrinsic parameter matrix K (including focal lengths f_x, f_y and principal point coordinates c_x, c_y) and distortion coefficients of the monocular camera are obtained in advance using the Zhang Zhengyou calibration method combined with the chessboard grid.
[0080] Secondly, using the standard geometric dimensions of the tennis court (such as the physical length and intersection of the service line and sidelines), the corresponding pixel corner points in the image are extracted. The PnP (Perspective-n-Point) algorithm is used to calculate the external parameters of the camera relative to the tennis court (rotation matrix R and translation vector T) to establish a unified 3D world coordinate system.
[0081] 2. Generation of high-precision ground truth training data
[0082] Implementation plan: Set up a temporary "high-precision, high-frame-rate binocular / multi-view vision system" as the acquisition end. Simultaneously acquire video data during real tennis matches or training.
[0083] Using the high-precision triangulation of a binocular system, the high-precision 3D spatial coordinates of the tennis ball in each frame (frame t) are calculated, denoted as P. t_gt =(X t_gt Y t_gt Z t_gt ), serving as the Ground Truth (truth label) for network training. gt is an abbreviation for Ground Truth, P t_gt This represents the extremely accurate 3D spatial coordinates of the tennis ball in frame t in the real world, used to provide the model with the "correct answer" for the learning standard.
[0084] 3. Lightweight BEVFormer Network Construction
[0085] Implementation solution: The BEVFormer architecture, originally designed for autonomous driving, is pruned and modified to be lightweight and adaptable to devices with limited computing power.
[0086] BEVFormer is a landmark model in the field of autonomous driving for pure vision-based multi-camera 3D perception. It uses a spatiotemporal Transformer to directly convert multi-camera images into unified bird's-eye view (BEV) features, enabling end-to-end 3D detection, map segmentation, and other tasks. In this embodiment, a lightweight BEVFormer network structure is constructed based on this, as shown in the attached figure. Figure 2 As shown.
[0087] Input: A stream of N consecutive monocular video images.
[0088] Spatiotemporal attention mechanism:
[0089] Temporal attention: Instead of viewing single frames in isolation, the network integrates historical bird's-eye view features from the previous N-1 frames to "understand" the movement trend of the tennis ball. Even if the tennis ball in the current frame is severely blurred or occluded due to high-speed movement, the network can still "track" and extract it using historical trajectory features.
[0090] Spatial attention: Combining the camera intrinsic and extrinsic parameter matrices obtained by the calibration module, 2D image features are mapped to 3D voxel space through a deformable attention mechanism, focusing on the spatial region where the tennis ball may exist.
[0091] Output: Directly output the 3D spatial coordinates of the tennis ball in the current frame (frame t, 0).
[0092] The lightweight BEVFormer network structure in this embodiment is not the conventional standard structure for autonomous driving. Instead, it has undergone deep customization to address the challenges of "edge-limited computing power" and "high-speed, small targets like tennis balls." Its structural modifications mainly include:
[0093] (1) Lightweight replacement of the backbone network: The standard BEVFormer usually uses a large ResNet-101 or Swin-Transformer as the feature extraction layer. In order to adapt to the limited computing power, the backbone is replaced with a lightweight convolutional neural network (such as a variant of MobileNetV3 or RepVGG), which significantly reduces the number of network parameters and floating-point operations (FLOPs), ensuring that the real-time input requirements of high frame rate (such as above 60fps) can be achieved on edge devices.
[0094] (2) Spatial Attention Reconstruction for Small Targets: Targets (vehicles, people) in autonomous driving are relatively large, while tennis balls often only occupy a few pixels. Therefore, in this embodiment, the sampling strategy of multi-scale feature maps in the spatial attention module is adjusted: low-resolution deep features (such as 1 / 32 and 1 / 64 downsampling rates) that are not helpful for tennis ball detection are removed, and the computing power is concentrated entirely on high-resolution shallow feature maps of 1 / 4 and 1 / 8 (the original 1 / 16 feature map is replaced with a 1 / 4 feature map, and the number of attention heads and sampling points are reduced to improve the resolution of small targets while controlling the amount of computation). At the same time, the number of attention heads is reduced from the usual 8 to 4, and the number of sampling points per reference point is reduced from 4 to 2. Through these quantization prunings, the computational complexity of this module is greatly reduced while maintaining the recall rate of small targets.
[0095] (3) Dimensional reduction and customization of BEV bird's-eye view space (Voxel Space): Autonomous driving needs to perceive the surrounding 100-meter range, while a monocular tennis scene only needs to cover a standard tennis court and its surroundings (such as a 24-meter * 12-meter area). In this embodiment, the XY plane resolution of the BEVGrid is refined to 0.1 meters / grid to capture the high-speed, minute displacement of the tennis ball.
[0096] More importantly, considering the dramatic changes in the tennis ball's height (Z-axis), this embodiment employs a "non-equidistant refined design" for voxel division along the Z-axis: a dense division of 0.1 meters is used in the 0-1 meter height range where the tennis ball has a high probability of hitting the ground; while a sparse division of 0.5 meters is used in the 1-4 meter flight range. This design significantly improves the 3D reverse calculation accuracy of the tennis ball's landing point and bounce height without increasing the total Voxel computation.
[0097] (4) Operator-level optimization for edge BPU / NPU adaptation: To enable deployment on real-world edge-constrained devices (such as heterogeneous computing platforms like the Horizon RDK series), this embodiment specifically replaces certain operators in the network. For example, the deformable convolution and GridSample operators in the native BEVFormer, which are extremely unfriendly to edge BPUs, are replaced with a bilinear interpolation that supports efficient hardware acceleration, combined with a conventional Conv2D approximation (a trade-off between accuracy and speed). Simultaneously, the GELU activation function in the network is replaced with ReLU, which is more compatible with INT8 quantization, and LayerNorm is fused and reconstructed into BatchNorm. These operator-level reconstructions enable the model to perfectly run the quantization deployment of the edge compiler, thereby strictly compressing the single-frame inference latency to the millisecond level.
[0098] 4. End-to-end network training
[0099] Implementation scheme: Monocular continuous images are used as input to the lightweight BEVFormer, and high-precision 3D data acquired by the binocular system in step 2 are used as supervision signals (Loss calculation benchmark). Through a large amount of data-driven learning, the monocular network implicitly learns perspective projection rules, aerodynamic characteristics, and depth inference capabilities.
[0100] To ensure the monocular network accurately learns 3D spatial position, this embodiment employs a joint loss function for end-to-end supervised training. The loss function not only constrains the absolute coordinate error of a single frame but also introduces a temporal velocity consistency constraint, as detailed in the following formula:
[0101] (1) 3D coordinate regression loss ( The error between the predicted coordinates and the stereo ground truth coordinates is calculated using Smooth L1 Loss.
[0102]
[0103] T represents the time span of the sliding window used to calculate coordinate regression loss or instantaneous velocity. If the calculation is based on a sliding window (e.g., 3-5 frames), it is the number of frames corresponding to that sliding window.
[0104] (2) Motion consistency / velocity loss To prevent the predicted 3D trajectory from exhibiting jitter that does not conform to physical laws, constraints are imposed on the displacement (i.e., the approximate velocity value) of adjacent frames:
[0105]
[0106] (3) Total loss function:
[0107]
[0108] in, and This is the weight hyperparameter. The network weights of the lightweight BEVFormer are continuously optimized by backpropagating this total loss.
[0109] 5. Real-time inference and applications
[0110] Implementation: Deployed on a device with limited computing power. It relies solely on a single camera to stream data in real-time, inputting it into a lightweight BEVFormer network, which outputs continuous 3D coordinate points in milliseconds. The continuous 3D coordinate points output by the lightweight BEVFormer network processing module are then input into a 3D trajectory output module. Based on timestamps and 3D coordinate differences, and after smoothing algorithms such as Kalman filtering, the trajectory of the tennis ball is calculated to obtain its instantaneous velocity and landing point.
[0111] (1) Instantaneous velocity calculation: Since single-frame calculation at high frame rates is easily affected by small noise, this method adopts a time sliding window. Calculate the instantaneous velocity (e.g., at intervals of 3-5 frames). Assume the video frame rate is FPS, and the time interval... The instantaneous spatial velocity at frame t. The calculation formula is:
[0112]
[0113] (2) High-precision calculation of impact point (point of impact):
[0114] The actual landing point of a tennis ball often occurs between two frames, and simply taking the frame with the lowest Z-axis value as the landing point results in a large error. This system uses a method of Z-axis extreme value detection combined with fitting to find the frame tt with the lowest Z-axis value, and performs 3D curve fitting on the trajectory of the m frames before and n frames after tt respectively. The intersection point obtained from the two curve fittings is taken as the landing point.
[0115] In this embodiment, with only "one eye" and "limited brain computing power" (limited computing power device), a special memory and attention mechanism (BEVFormer network) is used to not only locate the afterimage but also accurately determine its 3D spatial location.
[0116] The above technical solutions are merely exemplary embodiments of the present invention. For those skilled in the art, based on the application methods and principles disclosed in the present invention, it is easy to make various types of improvements or modifications, and not limited to the methods described in the specific embodiments of the present invention. Therefore, the methods described above are merely preferred and not restrictive.
Claims
1. A method for high-speed tennis trajectory recognition and 3D reconstruction based on monocular vision under limited computing power, characterized in that, Includes the following steps: S1, perform a one-time offline calibration before running; S2, For devices with limited computing power, a lightweight BEVFormer network is constructed. The lightweight BEVFormer network takes a monocular image as input and the 3D spatial coordinates of the tennis ball in the current frame as output. S3, through a single-channel image acquisition device, pushes the stream in real time, inputs the trained lightweight BEVFormer network, and outputs continuous 3D coordinate points; S4. After filtering and smoothing the continuous 3D coordinate points output in step S3, calculate the flight trajectory of the tennis ball to obtain the instantaneous speed and landing point. It also includes the following steps: Acquire high-precision stereo data as ground truth data for network training; The lightweight BEVFormer network is trained using monocular continuous images as input and high-precision binocular 3D data as supervision signals. Specific training methods include: A joint loss function is used for end-to-end supervised training. The specific formula for calculating the loss function is as follows: 3D coordinate regression loss The error between the predicted coordinates and the stereo ground truth coordinates is calculated using Smooth L1 Loss. , T represents the time span of the sliding window used to calculate coordinate regression loss or instantaneous velocity, P t_gt gt represents the true 3D spatial coordinates of the tennis ball in the real world at frame t; t represents frame t, and gt represents the true value. Motion consistency / speed loss To prevent the predicted 3D trajectory from exhibiting jitter that does not conform to physical laws, constraints are imposed on the approximate displacement and velocity values of adjacent frames. , Total loss function: , in, and These are weight hyperparameters; By backpropagating the total loss, the network weights of the lightweight BEVFormer are continuously optimized.
2. The method according to claim 1, characterized in that, Step S1 specifically includes: The intrinsic parameter matrix K of the monocular camera, including the focal length f, was obtained in advance using Zhang Zhengyou's calibration method combined with a checkerboard pattern. x f y and principal point coordinates c x c y and distortion coefficient; Using the standard geometric dimensions of a tennis court, the corresponding pixel corners in the image are extracted. The PnP algorithm is used to calculate the camera's extrinsic parameters relative to the tennis court, including the rotation matrix R and translation vector T, to establish a unified 3D world coordinate system.
3. The method according to claim 1, characterized in that, Step S2 specifically includes: The BEVFormer standard architecture used for autonomous driving is pruned and modified to be lightweight and adapted to devices with limited computing power. The input of the lightweight BEVFormer network is a continuous stream of N frames of monocular video images, and the output is the 3D spatial coordinates of the tennis ball in the current frame. The lightweight BEVFormer network employs a spatiotemporal attention mechanism, wherein: Temporal attention: By fusing features from the historical bird's-eye view of the previous N-1 frames, the movement trend of the tennis ball can be obtained; Spatial attention: Combining the camera intrinsic and extrinsic parameter matrices obtained in calibration step S1, 2D image features are mapped to 3D voxel space through a deformable attention mechanism, focusing on the spatial region where the tennis ball may exist.
4. The method according to claim 3, characterized in that, In step S2, the method for lightweight pruning and modification of the BEVFormer standard architecture includes one or more of the following: (1) Lightweight replacement of the backbone network, including: replacing the feature extraction layer Backbone of the standard BEVFormer with a lightweight convolutional neural network; (2) For small targets, perform "spatial attention" reconstruction, including: in the spatial attention module, remove low-resolution deep features and retain high-resolution shallow features; at the same time, reduce the number of attention heads and the number of sampling points for each reference point; (3) The BEV bird's-eye view space is reduced in size and customized, including: the perception range only covers the standard tennis court and its surroundings; the resolution of the XY plane of the BEVGrid is refined; and the voxel division of the Z axis adopts "non-equidistant refined design". (4) For edge devices with limited computing power, special operators in the network were replaced to compress the single-frame inference delay time.
5. The method according to claim 1, characterized in that, In step S4, the specific methods for calculating the instantaneous velocity and predicting the landing point include: (1) Instantaneous velocity calculation: Instantaneous velocity is calculated using a time sliding window. Assume the video frame rate is FPS and the time interval is... The instantaneous spatial velocity at frame t The calculation formula is: , Indicates the number of frames contained in the sliding window; (2) Calculation of landing point: The method of Z-axis extreme value detection combined with fitting is used to find the frame tt with the lowest z-axis, and 3D curve fitting is performed on the trajectory of the m frames before and n frames after tt respectively. The intersection point obtained by the two curve fittings is taken as the landing point.
6. The method according to claim 1, characterized in that, The method for acquiring high-precision stereo data specifically includes: In real tennis matches or training, a binocular vision system is used to simultaneously collect binocular video data. Using the high-precision triangulation of a binocular system, the high-precision 3D spatial coordinates P of each frame of tennis ball are calculated. t_gt =(X t_gt ,Y t_gt Z t_gt ), which serves as the ground truth label for network training, where t represents the t-th frame and gt represents the ground truth.
7. A device for high-speed tennis trajectory recognition and 3D reconstruction based on monocular vision under limited computing power, using the method described in any one of claims 1-6, characterized in that, include: The camera and site offline calibration module is used to perform a one-time offline calibration before the system is run. The lightweight BEVFormer network module is suitable for devices with limited computing power. It takes a monocular image as input and outputs the 3D spatial coordinates of the tennis ball in the current frame. Real-time inference and application module: Deployed on computing-limited devices, it relies solely on a single camera to push live data in real time, inputting a lightweight BEVFormer network and outputting continuous 3D coordinate points; 3D trajectory output module: Filters and smooths the continuous 3D coordinate points output by the lightweight BEVFormer network processing module, calculates the flight trajectory of the tennis ball, and obtains the instantaneous speed and landing point; The high-precision ground truth training data generation module is used to acquire high-precision stereo data as ground truth data for network training. The end-to-end network training module uses monocular continuous images as input and high-precision 3D data acquired by a binocular system as supervision signals to train the lightweight BEVFormer network end-to-end.
Citation Information
Patent Citations
Intelligent court motion information acquisition system and method based on monocular vision
CN110910489A
Three-dimensional tennis track real-time reconstruction method and system based on multiple cameras
CN120997256A