A Dual Vision Video Processing Method and System
Through the dual-visual video processing method and combined with billiards movement trajectory prediction, accurate auxiliary judgment of billiards collisions is achieved, the judgment disputes in the existing technology are solved, and the image processing effect is improved.
Patent Information
- Application Number
- CN202510246736.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The prior art cannot fully realize the collision monitoring between balls in billiards, resulting in controversy among the referees.
The dual-visual video processing method is used to acquire images through binocular cameras, and image preprocessing, dual-visual three-dimensional reconstruction, motion direction and velocity acquisition, motion trajectory prediction and matching degree verification are performed to assist in judging collisions.
Accurate auxiliary judgment of billiards collisions is achieved, the accuracy and reliability of the judges are improved, and the deblurring effect and robustness are improved during image processing.
Smart Images

Figure CN119741334B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of dual-vision video processing, and in particular to a dual-vision video processing method and system. Background Art
[0002] Billiards is an indoor sport that uses a cue to strike a cue ball to collide with target balls and sink the target balls into the pockets, which is very popular among people. The billiard collision rule is one of the core rules in billiards, mainly involving the collision behavior between the cue ball and target balls or other balls.
[0003] Currently, in the adjudication of billiards, a method combining manual visual inspection and video adjudication is adopted. Although the current video system can track the movement trajectory of the balls, due to the camera viewfinder technology used in the system, ordinary cameras cannot fully monitor the collision of balls due to the limitations of the acquisition speed and angle. Therefore, there are disputes in the adjudication of whether some balls collide. Summary of the Invention
[0004] To solve the technical problems of billiard collision adjudication in the prior art, the present invention provides a dual-vision video processing method and system.
[0005] The present invention is realized through the following technical solutions:
[0006] A dual-vision video processing method for billiard collision monitoring, comprising the following steps:
[0007] S1: Video image acquisition; obtaining an image on the billiard table using a binocular camera;
[0008] S2: Image preprocessing; including denoising and clarification, image segmentation, and target extraction;
[0009] S3: Dual-vision three-dimensional reconstruction; which includes:
[0010] S31: Camera calibration, taking multiple groups of images using a calibration board, and calculating camera parameters through a calibration algorithm;
[0011] S32: Image correction, stereo matching;
[0012] S33: Optimization of matching model parameters;
[0013] S34: Depth calculation, generating a point cloud;
[0014] S35: Three-dimensional reconstruction and visualization;
[0015] S4: Obtaining the movement direction and speed of the billiard balls according to the image after three-dimensional reconstruction and the number of frames of the camera;
[0016] S5: Predict the motion trajectory, perform trajectory matching degree verification, and then perform collision auxiliary judgment.
[0017] Further, the step S2 includes:
[0018] S21: Image denoising and clarification processing; including non-local means denoising, establishing a dynamic position estimation model for estimating the dynamic position offset and deblurring the image based on the dynamic position offset.
[0019] S22: Image segmentation to obtain the image of the billiard table part.
[0020] S23: Target extraction, obtaining the ball contour on the billiard table and locating the center point of the ball.
[0021] Further, the dynamic position estimation model consists of two neural networks, the first network is used to initially estimate the offset of the dynamic position, and the second network is used to accurately correct the dynamic position offset; the loss function in the training process of the dynamic position estimation model is as follows:
[0022]
[0023] Among them, are the regularization loss, cross-entropy loss, total variation loss, and mean square error loss respectively, are the corresponding weights;
[0024] The deblurring process uses a self-supervised convolutional neural network model, and the loss function of the model is as follows:
[0025]
[0026] Among them, is the multi-scale content loss function, is the multi-scale frequency reconstruction loss function, is the weight coefficient.
[0027] Further, the step S3 includes a stereo matching process, using the semi-global algorithm for stereo matching, and combining the Census transform and mutual information to determine the matching cost function.
[0028] Further, the matching cost function is expressed as follows:
[0029]
[0030] Among them, and are the control parameters of the Census transform and mutual information respectively, is the cost function based on the Census transform, is a cost function based on mutual information, is a normalization function.
[0031] Furthermore, the optimization of the matching model parameters includes the following steps:
[0032] a. Randomly select n samples from the dataset P as the initial subset to initialize the stereo matching model;
[0033] b. Set an error threshold, and in the remainder of the dataset P, filter out the samples with an error less than the threshold from the model M, denoted as the set S, and merge S with the initial subset to form the set C;
[0034] c. Based on the set C, re-estimate the stereo matching parameters and update the stereo matching; repeat the above steps to iteratively optimize the model;
[0035] d. When the number of sampling times reaches the preset value, select the model with the largest inlier set C from all the generated models as the final result.
[0036] Furthermore, the step S5 includes:
[0037] S51: Predict the motion trajectory according to the obtained motion direction and speed of the ball;
[0038] S52: Calculate the matching degree between the predicted motion trajectory and the actual motion trajectory. The matching degree includes the matching degree of the number of trajectory points and the matching degree of the trajectory length;
[0039] S53: Perform auxiliary collision judgment according to the matching degree verification result;
[0040] Judge whether the target ball has a collision contact according to the comparison between the matching degree verification result and the set threshold.
[0041] Furthermore, the step S51 includes: Based on the LSTM network model, input the motion direction, speed, and position of all the balls at the previous n moments obtained into the LSTM network model, and output the motion direction, speed, and position of the balls at the next m moments until the speeds of all the balls are 0, so as to form the motion trajectory prediction result of all the balls.
[0042] The present invention also provides a dual-vision video processing system, based on the above-mentioned dual-vision video processing method, which includes:
[0043] A video image acquisition module, which is used to obtain an image on the billiard table by using a binocular camera;
[0044] An image preprocessing module, which is used to denoise and clarify the image, perform image segmentation, and extract the target;
[0045] A dual-vision three-dimensional reconstruction module, which is used to perform three-dimensional reconstruction on the image information collected by a binocular camera;
[0046] A billiard motion parameter acquisition module, which is used to acquire the motion direction and speed of a billiard ball;
[0047] A collision assistance judgment module, which is used to predict the motion trajectory, perform trajectory matching degree verification, and then perform collision assistance judgment.
[0048] In addition, to achieve the above object, the present invention also provides a computer-readable storage medium, on which program instructions for a dual-vision video processing method are stored, and the program instructions for the dual-vision video processing method can be executed by one or more processors to implement the steps of the dual-vision video processing method as described above.
[0049] Compared with the prior art, the beneficial effects of the present invention are:
[0050] The present invention realizes the auxiliary judgment for billiard collisions by optimizing the processing of dual-vision videos and combining the prediction of billiard motion trajectories to help the referee make judgments; moreover, in the process of processing video images, the present invention uses a dynamic position estimation model to deblur the images, strengthening the deblurring effect; the present invention also optimizes the stereo matching, adopts an energy function combining the Census transform and mutual information, improves the robustness under illumination changes, and improves the result accuracy. Description of the Drawings
[0051] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0052] Figure 1 It is a schematic flowchart of the dual-vision video processing method according to an embodiment of the present application. Detailed Embodiments
[0053] The embodiments of the present invention will be described in detail below with reference to the drawings.
[0054] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0055] It should also be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. The diagrams only show the components related to the present invention, rather than being drawn according to the number, shape, and size of the components in actual implementation. The form, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the layout form of its components may also be more complex.
[0056] As Figure 1 shown, a dual-vision video processing method includes the following steps:
[0057] S1: Video image acquisition; using a binocular camera to obtain an image of a billiard table;
[0058] S2: Image preprocessing; including denoising and sharpening, image segmentation, and target extraction;
[0059] The specific steps include:
[0060] S21: Image denoising and sharpening processing; due to factors such as the exposure environment, the camera inevitably generates noise, so it is necessary to denoise the image; since the billiard movement process will cause motion blur in the image, it is also necessary to perform sharpening processing; the steps are as follows:
[0061] a. Based on non-local means denoising, the calculation steps are as follows:
[0062]
[0063] where N( ) is the non-local means denoising function, q is the image before processing, u is the denoised image, is the similarity coefficient, indicating the similarity between pixel points i and j. Optionally, the similarity is obtained through the geometric distance;
[0064] The above denoising algorithm utilizes the redundant information in the image, can well preserve the texture details in the image while removing the noise, and has the advantages of simplicity, good performance, and easy improvement.
[0065] Optionally, in order to further improve the denoising effect, the present invention performs multi-scale decomposition on the image by using wavelet transform, and then applies non-local means denoising to each scale. Specifically, multi-level wavelet decomposition is performed on the noisy image, non-local means denoising is applied to each scale after decomposition, and finally the image is reconstructed. The multi-level wavelet threshold processing method can effectively eliminate the high-frequency noise in the noisy image.
[0066] b. Establish a dynamic position estimation model for estimating the dynamic position offset situation;
[0067] The dynamic position estimation model consists of two neural networks, a first network and a second network. The first network is used to preliminarily estimate the offset of the dynamic position, and the second network is used to accurately correct the dynamic position offset;
[0068] Specifically, the first network is a cross-attention neural network, which adaptively learns the important features in different channels and spaces and estimates the dynamic position information of the motion-blurred image; the second network is a BP neural network.
[0069] The training process of the network is to give a blurred image in a training set and a corresponding clear image. The dynamic position estimation model takes the blurred image as the input and outputs the dynamic position offset of the clear image, including horizontal and vertical directions.
[0070] Optionally, the loss function in the model training process is as follows:
[0071]
[0072] Among them, are the regularization loss, cross-entropy loss, total variation loss, and mean square error loss respectively, are the corresponding weights.
[0073] c. Deblur the image based on the dynamic position offset;
[0074] The deblurring process uses a self-supervised convolutional neural network model, and the loss function of the model is as follows:
[0075]
[0076] Among them, is the multi-scale content loss function, is the multi-scale frequency reconstruction loss function, is the weight coefficient.
[0077] S22: Image segmentation to obtain the image of the billiard table part;
[0078] Optionally, segment the preprocessed image based on the watershed image segmentation algorithm to remove the background and obtain the image of the billiard table part.
[0079] S23: Target extraction, obtain the ball contours on the billiard table and locate the center points of the balls.
[0080] S3: Dual-vision three-dimensional reconstruction, including the following steps:
[0081] S31: Camera calibration, take multiple groups of images using a calibration board, and calculate the camera parameters through the calibration algorithm.
[0082] S32: Image correction, stereo matching;
[0083] Calculate the correction mapping using the calibration result, apply the correction transformation to the left and right images; find corresponding points in the left and right images, calculate the disparity, and obtain the disparity map;
[0084] Optionally, the stereo matching adopts the semi-global algorithm (Semi-Global Matching), including matching cost calculation, matching cost aggregation, disparity calculation, and disparity post-processing.
[0085] Furthermore, the present invention adopts the following energy function:
[0086]
[0087] Where r is the direction of cost traversal calculation, n is all directions, represents the matching cost in each direction when the disparity value at point p is d, and is specifically expressed as:
[0088]
[0089] Where, is the matching cost function with depth d at point p, is the traversal cost aggregation in the direction r with depth d at point p, is the variable cost aggregation of the previous pixel point of point p in the direction r with depth d, is the variable cost aggregation of the previous pixel point of point p in the direction r with depth d + 1, is the variable cost aggregation of the previous pixel point of point p in the direction r with depth d - 1, is the variable cost aggregation of the previous pixel point of point p in the direction r with depth i, is the minimum value of the variable cost aggregation of the previous pixel point of point p in the direction r; , are the penalty factors for the smooth constraint and the edge constraint respectively, and k and i are variable parameters.
[0090] The present invention also combines the Census transform and mutual information to determine the matching cost function, which is expressed as follows:
[0091]
[0092] Among them, , are the control parameters of the Census transform and mutual information respectively, is the cost function based on the Census transform, is the cost function based on mutual information, is the normalization function, and the expression is as follows:
[0093]
[0094] When both c and are positive numbers, the value range of the function is [0, 1].
[0095]
[0096] is the Hamming distance with depth d, , are two bit strings after the census transform respectively.
[0097]
[0098]
[0099] Among them, is the cost of pixel point p for disparity d, is the gray value of point p in the reference image b and point q in the matching image m corrected by the disparity map D, is the unit of gray mutual information; , are the gray distributions of the images, is the gray joint entropy.
[0100] The present invention uses gray mutual information as the cost metric for image matching. This method extracts the matching cost value from the gray entropy list and gray joint entropy distribution by analyzing the pixel gray value data, and transforms the macroscopic parameters into effective evaluation indicators of pixel-level data, thereby significantly improving the matching accuracy.
[0101] The present invention adopts an energy function combining the Census transform and mutual information, which improves the robustness under illumination changes.
[0102] S33: Matching model parameter optimization
[0103] Since there will be abnormal data in stereo matching, in order to improve the result accuracy, the present invention performs matching optimization based on the following steps:
[0104] a. Randomly select n samples from the dataset P as the initial subset for initializing the stereo matching model.
[0105] b. Set the error threshold, and in the complement of the dataset P, screen out the samples with an error less than the threshold from the model M, denoted as the set S. Combine S with the initial subset to form the set C.
[0106] c. Based on the set C, re-estimate the stereo matching parameters and update the stereo matching. Repeat the above steps to iteratively optimize the model.
[0107] d. When the sampling times reach the preset value, select the model with the largest inlier set C from all the generated models as the final result.
[0108] S34: Depth calculation to generate point cloud;
[0109] Calculate the depth value of each pixel according to the disparity map and convert the depth map into a three-dimensional point cloud.
[0110] S35: 3D reconstruction and visualization
[0111] Convert the point cloud into a 3D model and perform visualization.
[0112] Optionally, use surface reconstruction algorithms such as Poisson reconstruction and Delaunay triangulation to generate a mesh model, and use image processing tools for visualization.
[0113] S4: Obtain the movement direction and speed of the billiard ball
[0114] According to the image after 3D reconstruction combined with the number of frames of the camera, obtain the movement direction and speed of the billiard ball; including the following steps:
[0115] a. Initialize the parameters, and perform brightness equalization processing and normalization on the image;
[0116] b. Construct the loss function and optimize the YOLO network parameters through backpropagation for network training;
[0117] c. Input the image after 3D reconstruction into the convolutional layer of the YOLO network to obtain the features of each frame of the image and classify the feature information;
[0118] d. Obtain the movement direction and speed of the billiard ball on each frame of the image according to the network output.
[0119] S5: Predict the motion trajectory, perform trajectory matching degree verification, and then perform collision auxiliary judgment;
[0120] S51: Predict the motion trajectory based on the obtained motion direction and speed;
[0121] Specifically, based on the LSTM network model, input the motion directions, speeds, and positions of all balls at the previous n moments obtained into the LSTM network model, and output the motion directions, speeds, and positions of the balls at the next m moments until the speeds of all balls are 0, thereby forming the prediction result of the motion trajectories of all balls.
[0122] S52: Calculate the matching degree between the predicted motion trajectory and the actual motion trajectory; Specifically, calculate the matching degree according to the following formula, which includes the matching degree of the number of trajectory points and the matching degree of the trajectory length :
[0123]
[0124] where the number of matching trajectory points is the number of trajectory points that match between the predicted motion trajectory and the actual motion trajectory; the total number of trajectory points in the entire path refers to the total number of points of the actual motion trajectory.
[0125]
[0126] S53: Perform auxiliary collision judgment according to the matching degree verification result;
[0127] Compare the matching degree verification result with the set threshold to judge whether the target ball has a collision contact.
[0128] In this embodiment, it realizes providing auxiliary judgment for billiard collisions through the processing of dual-vision videos and combining with the prediction of billiard motion trajectories to help the referee make judgments; moreover, in the process of processing video images, the present invention uses a dynamic position estimation model to deblur the images, strengthening the deblurring effect; the present invention also optimizes the stereo matching, adopts an energy function combining the Census transform and mutual information, improves the robustness under illumination changes, and improves the result accuracy.
[0129] The embodiment of the present invention also proposes a dual-vision video processing system, based on the above-mentioned dual-vision video processing method, including:
[0130] A video image acquisition module, which is used to obtain images on the billiard table by using a binocular camera;
[0131] An image preprocessing module, which is used to denoise and clarify the images, perform image segmentation, and extract targets;
[0132] A dual-vision three-dimensional reconstruction module, which is used to perform three-dimensional reconstruction on the image information collected by the binocular camera;
[0133] A billiard motion parameter acquisition module, which is used to acquire the motion direction and speed of the billiard ball;
[0134] A collision assistance judgment module, which is used to predict the motion trajectory, perform trajectory matching degree verification, and then perform collision assistance judgment.
[0135] In addition, an embodiment of the present invention further provides a computer-readable storage medium, on which program instructions of the dual-vision video processing method are stored, and the program instructions of the dual-vision video processing method can be executed by one or more processors to implement the steps of the dual-vision video processing method as described above.
[0136] The above embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A dual-vision video processing method, characterized in that: For billiards video processing, including the following steps: S1: Video image acquisition; using a binocular camera to obtain images on the billiard table; S2: Image preprocessing; Including denoising and clarity, image segmentation, and target extraction; S3: Dual-vision 3D reconstruction; including: S31: Camera calibration, using the calibration board to take multiple sets of images and calculate the camera parameters through the calibration algorithm; S32: Image correction, stereo matching; S33: matching model parameter optimization; S34: depth calculation, generating point cloud; S35: 3D reconstruction and visualization; S4: Obtain the moving direction and speed of the billiard ball according to the three-dimensional reconstructed image and the frame number of the camera; S5: Predict the motion trajectory, perform trajectory matching verification, and then perform collision auxiliary judgment; The step S2 comprises: S21: image denoising and clarity processing; including denoising based on non-local mean, establishing a dynamic position estimation model for estimating the dynamic position offset and deblurring the image based on the dynamic position offset; S22: image segmentation, obtaining a partial image of the billiard table top; S23: target extraction, obtaining the ball outline on the billiard table and locating the center point of the ball; The dynamic position estimation model includes two neural networks, a first network and a second network. The first network is used to preliminarily estimate the dynamic position offset, and the second network is used to accurately correct the dynamic position offset. The loss function of the dynamic position estimation model training process is as follows: Among them, L1~L4 are regularization loss, cross entropy loss, total variation loss and mean square error loss, respectively, α i is the corresponding weight; The deblurring process uses a self-supervised convolutional neural network model, and the loss function of the model is as follows: L v =L CON +βL MS Among them, L CON is the multi-scale content loss function, L MS is the multi-scale frequency reconstruction loss function, and β is the weight coefficient.
2. The dual vision video processing method according to claim 1, characterized in that: The step S3 includes a stereo matching process, which uses a semi-global algorithm to perform stereo matching and combines Census transformation and mutual information to determine a matching cost function.
3. The dual vision video processing method according to claim 2, characterized in that: The matching cost function is expressed as follows: C(p,d)=δ(C census (p,d),e census )+δ(C mi (p,d),e mi ) Among them, ε census , ε mi are the control parameters of Census transformation and mutual information, C census (p, d) is the cost function based on Census transformation, C mi (p,d) is the cost function based on mutual information, and δ() is the normalization function.
4. The dual vision video processing method according to claim 1, characterized in that: The matching model parameter optimization comprises the following steps: a. Randomly select n samples from the data set P as the initial subset to initialize the stereo matching model; b. Set an error threshold, and select samples whose error with model M is less than the threshold from the residual set of data set P, record it as set S, and merge S with the initial subset to form set C; c. Based on set C, re-estimate stereo matching parameters and update stereo matching; repeat the above steps to iteratively optimize the model; d. When the sampling times reaches the preset value, the model with the largest internal point set C is selected from all generated models as the final result.
5. The dual vision video processing method according to claim 1, characterized in that: The step S5 comprises: S51: predicting the movement trajectory according to the obtained movement direction and speed of the ball; S52: Calculating the matching degree between the predicted motion trajectory and the actual motion trajectory, where the matching degree includes the trajectory point matching degree and the trajectory length matching degree; S53: Perform auxiliary collision judgment according to the matching degree verification result; By comparing the matching verification result with the set threshold, it is determined whether the target ball has collided or not.
6. The dual vision video processing method according to claim 5, characterized in that: The step S51 includes: based on the LSTM network model, the movement direction, speed and position of all balls in the previous n moments are input into the LSTM network model, and the movement direction, speed and position of the balls in the next m moments are output until the speed of all balls is 0, thereby forming the movement trajectory prediction results of all balls.
7. A dual-vision video processing system, based on the dual-vision video processing method according to any one of claims 1 to 6, comprising: A video image acquisition module, which is used to obtain images on the billiard table using a binocular camera; Image preprocessing module, which is used to remove noise, clarify images, segment images, and extract targets; A dual-vision 3D reconstruction module, which is used to perform 3D reconstruction of image information collected by the binocular camera; A billiards motion parameter acquisition module, which is used to obtain the motion direction and speed of the billiards; The collision auxiliary judgment module is used to predict the motion trajectory, perform trajectory matching verification, and then perform collision auxiliary judgment.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program instructions of the dual-vision video processing method, and the program instructions of the dual-vision video processing method can be executed by one or more processors to implement the steps of the dual-vision video processing method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Rigid-body collision track prediction display unit
CN104376154A
Ball motion trail prediction method and system, electronic equipment and storage medium
CN115965658A