Video acquisition method and system based on deep learning acceleration
By collecting multi-dimensional data in real time and using deep learning acceleration models, combined with dynamic threshold adjustment, it can accurately determine shooting intentions and events of interest, and screen out videos of exciting moments. This solves the problems of low acquisition efficiency and low accuracy in existing technologies, and achieves efficient and accurate video acquisition.
Patent Information
- Application Number
- CN202510814499.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-06-18
AI Technical Summary
Existing technologies have low efficiency and accuracy in video acquisition during football matches, and are unable to meet real-time and comprehensive requirements. They rely on complex models and single data, resulting in low acquisition efficiency and accuracy.
By collecting multi-dimensional key data in real time, such as the coordinates of the key points of the dribbling player's legs, the speed of the football, the shooting distance, etc., combined with the deep learning acceleration model and dynamic threshold adjustment mechanism, the shooting intention and the type of event of interest are determined, and the video of the wonderful moments is screened out.
It achieves accurate determination of shooting intentions and event types of interest, screens out video clips that match the actual level of excitement, improves acquisition efficiency and accuracy, avoids the one-sidedness of single-factor judgment, and enhances the pertinence and effectiveness of video screening.
Smart Images

Figure CN120708129A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a video acquisition method and system based on deep learning acceleration. Background Art
[0002] In football matches, accurately capturing highlights is crucial for enhancing the audience experience, optimizing content production, and facilitating tactical analysis. As football continues to expand its global influence, audience demand for high-quality, real-time highlights is growing. Traditional methods, which rely on manual labor or simple algorithms, suffer from low efficiency and accuracy, making them inadequate for modern competitions. The development of deep learning technology offers a new approach to addressing this challenge.
[0003] Patent document CN110378245A discloses a deep learning-based football match behavior recognition method, device, and terminal device, including: obtaining a football match video to be recognized; dividing the football match video into N video segments, and extracting a frame of image from each video segment as an input image, where N is an integer greater than 1; using a preset deep learning network model to process the input image to obtain a behavior recognition result corresponding to the football match video, wherein the deep learning network model is composed of a cascade of an Inception network model and a three-dimensional ResNet network model, the Inception network model is used to learn the relationship between pixels in each frame of the input image, and the three-dimensional ResNet network model is used to learn the relationship between each frame of the input image.
[0004] It can be seen that the football match behavior recognition method based on deep learning has the following problems: the deep learning network model is composed of a cascade of the Inception network model and the three-dimensional ResNet network model. Such a complex model requires a lot of computing resources and time; it cannot meet the real-time requirements when performing video processing and behavior recognition; insufficient training data or low data quality will affect the recognition accuracy of the model; it relies on a single video screen analysis, the analysis is incomplete and the accuracy is insufficient. Summary of the Invention
[0005] To this end, the present invention provides a video acquisition method and system based on deep learning acceleration, which is used to overcome the problems of low acquisition efficiency and low accuracy in the existing technology due to over-reliance on complex models and single data through multi-dimensional data analysis, deep learning acceleration model, and dynamic threshold adjustment mechanism.
[0006] To achieve the above objectives, the present invention provides a video acquisition method based on deep learning acceleration, comprising:
[0007] Real-time data collection of several key leg coordinates of the dribbling player, the speed of the ball, the shooting distance, the movement angle, the goalkeeper's distance from the goal line, the moving speed, the audience's sound frequency, and the commentator's voice pitch during a football match;
[0008] determining the presence of a shooting intention based on the leg key points and the shooting distance;
[0009] Determine the event type of interest based on the shooting distance, the preset goal width, the movement speed, the movement angle, the moving speed, the goalkeeping distance, and a preset comprehensive determination threshold;
[0010] Collect a number of videos of interest based on the deep learning acceleration model and the type of the event of interest;
[0011] Filtering a number of wonderful moment videos from all the videos of interest according to the sound frequency, the pitch, and a preset relevance threshold;
[0012] Adjusting the preset relevance threshold according to the number of all the videos of interest, the number of all the videos of wonderful moments within the preset adjustment time, and the comprehensive determination factor to obtain an adjusted relevance threshold, or adjusting the preset comprehensive determination threshold to obtain an adjusted comprehensive determination threshold;
[0013] The wonderful moment video re-determined based on the adjusted relevance threshold or the adjusted comprehensive determination threshold is stored in a video database.
[0014] Furthermore, determining the event type of interest based on the shooting distance, the preset goal width, the movement speed, the movement angle, the moving speed, the goalkeeping distance, and a preset comprehensive determination threshold includes:
[0015] Calculating a goal angle based on the shooting distance and a preset goal width;
[0016] Calculating a goal matching degree according to the movement speed, the movement angle, and the goal angle;
[0017] When the goal matching degree is greater than a preset matching degree threshold, determining that a concern event occurs based on the shooting distance, the goal matching degree, the moving speed, the goalkeeping distance, and a preset comprehensive determination threshold;
[0018] The type of event of interest is determined according to the goal matching degree and the defensive capability factor.
[0019] Furthermore, determining whether an event of interest occurs based on the shooting distance, the goal matching degree, the moving speed, the goalkeeping distance, and a preset comprehensive determination threshold includes:
[0020] Calculate a defensive capability factor based on the moving speed and the goalkeeping distance;
[0021] A comprehensive determination factor is calculated based on the goal matching degree, the defensive ability factor, and the shooting distance;
[0022] When the comprehensive determination factor is greater than the preset comprehensive determination threshold, it is determined that an event of interest occurs.
[0023] Furthermore, determining the event type of interest based on the goal matching degree and the defensive capability factor includes:
[0024] Normalizing the goal matching degree to obtain a matching degree normalization value;
[0025] Normalizing the defense capability factor to obtain a defense normalized value;
[0026] When the matching degree normalized value is greater than the defense normalized value, determining that the type of the event of interest is a shooting event;
[0027] When the matching degree normalized value is less than or equal to the defense normalized value, the type of the event of interest is determined to be a defense event.
[0028] Furthermore, a number of wonderful moment videos are screened out from all the videos of interest based on the sound frequency, the pitch, and a preset relevance threshold, including:
[0029] Calculating the standard deviation of the sound frequency from the initial moment to each moment within a preset screening time period to obtain a number of frequency fluctuation values;
[0030] Calculating the average of all the frequency fluctuation values to obtain a frequency mean;
[0031] When the frequency mean is greater than a preset frequency mean threshold, calculating the standard deviation of the pitch at each moment from the initial moment to the preset screening time length to obtain a plurality of pitch fluctuation values;
[0032] A number of wonderful moment videos are screened out from all the videos of interest based on all the frequency fluctuation values, all the pitch fluctuation values, and a preset relevance threshold.
[0033] Furthermore, a number of wonderful moment videos are screened out from all the videos of interest based on all the frequency fluctuation values, all the pitch fluctuation values, and a preset relevance threshold, including:
[0034] Normalizing all the frequency fluctuation values within the preset screening time to obtain a frequency normalized data set;
[0035] Normalizing all the pitch fluctuation values within the preset screening time to obtain a pitch normalized data set;
[0036] Calculating the correlation coefficient between the frequency-normalized data set and the pitch-normalized data set to obtain a change correlation;
[0037] When the change correlation is greater than the preset correlation threshold, the video of interest is determined to be the wonderful moment video, so as to screen out a number of wonderful moment videos.
[0038] Further, adjusting the preset relevance threshold according to the number of all the videos of interest, the number of all the videos of wonderful moments, and the comprehensive determination factor within the preset adjustment time to obtain an adjusted relevance threshold, or adjusting the preset comprehensive determination threshold to obtain an adjusted comprehensive determination threshold, includes:
[0039] Calculate the proportion of the wonderful moment videos in all the videos of interest within the preset adjustment time to obtain the video accuracy rate;
[0040] When the video accuracy is greater than the maximum value of the preset accuracy range, increasing the preset relevance threshold according to the relative deviation between the video accuracy and the maximum value of the preset accuracy range and the preset relevance adjustment coefficient to obtain an adjusted relevance threshold;
[0041] When the video accuracy is less than a minimum value of a preset accuracy range, the preset comprehensive determination threshold is adjusted according to the comprehensive determination factors of all the wonderful moment videos to obtain an adjusted comprehensive determination threshold.
[0042] Furthermore, adjusting a preset comprehensive determination threshold according to the comprehensive determination factors of all the wonderful moment videos to obtain an adjusted comprehensive determination threshold includes:
[0043] Calculating an average value of the comprehensive determination factors of all the wonderful moment videos to obtain a comprehensive mean value;
[0044] Calculating the relative deviation between the comprehensive mean and the preset comprehensive determination threshold to obtain a comprehensive deviation;
[0045] The preset comprehensive determination threshold is increased according to the comprehensive deviation and the preset comprehensive factor adjustment coefficient to obtain an adjusted comprehensive determination threshold.
[0046] Further, determining whether there is a shooting intention based on the leg key points and the shooting distance includes:
[0047] When the shooting distance is less than a preset distance threshold, calculating the arc cosine function of the key point according to the coordinates of the key point of the leg to obtain the leg swing amplitude;
[0048] When the leg swing amplitude is greater than a preset swing threshold, it is determined that there is a shooting intention.
[0049] On the other hand, the present invention also provides a video acquisition system based on deep learning acceleration, comprising:
[0050] The data acquisition module is used to collect in real time the coordinates of several key points of the dribbling player's legs during a football match, the speed of the football, the shooting distance, the movement angle, the goalkeeper's distance from the goal line, the movement speed, the frequency of the audience's voices, and the commentator's pitch;
[0051] a first determination module, connected to the data acquisition module, for determining whether there is a shooting intention based on the key points of the leg and the shooting distance;
[0052] a second determination module, connected to the data acquisition module and the first determination module respectively, for determining a type of event of interest based on the shooting distance, the preset goal width, the movement speed, the movement angle, the moving speed, the goalkeeping distance, and a preset comprehensive determination threshold;
[0053] a video acquisition module, connected to the second determination module, for acquiring a plurality of videos of interest based on the deep learning acceleration model and the type of the event of interest;
[0054] a screening module, connected to the data acquisition module and the video acquisition module respectively, for screening out a number of wonderful moment videos from all the videos of interest based on the sound frequency, the pitch, and a preset relevance threshold;
[0055] an adjustment module, connected to the video acquisition module and the screening module respectively, for adjusting a preset relevance threshold according to the number of all the videos of interest, the number of all the videos of highlights within a preset adjustment time period, and a comprehensive determination factor to obtain an adjusted relevance threshold, or adjusting a preset comprehensive determination threshold to obtain an adjusted comprehensive determination threshold;
[0056] A storage module is connected to the screening module and stores the wonderful moment video re-determined based on the adjusted relevance threshold or the adjusted comprehensive determination threshold in a video database.
[0057] Compared with the existing technology, the beneficial effect of the present invention is that, by real-time collection and analysis of multi-dimensional key data, it can accurately determine the shooting intention and the type of event of interest, and screen out the wonderful moment video in combination with sound-related information; it can judge whether the player has the intention to shoot based on the key points of the legs and the shooting distance, and the comprehensive data analysis of football and goalkeepers can determine the type of event of interest. The sound frequency of the audience and the pitch of the commentator can assist in determining whether the video clip is wonderful. By adjusting the threshold, the screening process of the wonderful moment video can be dynamically optimized, so that the wonderful moment video finally stored in the video database is more in line with the actual degree of excitement, effectively solving the problems of low acquisition efficiency and low accuracy caused by over-reliance on complex models and single data.
[0058] Furthermore, through comprehensive consideration, the type of event of interest can be determined comprehensively and accurately, avoiding the one-sidedness brought about by single-factor judgment, because the shooting distance and goal width are directly related to the goal angle, and the movement speed, movement angle and goal angle jointly affect the goal matching degree. When the goal matching degree exceeds the preset matching degree threshold, the event of interest is comprehensively judged, taking into full consideration multiple factors such as the difficulty of the shot, the movement trend of the football, and the goalkeeper's defense. The type of event of interest is determined according to the goal matching degree and defensive ability factor, and the classification of different types of shooting events is further refined, which facilitates the subsequent targeted collection of videos of interest and screening of videos of wonderful moments, thereby improving the pertinence and effectiveness of video screening.
[0059] Furthermore, the defensive ability factor is calculated by comprehensively considering the moving speed and goalkeeping distance, and the comprehensive judgment factor is calculated by combining the goal matching degree and shooting distance. The comprehensive judgment factor is compared with the preset comprehensive judgment threshold to determine the focus event. This closely links the ability of the defender in the game with the shooting conditions of the attacker, and comprehensively evaluates the potential threat level and excitement of the current shooting scene from the perspectives of both the attacker and the defender. The comprehensive judgment factor comprehensively reflects the possibility of a successful shot and its viewing value. Comparing it with the preset comprehensive judgment threshold can effectively screen out game moments worthy of attention, ensuring that the collected video clips are more valuable and representative.
[0060] Furthermore, by normalizing the goal-matching and defensive ability factors separately, we can eliminate differences in their dimensions and magnitudes, making the different parameters comparable. Comparing the normalized match value with the normalized defense value to determine the type of event of interest effectively takes into account both the likelihood of a shot and the strength of the defense. When the normalized match value is greater than the normalized defense value, it indicates favorable conditions for a shot and a relatively high probability of a goal, thus determining it as a shot event. Conversely, when the normalized match value is less than or equal to the normalized defense value, it indicates that the defender's blocking ability is more advantageous in the current situation, making it difficult to create an effective shot threat, thus determining it as a defensive event.
[0061] Furthermore, by calculating the frequency fluctuation value and then the frequency mean, we can initially determine the emotional ups and downs of the live audience. This is because the fluctuation of sound frequency reflects the audience's reaction. When the frequency mean is greater than the preset frequency mean threshold, it indicates that the audience is quite excited and there may be a highlight moment. Next, the pitch fluctuation value is calculated. Pitch fluctuation can reflect the commentator's emotional changes. Commentators often have noticeable pitch changes during exciting moments. Finally, by combining the frequency fluctuation value, pitch fluctuation value, and the preset correlation threshold, we can filter the highlights video, reflecting the excitement of the game from different angles and more comprehensively capturing the climax of the game.
[0062] Furthermore, through normalization processing, a frequency-normalized data set and a pitch-normalized data set are obtained, and then the correlation coefficient of the two normalized data sets is calculated to obtain the change correlation. When the change correlation is greater than the preset correlation threshold, the corresponding focus video is determined to be a wonderful moment video. This is because the fluctuations in sound frequency and pitch are closely related to the excitement of the game. Frequency fluctuations reflect the audience's reaction, and pitch fluctuations reflect the commentator's emotional changes. The change correlation between the two can comprehensively reflect the intensity of the game and the on-site atmosphere.
[0063] Furthermore, by calculating the video accuracy, the effectiveness of video screening under the current threshold setting can be intuitively reflected. When the video accuracy is higher than or equal to the maximum value of the preset accuracy range, it indicates that the proportion of screened wonderful moment videos is too high, and there may be over-screening, and some general videos may be misjudged as wonderful moments. At this time, increasing the preset relevance threshold can improve the screening standard and reduce misjudgment; conversely, when the video accuracy is lower than the minimum value of the preset accuracy range, it means that a large number of wonderful moment videos have been missed. At this time, the preset comprehensive judgment threshold is adjusted according to the comprehensive judgment factor of the wonderful moment video.
[0064] Furthermore, by calculating the average value of the comprehensive judgment factors for all highlight moment videos, a comprehensive mean is calculated. This comprehensive mean reflects the overall level of the comprehensive judgment factors for the currently selected highlight moment videos. The relative deviation between the comprehensive mean and the preset comprehensive judgment threshold is then calculated to obtain a comprehensive deviation. This deviation reflects the difference between the threshold and the comprehensive judgment factors of the actual highlight videos. Increasing the preset comprehensive judgment threshold based on the comprehensive deviation and the preset comprehensive factor adjustment coefficient can dynamically improve the screening criteria, making subsequent video screening more rigorous and accurate.
[0065] Furthermore, by determining whether the shooting distance is less than a preset distance threshold, we can initially screen players in a range of possible shooting positions. This is because the success rate and likelihood of a shot only increase significantly when the player is a certain distance from the goal. Next, the leg swing amplitude is calculated using the coordinates of key leg points to quantify the player's leg swing. Leg swing amplitude is a key characteristic of shooting and is directly related to shooting intent. Finally, the leg swing amplitude is compared with a preset swing threshold. If it exceeds the threshold, it is determined that there is shooting intent.
[0066] Furthermore, by collecting real-time data such as the coordinates of key points on the player's legs and combining it with shot distance to determine shooting intent, the system provides a precise basis for subsequent event type determination. Next, the event type is determined by integrating multiple dimensions of data, including shot distance, goal width, and speed. These data collectively determine whether a shot scene has the potential to be exciting. Subsequently, relevant videos are collected based on the deep learning acceleration model and the event type of interest. The deep learning model can quickly identify and extract images related to these events, ensuring that the collected videos are targeted and timely. Highlight moment videos are then filtered based on the frequency of audience sounds and the pitch of commentators. These sound data can reflect the atmosphere and the level of emotion at the scene, and are closely related to exciting moments. Finally, by dynamically adjusting the preset relevance threshold and comprehensive judgment threshold, the system adaptively optimizes the selection criteria based on the number of videos of interest and highlights within the preset adjustment time period and the comprehensive judgment factor, ensuring that the video database contains both rich and accurate videos of exciting moments. This effectively addresses the issues of low acquisition efficiency and accuracy caused by over-reliance on complex models and single data sets. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 This is a flow chart of the video acquisition method based on deep learning acceleration in this embodiment;
[0068] Figure 2 This is a logic diagram for determining the occurrence of an event of interest in this embodiment;
[0069] Figure 3 A decision logic diagram for determining the event type of interest in this embodiment;
[0070] Figure 4 Schematic diagram of the video acquisition system based on deep learning acceleration in this embodiment. DETAILED DESCRIPTION
[0071] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.
[0072] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0073] See also Figure 1 As shown, it is a flow chart of the video acquisition method based on deep learning acceleration in this embodiment;
[0074] On the one hand, this embodiment provides a video acquisition method based on deep learning acceleration, including:
[0075] Real-time data collection of several key leg coordinates of the dribbling player, the speed of the ball, the shooting distance, the movement angle, the goalkeeper's distance from the goal line, the moving speed, the audience's sound frequency, and the commentator's voice pitch during a football match;
[0076] determining the presence of a shooting intention based on the leg key points and the shooting distance;
[0077] Determine the event type of interest based on the shooting distance, the preset goal width, the movement speed, the movement angle, the moving speed, the goalkeeping distance, and a preset comprehensive determination threshold;
[0078] Collect a number of videos of interest based on the deep learning acceleration model and the type of the event of interest;
[0079] Filtering a number of wonderful moment videos from all the videos of interest according to the sound frequency, the pitch, and a preset relevance threshold;
[0080] Adjusting the preset relevance threshold according to the number of all the videos of interest, the number of all the videos of wonderful moments within the preset adjustment time, and the comprehensive determination factor to obtain an adjusted relevance threshold, or adjusting the preset comprehensive determination threshold to obtain an adjusted comprehensive determination threshold;
[0081] The wonderful moment video re-determined based on the adjusted relevance threshold or the adjusted comprehensive determination threshold is stored in a video database.
[0082] The coordinates of several key points on the dribbling player's legs in this embodiment refer to the position coordinates of several key parts of the dribbling player's legs (such as the hip joint, knee joint, ankle joint, etc.). Their specific movement patterns are often associated with shooting movements and can be accurately extracted using computer vision technology by deploying multiple high-definition cameras around the playing field. The speed of the ball refers to the speed of the ball on the field. An accelerometer can be embedded in the ball to directly measure and transmit the ball's speed data. The shooting distance refers to the distance between the ball and the goal when the player is about to shoot. This can be directly measured using a laser rangefinder. The movement angle refers to the direction of the ball's movement and the goal line. The movement angle of the ball can be calculated by analyzing the ball's trajectory to determine its movement direction vector, simultaneously obtaining the direction vector of the goal line, and using the dot product formula between the two vectors to calculate the angle between the two vectors. The goalkeeper's goalkeeping distance refers to the vertical distance from the goalkeeper's position when guarding the goal to the goal line. High-precision cameras can be installed in the goal area to capture the goalkeeper's position in real time. The goalkeeping distance is calculated based on the field coordinate system. Movement speed refers to the speed at which the goalkeeper moves on the field. This can be measured in real time using a GPS device worn by the goalkeeper. Audience sound frequency reflects the pitch and variability of sounds such as shouts and cheers. Microphone arrays can be installed in different areas of the stadium to collect audience sound signals. Fourier transforms can then be used to convert the sound signals from the time domain to the frequency domain to obtain the sound frequency. The commentator's pitch refers to the high or low pitch of the commentator's voice. By placing a professional microphone in the commentator's work area, the pitch detection algorithm in audio analysis tools can be used to determine the pitch of the commentator's voice.
[0083] The preset comprehensive judgment threshold is a benchmark value used to determine the type of event of interest. It depends on the characteristics of the football game, the definition of a highlight moment, historical game data, and expert experience, and is usually set between 0.5 and 1. In this embodiment, it is set to 0.7, which ensures that key moments are not missed while avoiding excessive interference from irrelevant information.
[0084] Different video acquisition times are used for different types of events of interest. For shooting events, a relatively long acquisition time may be required to fully record the entire process from the player's kick to the ball entering the goal or being saved; for defensive events, the acquisition time can be slightly shorter, focusing on recording the key actions and interception moments of the defensive players; then, based on the deep learning acceleration model, high-speed image sensors and professional acquisition equipment are used to achieve high-frame, high-refresh rate video acquisition, and obtain several videos of interest.
[0085] The deep learning acceleration model selects different video collection times according to the type of event after the event of interest occurs, records the event of interest, and obtains the video of interest.
[0086] 1. Model
[0087] Convolutional Neural Networks (ResNet): Suitable for processing image and video data. Residual blocks are introduced to address the vanishing gradient problem during deep network training. This allows for the construction of very deep network structures, enabling better extraction of complex features. ResNet excels in tasks such as image classification and object detection. It can also be used for video analysis, taking video frames as input and extracting features from them to identify the event type of interest.
[0088] 2. Initial model parameters
[0089] 2.1 Weights: He initialization is typically used. For example, in a convolutional layer in a residual block, assuming a layer has 64 input channels and 128 output channels, the weights of this layer are initialized to follow a random normal distribution with a mean of 0 and a standard deviation of 64 + 1282. This ensures effective signal propagation in the neural network and avoids vanishing or exploding gradients.
[0090] 2.2 Bias: 0.
[0091] 3. Training process
[0092] 3.1 Data Preparation: We collected a large amount of football match video data and annotated it. For example, we labeled video clips containing shooting as "shooting events" and video clips containing defensive actions as "defensive events." We then preprocessed the video data, including cropping and scaling, to meet the model input size requirements. Video frames were uniformly cropped to 224×224 pixels.
[0093] 3.2 Build a training environment: Divide the dataset into a training set (used for model training, accounting for, for example, 70%) and a validation set (used to evaluate model performance, accounting for, for example, 30%), determine the loss function as the cross-entropy loss function, use the Adam optimizer as the optimizer, and set the initial learning rate to 0.001.
[0094] 3.3 Training Steps: The training set data is input into the ResNet model. The model first performs forward propagation to calculate the predicted probability for each event type, and then calculates the cross-entropy loss between the predicted results and the true labels. Through the backpropagation algorithm, the model weights and bias parameters are updated according to the gradient of the loss function. Each iteration adjusts the parameters in a direction that reduces the loss function. This process is repeated continuously. After multiple epochs (training cycles), the model gradually learns how to distinguish different types of events of interest.
[0095] 4. Model after training
[0096] The trained model achieved high performance metrics on the validation set, with accuracy exceeding 90%, accurately classifying input video frames into corresponding event types. The model's parameters have been optimized to their optimal state, with weights and bias values enabling the network to effectively extract and classify video features, thereby accurately identifying the types of events of interest in football match videos.
[0097] The preset relevance threshold is used to select highlight moments from the videos of interest. It is determined by the correlation between the sound frequency and pitch and the highlight moments, and is usually set between 0.5 and 1. In this embodiment, it is set to 0.6, which ensures that most highlight moments are selected while avoiding over-selection due to accidental sound factors.
[0098] By collecting multi-dimensional key data during football games in real time, the presence of a shooting intention is determined based on the collected leg key points and shooting distance; the type of event of interest is determined based on the shooting distance, preset goal width, movement speed, movement angle, moving speed, goalkeeping distance and a preset comprehensive judgment threshold; a number of videos of interest are collected using a deep learning acceleration model and the determined type of event of interest; a number of wonderful moment videos are screened out from all videos of interest based on sound frequency, pitch and a preset relevance threshold; the preset relevance threshold or the preset comprehensive judgment threshold is adjusted based on the number of all videos of interest within a preset adjustment time, the number of all wonderful moment videos and a comprehensive judgment factor; the wonderful moment videos re-determined based on the adjusted relevance threshold or comprehensive judgment threshold are stored in a video database.
[0099] By collecting and analyzing multi-dimensional key data in real time, it is possible to accurately determine shooting intentions and types of events of interest, and filter out videos of exciting moments based on sound-related information; it is possible to determine whether a player has intentions to shoot based on key points on the legs and the shooting distance, and to determine the types of events of interest based on comprehensive football and goalkeeper data analysis. The frequency of audience voices and the pitch of commentators can assist in determining whether a video clip is exciting. By adjusting the threshold, the screening process of exciting moment videos can be dynamically optimized, so that the exciting moment videos finally stored in the video database are more in line with the actual level of excitement, effectively solving the problems of low acquisition efficiency and low accuracy caused by over-reliance on complex models and single data.
[0100] Specifically, the type of event of interest is determined based on the shooting distance, the preset goal width, the movement speed, the movement angle, the moving speed, the goalkeeping distance, and a preset comprehensive determination threshold, including:
[0101] Calculating a goal angle based on the shooting distance and a preset goal width;
[0102] Calculating a goal matching degree according to the movement speed, the movement angle, and the goal angle;
[0103] When the goal matching degree is greater than a preset matching degree threshold, determining that a concern event occurs based on the shooting distance, the goal matching degree, the moving speed, the goalkeeping distance, and a preset comprehensive determination threshold;
[0104] The type of event of interest is determined according to the goal matching degree and the defensive capability factor.
[0105] The preset goal width refers to the actual width of a soccer goal, which has strict uniform rules. In this embodiment, the standard goal size of 11-a-side soccer is 7.32 meters.
[0106] The preset match threshold is used to determine whether the goal match meets the criteria for determining a notable event. It depends on the intensity of the game and the skill level of the players and is usually set between 0.6 and 0.9. In this embodiment, it is set to 0.7, which ensures that shot events with a high probability of scoring are screened out without being too strict.
[0107] The goal angle is calculated using the shot distance and the preset goal width. Then, the goal fit is calculated based on the movement speed, movement angle, and goal angle. If the goal fit exceeds the preset fit threshold, a determination is made as to whether an event of interest has occurred based on the shot distance, goal fit, movement speed, goalkeeping distance, and a preset comprehensive determination threshold. Once an event of interest has been determined, the specific type of event is determined based on the goal fit and defensive ability factor.
[0108] Through comprehensive consideration, the type of event of interest can be determined comprehensively and accurately, avoiding the one-sidedness brought about by single-factor judgment, because the shooting distance and goal width are directly related to the goal angle, and the movement speed, movement angle and goal angle jointly affect the goal matching degree. When the goal matching degree exceeds the preset matching degree threshold, the event of interest is comprehensively judged, taking into full consideration multiple factors such as the difficulty of the shot, the movement trend of the football, and the goalkeeper's defense. The type of event of interest is determined based on the goal matching degree and defensive ability factor, and the classification of different types of shooting events is further refined, which facilitates the subsequent targeted collection of videos of interest and screening of videos of wonderful moments, thereby improving the pertinence and effectiveness of video screening.
[0109] See also Figure 2 As shown, it is a decision logic diagram for determining the occurrence of a concern event in this embodiment;
[0110] Determining the occurrence of a concern event according to the shooting distance, the goal matching degree, the moving speed, the goalkeeping distance, and a preset comprehensive determination threshold includes:
[0111] Calculate a defensive capability factor based on the moving speed and the goalkeeping distance;
[0112] A comprehensive determination factor is calculated based on the goal matching degree, the defensive ability factor, and the shooting distance;
[0113] When the comprehensive determination factor is greater than the preset comprehensive determination threshold, it is determined that an event of interest occurs.
[0114] By calculating the defensive ability factor based on the moving speed and the goalkeeping distance, and then combining the goal matching degree, the defensive ability factor and the shooting distance, a comprehensive determination factor is further calculated. This comprehensive determination factor is then compared with the preset comprehensive determination threshold. If the comprehensive determination factor exceeds this preset threshold, it is determined that an event of concern has occurred.
[0115] The defensive ability factor is calculated by comprehensively considering the moving speed and goalkeeping distance, and the comprehensive judgment factor is calculated based on the goal matching degree and shooting distance. The factor is compared with the preset comprehensive judgment threshold to determine the events of interest. This closely links the ability of the defender in the game with the shooting conditions of the attacker, and comprehensively evaluates the potential threat level and excitement of the current shooting scene from the perspectives of both the attacker and the defender. The comprehensive judgment factor comprehensively reflects the possibility of a successful shot and its viewing value. Comparing it with the preset comprehensive judgment threshold can effectively screen out game moments worthy of attention, ensuring that the collected video clips are more valuable and representative.
[0116] Please continue reading Figure 3 As shown, it is a decision logic diagram for determining the type of event of interest in this embodiment;
[0117] The event type of interest is determined based on the goal matching degree and defensive capability factor, including:
[0118] Normalizing the goal matching degree to obtain a matching degree normalization value;
[0119] Normalizing the defense capability factor to obtain a defense normalized value;
[0120] When the matching degree normalized value is greater than the defense normalized value, determining that the type of the event of interest is a shooting event;
[0121] When the matching degree normalized value is less than or equal to the defense normalized value, the type of the event of interest is determined to be a defense event.
[0122] By normalizing the goal match degree and the defensive ability factor respectively, we obtain the match degree normalization value and the defense normalization value, that is, scaling the goal match degree and the defensive ability factor to between 0 and 1, eliminating the dimension and order of magnitude differences between different data. This is a prior art and will not be described in detail here. Then, by comparing these two normalized values, we determine the type of event of interest. Specifically, if the match degree normalization value is greater than the defense normalization value, the event of interest is determined to be a shooting event; conversely, if the match degree normalization value is less than or equal to the defense normalization value, the event of interest is determined to be a defensive event.
[0123] By normalizing the goal-matching and defensive ability factors separately, we can eliminate differences in their dimensions and magnitudes, making the different parameters comparable. Comparing the normalized match value with the normalized defense value to determine the type of event of interest effectively considers both the likelihood of a shot and the strength of the defense. When the normalized match value is greater than the normalized defense value, it indicates favorable conditions for a shot and a relatively high probability of a goal, thus determining it as a shot event. Conversely, when the normalized match value is less than or equal to the normalized defense value, it indicates that the defender's blocking ability is superior in the current situation, making it difficult to create an effective shot threat, thus determining it as a defensive event.
[0124] Specifically, a number of wonderful moment videos are screened out from all the videos of interest based on the sound frequency, the pitch, and a preset relevance threshold, including:
[0125] Calculating the standard deviation of the sound frequency from the initial moment to each moment within a preset screening time period to obtain a number of frequency fluctuation values;
[0126] Calculating the average of all the frequency fluctuation values to obtain a frequency mean;
[0127] When the frequency mean is greater than a preset frequency mean threshold, calculating the standard deviation of the pitch at each moment from the initial moment to the preset screening time length to obtain a plurality of pitch fluctuation values;
[0128] A number of wonderful moment videos are screened out from all the videos of interest based on all the frequency fluctuation values, all the pitch fluctuation values, and a preset relevance threshold.
[0129] The preset screening time is the length of time used to calculate the sound frequency fluctuations and pitch fluctuations. It depends on the average duration of the highlights in the football match and the system's real-time requirements. It is usually set between 5 and 10 seconds. In this embodiment, it is set to 7 seconds, which can cover the key moments of the highlights while avoiding too long a time.
[0130] The preset frequency mean threshold is a critical value used to reflect the excitement of the scene. It is determined by the frequency variation of sound when the audience is excited and the typical intensity of football matches. It is usually set between 50Hz and 200Hz. In this embodiment, it is set to 120Hz, which can effectively distinguish between ordinary moments and those with high audience excitement, thereby more accurately screening videos of potentially exciting moments.
[0131] By calculating the standard deviation of the sound frequency at each moment from the initial moment to the preset screening duration, multiple frequency fluctuation values are obtained. Then, the average of all frequency fluctuation values is calculated to obtain the frequency mean. If the frequency mean exceeds the preset frequency mean threshold, the system further calculates the standard deviation of the pitch at each moment from the initial moment to the preset screening duration to obtain multiple pitch fluctuation values. Finally, based on all frequency fluctuation values, pitch fluctuation values, and the preset relevance threshold, a number of videos with highlights are selected from all videos of interest.
[0132] By calculating the frequency fluctuation value and then the frequency mean, we can initially determine the emotional ups and downs of the live audience. This is because the fluctuation of sound frequency reflects the audience's reaction. When the frequency mean exceeds the preset frequency mean threshold, it indicates that the audience is quite excited and there may be a highlight moment. Next, we calculate the pitch fluctuation value. Pitch fluctuation can reflect the commentator's emotional changes. Commentators often have noticeable pitch changes during exciting moments. Finally, by combining the frequency fluctuation value, pitch fluctuation value, and the preset correlation threshold, we can filter the highlights video, reflecting the excitement of the game from different angles and more comprehensively capturing the climax of the game.
[0133] Specifically, a number of wonderful moment videos are screened out from all the videos of interest based on all the frequency fluctuation values, all the pitch fluctuation values, and a preset relevance threshold, including:
[0134] Normalizing all the frequency fluctuation values within the preset screening time to obtain a frequency normalized data set;
[0135] Normalizing all the pitch fluctuation values within the preset screening time to obtain a pitch normalized data set;
[0136] Calculating the correlation coefficient between the frequency-normalized data set and the pitch-normalized data set to obtain a change correlation;
[0137] When the change correlation is greater than the preset correlation threshold, the video of interest is determined to be the wonderful moment video, so as to screen out a number of wonderful moment videos.
[0138] That is, the goal matching degree and the defensive ability factor are scaled to between 0 and 1 respectively, eliminating the dimension and order of magnitude differences between different data. This is an existing technology and will not be elaborated on here.
[0139] By normalizing all frequency fluctuation values and all pitch fluctuation values within the preset screening time, a frequency normalized data set and a pitch normalized data set are obtained, that is, the frequency fluctuation values and the pitch fluctuation values are scaled to between 0 and 1, respectively, to eliminate the dimension and order of magnitude differences between different data. This is a prior art and will not be elaborated on here. Next, the correlation coefficient of the frequency normalized data set and the pitch normalized data set is calculated to obtain the change correlation. If the change correlation is greater than the preset correlation threshold, the corresponding video of interest is determined to be a wonderful moment video, thereby screening out several wonderful moment videos.
[0140] Through normalization processing, we obtain a frequency-normalized data set and a pitch-normalized data set. Then, we calculate the correlation coefficient of these two normalized data sets to obtain the change correlation. When the change correlation is greater than the preset correlation threshold, the corresponding video of interest is determined to be a wonderful moment video. This is because the fluctuations in sound frequency and pitch are closely related to the excitement of the game. Frequency fluctuations reflect the audience's reaction, and pitch fluctuations reflect the commentator's emotional changes. The change correlation between the two can comprehensively reflect the intensity of the game and the on-site atmosphere.
[0141] Specifically, adjusting the preset relevance threshold according to the number of all the videos of interest, the number of all the wonderful moment videos, and the comprehensive determination factor within the preset adjustment time to obtain the adjusted relevance threshold, or adjusting the preset comprehensive determination threshold to obtain the adjusted comprehensive determination threshold includes:
[0142] Calculate the proportion of the wonderful moment videos in all the videos of interest within the preset adjustment time to obtain the video accuracy rate;
[0143] When the video accuracy is greater than the maximum value of the preset accuracy range, increasing the preset relevance threshold according to the relative deviation between the video accuracy and the maximum value of the preset accuracy range and the preset relevance adjustment coefficient to obtain an adjusted relevance threshold;
[0144] When the video accuracy is less than a minimum value of a preset accuracy range, the preset comprehensive determination threshold is adjusted according to the comprehensive determination factors of all the wonderful moment videos to obtain an adjusted comprehensive determination threshold.
[0145] The preset adjustment period is the time window used to evaluate and adjust the relevance threshold and the comprehensive determination threshold. It depends on the video acquisition system's requirements for real-time event processing and the need for data statistical stability, and is typically set between 1 and 3 hours. In this embodiment, it is set to 2 hours to ensure that the data sample size is representative and to avoid excessive data fluctuations caused by a shorter period.
[0146] The preset accuracy range is used to measure whether the proportion of highlight videos in the focus video is reasonable. It depends on the requirements for video acquisition quality and the definition of highlight moments in the game, and is usually set between 50% and 90%. In this embodiment, it is set to [60%, 80%] to ensure that the selected highlight videos are of high quality while avoiding overly strict screening that may miss some potential highlights.
[0147] The preset relevance adjustment coefficient determines the magnitude of the threshold adjustment. It is determined by the sensitivity to changes in video accuracy and the stability requirements for the screening results, and is typically set between 0.1 and 0.5. In this embodiment, it is set to 0.3 to ensure that the system is sufficiently responsive to changes in video accuracy while avoiding instability caused by excessively large threshold adjustments due to an excessively large coefficient.
[0148] The video accuracy rate is calculated by calculating the proportion of highlight moment videos among all watched videos. Then, if the video accuracy rate is greater than or equal to the maximum value of a preset accuracy rate range, the preset relevance threshold is increased based on the relative deviation between the video accuracy rate and the maximum value and the preset relevance adjustment coefficient, thereby obtaining an adjusted relevance threshold. If the video accuracy rate is less than the minimum value of the preset accuracy rate range, the preset comprehensive judgment threshold is adjusted based on the comprehensive judgment factor of all highlight moment videos, thereby obtaining an adjusted comprehensive judgment threshold.
[0149] By calculating the video accuracy, the effectiveness of video screening under the current threshold setting can be intuitively reflected. When the video accuracy is higher than or equal to the maximum value of the preset accuracy range, it indicates that the proportion of screened wonderful moment videos is too high, and there may be over-screening, and some general videos may be misjudged as wonderful moments. At this time, increasing the preset relevance threshold can improve the screening standard and reduce misjudgment; conversely, when the video accuracy is lower than the minimum value of the preset accuracy range, it means that a large number of wonderful moment videos have been missed. At this time, the preset comprehensive judgment threshold is adjusted according to the comprehensive judgment factor of the wonderful moment video.
[0150] Specifically, adjusting the preset comprehensive determination threshold according to the comprehensive determination factors of all the wonderful moment videos to obtain the adjusted comprehensive determination threshold includes:
[0151] Calculating an average value of the comprehensive determination factors of all the wonderful moment videos to obtain a comprehensive mean value;
[0152] Calculating the relative deviation between the comprehensive mean and the preset comprehensive determination threshold to obtain a comprehensive deviation;
[0153] The preset comprehensive determination threshold is increased according to the comprehensive deviation and the preset comprehensive factor adjustment coefficient to obtain an adjusted comprehensive determination threshold.
[0154] The preset comprehensive factor adjustment coefficient is used to adjust the preset comprehensive judgment threshold based on the comprehensive deviation. It depends on the fluctuation range of the comprehensive judgment factor, the rate of change of the game data, and the required sensitivity of the threshold adjustment. It is usually set between 0.1 and 0.5. In this embodiment, it is set to 0.3, which can adapt to data changes in a timely manner without causing threshold instability due to excessive adjustment.
[0155] The comprehensive mean is calculated by averaging the comprehensive judgment factors of all highlight moment videos. Next, the relative deviation between the comprehensive mean and a preset comprehensive judgment threshold is calculated to obtain a comprehensive deviation. Then, based on the comprehensive deviation and the preset comprehensive factor adjustment coefficient, the preset comprehensive judgment threshold is increased to obtain an adjusted comprehensive judgment threshold.
[0156] The comprehensive mean is calculated by averaging the comprehensive factors of all highlight videos. This means it reflects the overall performance of the selected highlight videos in terms of these factors. The relative deviation between the comprehensive mean and the preset comprehensive threshold is then calculated to obtain the comprehensive deviation. This deviation reflects the difference between the threshold and the actual comprehensive factors of the highlight videos. Increasing the preset comprehensive threshold based on the comprehensive deviation and the preset comprehensive factor adjustment coefficient dynamically improves the screening criteria, making subsequent video screening more rigorous and accurate.
[0157] Specifically, determining whether there is a shooting intention based on the key points of the leg and the shooting distance includes:
[0158] When the shooting distance is less than a preset distance threshold, calculating the arc cosine function of the key point according to the coordinates of the key point of the leg to obtain the leg swing amplitude;
[0159] When the leg swing amplitude is greater than a preset swing threshold, it is determined that there is a shooting intention.
[0160] The preset distance threshold is used to determine whether a shot is close enough to indicate a potential shot attempt. It depends on the player's shooting range and the goalkeeper's defensive coverage and is typically set between 5 and 15 meters. In this embodiment, it is set to 10 meters to cover most threatening shooting areas while avoiding misidentifying shots that are too far away and have a low probability of being shot as a potential shot attempt.
[0161] The preset leg swing threshold is a standard value used to determine whether the leg swing amplitude is large enough to confirm the presence of a shooting intention. It is determined by the typical leg swing amplitude during a shooting action and individual differences between players and is typically set between 0.5 and 1.5 radians. In this embodiment, it is set to 0.8 radians, which can effectively distinguish between ordinary running or passing actions and leg swings with shooting intentions, thereby improving the accuracy of the judgment.
[0162] When the shooting distance is less than a preset distance threshold, the inverse cosine function of the key points on the legs is calculated based on their coordinates to determine the leg swing amplitude. The calculated leg swing amplitude is then compared with a preset swing threshold. If the leg swing amplitude exceeds the preset swing threshold, a shooting intention is determined.
[0163] By determining whether the shooting distance is less than a preset distance threshold, we can initially identify players in a range of possible shooting positions. This is because the success rate and likelihood of a shot significantly increase only when the player is a certain distance from the goal. Next, the leg swing amplitude is calculated using the coordinates of key leg points to quantify the player's leg swing. Leg swing amplitude is a key characteristic of shooting and is directly related to shooting intent. Finally, the leg swing amplitude is compared with a preset swing threshold. If it exceeds the threshold, it is determined to be a shooting intent.
[0164] Please continue reading Figure 4 As shown, it is a schematic diagram of the video acquisition system based on deep learning acceleration in this embodiment;
[0165] On the other hand, this embodiment provides a video acquisition system based on deep learning acceleration, including:
[0166] The data acquisition module is used to collect in real time the coordinates of several key points of the dribbling player's legs during a football match, the speed of the football, the shooting distance, the movement angle, the goalkeeper's distance from the goal line, the movement speed, the frequency of the audience's voices, and the commentator's pitch;
[0167] a first determination module, connected to the data acquisition module, for determining whether there is a shooting intention based on the key points of the leg and the shooting distance;
[0168] a second determination module, connected to the data acquisition module and the first determination module respectively, for determining a type of event of interest based on the shooting distance, the preset goal width, the movement speed, the movement angle, the moving speed, the goalkeeping distance, and a preset comprehensive determination threshold;
[0169] a video acquisition module, connected to the second determination module, for acquiring a plurality of videos of interest based on the deep learning acceleration model and the type of the event of interest;
[0170] a screening module, connected to the data acquisition module and the video acquisition module respectively, for screening out a number of wonderful moment videos from all the videos of interest based on the sound frequency, the pitch, and a preset relevance threshold;
[0171] an adjustment module, connected to the video acquisition module and the screening module respectively, for adjusting a preset relevance threshold according to the number of all the videos of interest, the number of all the videos of highlights within a preset adjustment time period, and a comprehensive determination factor to obtain an adjusted relevance threshold, or adjusting a preset comprehensive determination threshold to obtain an adjusted comprehensive determination threshold;
[0172] A storage module is connected to the screening module and stores the wonderful moment video re-determined based on the adjusted relevance threshold or the adjusted comprehensive determination threshold in a video database.
[0173] The system collects real-time coordinates of several key leg points of the dribbler, the ball's speed, shot distance, angle of movement, goalkeeper's distance from the goal line, movement speed, audience voice frequency, and commentator's voice pitch. Shooting intent is determined based on the leg key points and shot distance. The event type of interest is determined based on the shot distance, preset goal width, speed, angle of movement, movement speed, goalkeeper distance, and a preset comprehensive judgment threshold. Based on the deep learning acceleration model and the event type of interest, several videos of interest are collected. Highlight moment videos are filtered from the videos of interest using audience voice frequency, commentator's voice pitch, and a preset relevance threshold. Based on the number of videos of interest within a preset adjustment time period, the number of highlight moment videos, and a comprehensive judgment factor, the preset relevance threshold or preset comprehensive judgment threshold is adjusted to obtain a new adjusted relevance threshold or adjusted comprehensive judgment threshold. Finally, the highlight moment videos re-determined based on the adjusted relevance threshold or comprehensive judgment threshold are stored in the video database.
[0174] By collecting real-time data such as the coordinates of key points on a player's legs and combining them with shot distance to determine shooting intent, the system provides a precise basis for subsequent event type determination. Next, the system determines the event type by integrating multiple dimensions, including shot distance, goal width, and speed. These data collectively determine whether a shot scene has the potential to be exciting. Subsequently, relevant videos are collected based on the deep learning acceleration model and the event type. The deep learning model can quickly identify and extract footage related to these events, ensuring that the collected videos are targeted and timely. Highlight moment videos are then filtered based on the frequency of audience sounds and the pitch of commentators' voices. These sound data can reflect the atmosphere and the level of emotion at the scene, and are closely related to exciting moments. Finally, the system dynamically adjusts the preset relevance threshold and comprehensive judgment threshold. Based on the number of videos of interest and highlights within the preset adjustment time period and the comprehensive judgment factor, the system adaptively optimizes the selection criteria, ensuring that the video database contains both rich and accurate videos of exciting moments. This effectively addresses the issues of low acquisition efficiency and accuracy caused by over-reliance on complex models and single-source data.
[0175] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A video acquisition method based on deep learning acceleration, characterized in that: include: Real-time data collection of several key leg coordinates of the dribbling player, the speed of the ball, the shooting distance, the movement angle, the goalkeeper's distance from the goal line, the moving speed, the audience's sound frequency, and the commentator's voice pitch during a football match; determining the presence of a shooting intention based on the leg key points and the shooting distance; Determine the event type of interest based on the shooting distance, the preset goal width, the movement speed, the movement angle, the moving speed, the goalkeeping distance, and a preset comprehensive determination threshold; Collect a number of videos of interest based on the deep learning acceleration model and the type of the event of interest; Filtering a number of wonderful moment videos from all the videos of interest according to the sound frequency, the pitch, and a preset relevance threshold; Adjusting the preset relevance threshold according to the number of all the videos of interest, the number of all the videos of wonderful moments within the preset adjustment time, and the comprehensive determination factor to obtain an adjusted relevance threshold, or adjusting the preset comprehensive determination threshold to obtain an adjusted comprehensive determination threshold; The wonderful moment video re-determined based on the adjusted relevance threshold or the adjusted comprehensive determination threshold is stored in a video database.
2. The video acquisition method based on deep learning acceleration according to claim 1, characterized in that Determining the event type of interest according to the shooting distance, the preset goal width, the movement speed, the movement angle, the moving speed, the goalkeeping distance, and a preset comprehensive determination threshold includes: Calculating a goal angle based on the shooting distance and a preset goal width; Calculating a goal matching degree according to the movement speed, the movement angle, and the goal angle; When the goal matching degree is greater than a preset matching degree threshold, determining that a concern event occurs based on the shooting distance, the goal matching degree, the moving speed, the goalkeeping distance, and a preset comprehensive determination threshold; The type of event of interest is determined according to the goal matching degree and the defensive capability factor.
3. The video acquisition method based on deep learning acceleration according to claim 2, characterized in that Determining the occurrence of a concern event according to the shooting distance, the goal matching degree, the moving speed, the goalkeeping distance, and a preset comprehensive determination threshold includes: Calculate a defensive capability factor based on the moving speed and the goalkeeping distance; A comprehensive determination factor is calculated based on the goal matching degree, the defensive ability factor, and the shooting distance; When the comprehensive determination factor is greater than the preset comprehensive determination threshold, it is determined that an event of interest occurs.
4. The video acquisition method based on deep learning acceleration according to claim 3, characterized in that The event type of interest is determined based on the goal matching degree and defensive capability factor, including: Normalizing the goal matching degree to obtain a matching degree normalization value; Normalizing the defense capability factor to obtain a defense normalized value; When the matching degree normalized value is greater than the defense normalized value, determining that the type of the event of interest is a shooting event; When the matching degree normalized value is less than or equal to the defense normalized value, the type of the event of interest is determined to be a defense event.
5. The video acquisition method based on deep learning acceleration according to claim 4, characterized in that: A number of highlight moment videos are screened out from all the videos of interest based on the sound frequency, the pitch, and a preset relevance threshold, including: Calculating the standard deviation of the sound frequency from the initial moment to each moment within a preset screening time period to obtain a number of frequency fluctuation values; Calculating the average of all the frequency fluctuation values to obtain a frequency mean; When the frequency mean is greater than a preset frequency mean threshold, calculating the standard deviation of the pitch at each moment from the initial moment to the preset screening time length to obtain a plurality of pitch fluctuation values; A number of wonderful moment videos are screened out from all the videos of interest based on all the frequency fluctuation values, all the pitch fluctuation values, and a preset relevance threshold.
6. The video acquisition method based on deep learning acceleration according to claim 5, characterized in that: A number of wonderful moment videos are screened out from all the videos of interest based on all the frequency fluctuation values, all the pitch fluctuation values, and a preset relevance threshold, including: Normalizing all the frequency fluctuation values within the preset screening time to obtain a frequency normalized data set; Normalizing all the pitch fluctuation values within the preset screening time to obtain a pitch normalized data set; Calculating the correlation coefficient between the frequency-normalized data set and the pitch-normalized data set to obtain a change correlation; When the change correlation is greater than the preset correlation threshold, the video of interest is determined to be the wonderful moment video, so as to screen out a number of wonderful moment videos.
7. The video acquisition method based on deep learning acceleration according to claim 6, characterized in that: Adjusting the preset relevance threshold according to the number of all the videos of interest, the number of all the videos of wonderful moments within the preset adjustment time, and the comprehensive determination factor to obtain an adjusted relevance threshold, or adjusting the preset comprehensive determination threshold to obtain an adjusted comprehensive determination threshold, including: Calculate the proportion of the wonderful moment videos in all the videos of interest within the preset adjustment time to obtain the video accuracy rate; When the video accuracy is greater than the maximum value of the preset accuracy range, increasing the preset relevance threshold according to the relative deviation between the video accuracy and the maximum value of the preset accuracy range and the preset relevance adjustment coefficient to obtain an adjusted relevance threshold; When the video accuracy is less than a minimum value of a preset accuracy range, the preset comprehensive determination threshold is adjusted according to the comprehensive determination factors of all the wonderful moment videos to obtain an adjusted comprehensive determination threshold.
8. The video acquisition method based on deep learning acceleration according to claim 7, characterized in that: Adjusting a preset comprehensive determination threshold according to the comprehensive determination factors of all the wonderful moment videos to obtain an adjusted comprehensive determination threshold includes: Calculating an average value of the comprehensive determination factors of all the wonderful moment videos to obtain a comprehensive mean value; Calculating the relative deviation between the comprehensive mean and the preset comprehensive determination threshold to obtain a comprehensive deviation; The preset comprehensive determination threshold is increased according to the comprehensive deviation and the preset comprehensive factor adjustment coefficient to obtain an adjusted comprehensive determination threshold.
9. The video acquisition method based on deep learning acceleration according to claim 8, characterized in that: Determining whether there is a shooting intention based on the leg key points and the shooting distance includes: When the shooting distance is less than a preset distance threshold, calculating the arc cosine function of the key point according to the coordinates of the key point of the leg to obtain the leg swing amplitude; When the leg swing amplitude is greater than a preset swing threshold, it is determined that there is a shooting intention.
10. A video acquisition system based on deep learning acceleration, based on the video acquisition method based on deep learning acceleration according to any one of claims 1 to 9, characterized in that: include: The data acquisition module is used to collect in real time the coordinates of several key points of the dribbling player's legs during a football match, the speed of the football, the shooting distance, the movement angle, the goalkeeper's distance from the goal line, the movement speed, the frequency of the audience's voices, and the commentator's pitch; a first determination module, connected to the data acquisition module, for determining whether there is a shooting intention based on the key points of the leg and the shooting distance; a second determination module, connected to the data acquisition module and the first determination module respectively, for determining a type of event of interest based on the shooting distance, the preset goal width, the movement speed, the movement angle, the moving speed, the goalkeeping distance, and a preset comprehensive determination threshold; a video acquisition module, connected to the second determination module, for acquiring a plurality of videos of interest based on the deep learning acceleration model and the type of the event of interest; a screening module, connected to the data acquisition module and the video acquisition module respectively, for screening out a number of wonderful moment videos from all the videos of interest based on the sound frequency, the pitch, and a preset relevance threshold; an adjustment module, connected to the video acquisition module and the screening module respectively, for adjusting a preset relevance threshold according to the number of all the videos of interest, the number of all the videos of highlights within a preset adjustment time period, and a comprehensive determination factor to obtain an adjusted relevance threshold, or adjusting a preset comprehensive determination threshold to obtain an adjusted comprehensive determination threshold; A storage module is connected to the screening module and stores the wonderful moment video re-determined based on the adjusted relevance threshold or the adjusted comprehensive determination threshold in a video database.
Citation Information
Patent Citations
Football match behavior recognition method and device based on deep learning and terminal equipment
CN110378245A
Method based on audio / video combination for detecting highlight events in football video
CN101650722A
Game spectacular moment recognition method, terminal and computer readable storage medium
CN110339566A
Method for extracting wonderful clip of badminton event video based on machine learning
CN111291617A
Video analysis method and device, electronic equipment and storage medium
CN115115985A