A video acquisition method and system based on deep learning acceleration
By collecting multi-dimensional data in real time and accelerating the model with deep learning, combined with dynamic threshold adjustment, the system accurately determines the shooting intention and the events of interest, and filters out the highlights videos. This solves the problems of low collection efficiency and insufficient accuracy in existing technologies, and achieves efficient and accurate video collection.
Patent Information
- Application Number
- CN202510814499.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-06-18
AI Technical Summary
Existing technologies for video capture in football matches suffer from low efficiency and insufficient accuracy, failing to meet the requirements for real-time performance and comprehensiveness. Their reliance on complex models and single data sources results in poor capture outcomes.
By collecting multi-dimensional key data in real time, such as the coordinates of key points on the dribbling player's legs, the speed of the football, the distance to the goalkeeper, and the frequency of audience voices, combined with a deep learning acceleration model and a dynamic threshold adjustment mechanism, the system determines the shooting intention and the type of event of interest, and selects exciting video moments.
It enables accurate determination of shooting intent and the type of event of interest, and filters out video clips that match the actual level of excitement, improving the efficiency and accuracy of data collection, avoiding the one-sidedness of judgment based on a single factor, and ensuring the value and representativeness of the exciting moments videos stored in the video database.
Smart Images

Figure CN120708129B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and in particular to a video acquisition method and system based on deep learning acceleration. BACKGROUND
[0002] In football events, accurate collection of highlight clips is of great significance for improving audience experience, optimizing content production and assisting tactical analysis. With the continuous expansion of the global influence of football events, the demand for high-quality, real-time highlight clips among audiences is growing. Traditional methods rely on manual or simple algorithms, which have low efficiency and insufficient accuracy, and are difficult to meet the needs of modern events. The development of deep learning technology provides a new way to solve this problem.
[0003] The patent document with publication number CN110378245A discloses a football match behavior recognition method, device and terminal equipment based on deep learning, which includes: obtaining a football match video to be recognized; dividing the football match video into N video segments, and extracting a frame of image from each video segment as an input image, N is an integer greater than 1; using a preset deep learning network model to process the input image to obtain a behavior recognition result corresponding to the football match video, wherein the deep learning network model is composed of an Inception network model and a three-dimensional ResNet network model in cascade, the Inception network model is used to learn the relationship between pixel points in each frame of the input image, and the three-dimensional ResNet network model is used to learn the relationship between frames of the input image.
[0004] As can be seen, the football match behavior recognition method based on deep learning has the following problems: the Inception network model and the three-dimensional ResNet network model are cascaded to form a deep learning network model, which requires a large amount of computing resources and time; when processing video and recognizing behavior, it cannot meet the real-time requirement; insufficient training data or low data quality will affect the recognition accuracy of the model; relying on single video picture analysis, the analysis is not comprehensive and the accuracy is insufficient. SUMMARY
[0005] Therefore, the present application provides a video acquisition method and system based on deep learning acceleration, which overcomes the problems of low acquisition efficiency and low accuracy caused by excessive reliance on complex models and single data in the prior art through multi-dimensional data analysis, deep learning acceleration model and dynamic threshold adjustment mechanism.
[0006] To achieve the above-mentioned purpose, on the one hand, the present application provides a video acquisition method based on deep learning acceleration, which includes:
[0007] Real-time collection of several leg key point coordinates of a ball player, movement speed of a football, shooting distance, movement angle, goalkeeper distance to the goal line, movement speed, sound frequency of the audience, and tone of the commentator during a football match;
[0008] Determination of shooting intention according to the leg key points and the shooting distance;
[0009] Determination of a focus event type according to the shooting distance, a preset goal width, the movement speed, the movement angle, the movement speed, the goalkeeper distance, and a preset comprehensive determination threshold;
[0010] Collection of several focus videos according to a deep learning acceleration model and the focus event type;
[0011] Screening of several highlight videos from all the focus videos according to the sound frequency, the tone, and a preset correlation threshold;
[0012] Adjustment of a preset correlation threshold according to the number of all the focus videos, the number of all the highlight videos, and a comprehensive determination factor within a preset adjustment duration, to obtain an adjusted correlation threshold, or adjustment of a preset comprehensive determination threshold to obtain an adjusted comprehensive determination threshold;
[0013] Storage of the highlight videos re-determined based on the adjusted correlation threshold or the adjusted comprehensive determination threshold into a video database.
[0014] Further, determination of a focus event type according to the shooting distance, a preset goal width, the movement speed, the movement angle, the movement speed, the goalkeeper distance, and a preset comprehensive determination threshold, comprises:
[0015] Calculation of a goal angle according to the shooting distance and a preset goal width;
[0016] Calculation of a goal matching degree according to the movement speed, the movement angle, and the goal angle;
[0017] Determination of a focus event according to the shooting distance, the goal matching degree, the movement speed, the goalkeeper distance, and a preset comprehensive determination threshold when the goal matching degree is greater than a preset matching degree threshold;
[0018] Determination of a focus event type according to the goal matching degree and a defense ability factor.
[0019] Further, determination of a focus event according to the shooting distance, the goal matching degree, the movement speed, the goalkeeper distance, and a preset comprehensive determination threshold, comprises:
[0020] a defense ability factor is calculated according to the moving speed and the guard distance;
[0021] a comprehensive judgment factor is calculated according to the goal matching degree, the defense ability factor and the shooting distance;
[0022] when the comprehensive judgment factor is greater than a preset comprehensive judgment threshold, it is determined that an attention event occurs.
[0023] Further, an attention event type is determined according to the goal matching degree and the defense ability factor, comprising:
[0024] the goal matching degree is normalized to obtain a matching degree normalized value;
[0025] the defense ability factor is normalized to obtain a defense normalized value;
[0026] when the matching degree normalized value is greater than the defense normalized value, it is determined that the type of the attention event is a shooting event;
[0027] when the matching degree normalized value is less than or equal to the defense normalized value, it is determined that the type of the attention event is a defense event.
[0028] Further, a number of highlight moment videos are screened from all the attention videos according to the sound frequency, the pitch and a preset correlation threshold, comprising:
[0029] a standard deviation of the sound frequency at each moment within an initial moment to a preset screening duration is calculated to obtain a number of frequency fluctuation values;
[0030] an average of all the frequency fluctuation values is calculated to obtain a frequency average value;
[0031] when the frequency average value is greater than a preset frequency average threshold, a standard deviation of the pitch at each moment within the initial moment to the preset screening duration is calculated to obtain a number of pitch fluctuation values;
[0032] a number of highlight moment videos are screened from all the attention videos according to all the frequency fluctuation values, all the pitch fluctuation values and a preset correlation threshold.
[0033] Further, a number of highlight moment videos are screened from all the attention videos according to all the frequency fluctuation values, all the pitch fluctuation values and a preset correlation threshold, comprising:
[0034] all the frequency fluctuation values within the preset screening duration are normalized to obtain a frequency normalized data set;
[0035] all the pitch fluctuation values within the preset screening duration are normalized to obtain a pitch normalized data set;
[0036] calculating a correlation coefficient of the frequency normalized dataset and the pitch normalized dataset, to obtain a change correlation degree;
[0037] when the change correlation degree is greater than the preset correlation threshold, determining that the video of interest is the highlight video, to screen a plurality of highlight videos.
[0038] Further, adjusting the preset correlation threshold according to the number of all the videos of interest, the number of all the highlight videos in a preset adjustment period, and a comprehensive determination factor, to obtain an adjusted correlation threshold, or adjusting the preset comprehensive determination threshold to obtain an adjusted comprehensive determination threshold, including:
[0039] calculating a proportion of the highlight videos in all the videos of interest in the preset adjustment period, to obtain a video accuracy rate;
[0040] when the video accuracy rate is greater than a maximum value of a preset accuracy range, increasing the preset correlation threshold according to a relative deviation of the video accuracy rate and the maximum value of the preset accuracy range and a preset correlation adjustment coefficient, to obtain an adjusted correlation threshold;
[0041] when the video accuracy rate is less than a minimum value of the preset accuracy range, adjusting the preset comprehensive determination threshold according to the comprehensive determination factor of all the highlight videos, to obtain an adjusted comprehensive determination threshold.
[0042] Further, adjusting the preset comprehensive determination threshold according to the comprehensive determination factor of all the highlight videos, to obtain an adjusted comprehensive determination threshold, including:
[0043] calculating an average value of the comprehensive determination factor of all the highlight videos, to obtain a comprehensive average value;
[0044] calculating a relative deviation of the comprehensive average value and the preset comprehensive determination threshold, to obtain a comprehensive deviation;
[0045] increasing the preset comprehensive determination threshold according to the comprehensive deviation and a preset comprehensive factor adjustment coefficient, to obtain an adjusted comprehensive determination threshold.
[0046] Further, determining the shooting intention according to the leg key point and the shooting distance, including:
[0047] when the shooting distance is less than a preset distance threshold, calculating an inverse cosine function of the leg key point according to the leg key point coordinates, to obtain a leg swing amplitude;
[0048] when the leg swing amplitude is greater than a preset swing threshold, determining that the shooting intention exists.
[0049] On the other hand, the present invention also provides a video acquisition system based on deep learning acceleration, comprising:
[0050] The data acquisition module is used to collect in real time the coordinates of several key points on the legs of the dribbling player, the speed of the ball, the shooting distance, the angle of movement, the goalkeeper's distance from the goal line, the movement speed, the frequency of the audience's voice, and the pitch of the commentator during a football match.
[0051] The first determination module is connected to the data acquisition module and is used to determine whether there is a shooting intention based on the key points of the leg and the shooting distance.
[0052] The second determination module is connected to the data acquisition module and the first determination module respectively, and is used to determine the type of event of interest based on the shooting distance, the preset goal width, the movement speed, the movement angle, the movement speed, the goalkeeping distance and the preset comprehensive determination threshold.
[0053] A video acquisition module, which is connected to the second determination module, is used to acquire several videos of interest based on the deep learning acceleration model and the type of the interest event.
[0054] A filtering module, which is connected to the data acquisition module and the video acquisition module respectively, is used to filter out a number of exciting moments from all the videos of interest based on the sound frequency, the pitch and a preset relevance threshold.
[0055] An adjustment module, which is connected to the video acquisition module and the filtering module respectively, is used to adjust a preset relevance threshold based on the number of all the videos of interest, the number of all the videos of highlights, and a comprehensive judgment factor within a preset adjustment period, to obtain an adjustment relevance threshold, or to adjust a preset comprehensive judgment threshold to obtain an adjustment comprehensive judgment threshold.
[0056] A storage module, connected to the filtering module, stores the highlight videos, which are re-determined based on the adjusted relevance threshold or the adjusted comprehensive judgment threshold, into the video database.
[0057] Compared with existing technologies, the beneficial effects of this invention are that by collecting and analyzing multi-dimensional key data in real time, it can accurately determine the shooting intention and the type of event of interest, and filter out highlight videos by combining sound-related information; it can determine whether a player has a shooting intention based on key leg points and shooting distance, and determine the type of event of interest by combining data analysis of football and goalkeeper; the audience's voice frequency and the commentator's pitch can help determine whether a video clip is exciting; by adjusting the threshold, the selection process of highlight videos can be dynamically optimized, so that the highlight videos finally stored in the video database are more in line with the actual level of excitement, effectively solving the problems of low collection efficiency and low accuracy caused by over-reliance on complex models and single data.
[0058] Furthermore, by taking a comprehensive approach, the types of events of interest can be determined accurately and comprehensively, avoiding the one-sidedness caused by judging a single factor. This is because the shooting distance and goal width are directly related to the goal angle, while the speed of movement, the angle of movement, and the goal angle all affect the goal matching degree. When the goal matching degree exceeds the preset matching degree threshold, the event of interest is comprehensively judged. This fully considers factors such as the difficulty of the shot, the movement trend of the ball, and the goalkeeper's defensive performance. The type of event of interest is determined based on the goal matching degree and defensive ability factor, which further refines the classification of different types of shooting events. This facilitates the subsequent targeted collection of videos of interest and the selection of highlights, improving the targeting and effectiveness of video selection.
[0059] Furthermore, by comprehensively considering movement speed and goalkeeping distance to calculate the defensive capability factor, and then combining goal matching degree and shooting distance to calculate the comprehensive judgment factor, and comparing it with the preset comprehensive judgment threshold to determine the events of interest, this method closely links the defensive team's ability and the attacking team's shooting conditions in the game. It comprehensively evaluates the potential threat level and excitement level of the current shooting scene from the perspectives of both the attacking and defensive teams. The comprehensive judgment factor comprehensively reflects the probability of a successful shot and its entertainment value. Comparing it with the preset comprehensive judgment threshold can effectively filter out noteworthy moments in the game, ensuring that the collected video clips are more valuable and representative.
[0060] Furthermore, by normalizing the goal matching degree and defensive capability factor separately, the differences in their dimensions and orders of magnitude can be eliminated, making different parameters comparable. Comparing the normalized matching degree value with the normalized defensive value to determine the type of event of interest effectively considers both the probability of a shot and the intensity of the defense. When the normalized matching degree value is greater than the normalized defensive value, it indicates that the conditions for a shot are more favorable, and the probability of scoring is relatively high; therefore, it is classified as a shooting event. Conversely, when the normalized matching degree value is less than or equal to the normalized defensive value, it means that the defending side's blocking ability is more advantageous in the current situation, making it difficult to create an effective shooting threat; therefore, it is classified as a defensive event.
[0061] Furthermore, by calculating the frequency fluctuation value and then the frequency mean, we can initially determine the fluctuations in the audience's emotions. This is because fluctuations in sound frequency reflect the audience's level of reaction; when the frequency mean exceeds a preset threshold, it indicates that the audience is highly excited, potentially indicating exciting moments. Next, we calculate the pitch fluctuation value, which reflects the commentator's emotional changes; commentators often exhibit significant pitch changes during exciting moments. Finally, by combining the frequency fluctuation value, pitch fluctuation value, and a preset relevance threshold, we can filter out videos of exciting moments, reflecting the excitement of the competition from different angles and more comprehensively capturing the climaxes of the match.
[0062] Furthermore, through normalization processing, frequency-normalized datasets and pitch-normalized datasets are obtained. Then, the correlation coefficient between these two normalized datasets is calculated to obtain the change correlation. When the change correlation is greater than the preset correlation threshold, the corresponding video of interest is determined to be a highlight video. This is because the fluctuations in sound frequency and pitch are closely related to the excitement of the competition. Frequency fluctuations reflect the audience's reaction, and pitch fluctuations reflect the commentator's emotional changes. The change correlation between the two can comprehensively reflect the intensity of the competition and the atmosphere of the scene.
[0063] Furthermore, by calculating the video accuracy rate, the effectiveness of video filtering under the current threshold setting can be intuitively reflected. When the video accuracy rate is higher than or equal to the maximum value of the preset accuracy rate range, it indicates that the proportion of selected highlight videos is too high, which may lead to over-filtering and misjudging some general videos as highlight videos. In this case, increasing the preset relevance threshold can improve the filtering standard and reduce misjudgments. Conversely, when the video accuracy rate is lower than the minimum value of the preset accuracy rate range, it means that a large number of highlight videos are missed. In this case, the preset comprehensive judgment threshold is adjusted based on the comprehensive judgment factor of the highlight videos.
[0064] Furthermore, the average of the comprehensive judgment factors for all highlight videos is calculated to obtain the comprehensive mean, which reflects the overall level of the currently selected highlight videos in terms of the comprehensive judgment factors. Next, the relative deviation between the comprehensive mean and the preset comprehensive judgment threshold is calculated to obtain the comprehensive deviation, which reflects the gap between the threshold and the actual comprehensive judgment factors of the highlight videos. Increasing the preset comprehensive judgment threshold based on the comprehensive deviation and the preset comprehensive factor adjustment coefficient dynamically improves the screening criteria, making subsequent video screening more rigorous and accurate.
[0065] Furthermore, by determining whether the shooting distance is less than a preset distance threshold, it's possible to initially filter players into potential shooting positions. This is because the success rate and probability of a shot only significantly increase when a player is close to the goal. Next, the inverse cosine function of the key leg coordinates is calculated to obtain the leg swing amplitude. This operation quantifies the player's leg swing, and the leg swing amplitude is one of the key characteristics of the shooting action, directly related to the shooting intention. Finally, the leg swing amplitude is compared with a preset swing threshold; if it exceeds the threshold, a shooting intention is determined.
[0066] Furthermore, by collecting real-time data such as the coordinates of key points on the player's legs and combining this with the shooting distance to determine the shooting intention, a precise basis is provided for subsequent determination of the type of event of interest. Next, the type of event of interest is determined by comprehensively considering multi-dimensional data such as shooting distance, goal width, and movement speed. These data collectively determine whether a shooting scene has the potential for excitement. Subsequently, relevant videos are collected based on a deep learning acceleration model and the type of event of interest. The deep learning model can quickly identify and extract images related to these events, ensuring that the collected videos are targeted and timely. Then, videos of exciting moments are filtered by audience voice frequency and commentator pitch. This sound data reflects the atmosphere and emotional intensity of the scene, which is closely related to exciting moments. Finally, preset relevance thresholds and comprehensive judgment thresholds are dynamically adjusted. Based on the number of videos of interest and exciting moments within a preset adjustment period and the comprehensive judgment factor, the system can adaptively optimize the selection criteria, ensuring that the video database stores both rich and accurate videos of exciting moments. This effectively solves the problems of low collection efficiency and low accuracy caused by over-reliance on complex models and single data sources. Attached Figure Description
[0067] Figure 1 This is a flowchart of the deep learning-accelerated video acquisition method in this embodiment;
[0068] Figure 2 This is a logic diagram for determining the occurrence of an event of interest in this embodiment;
[0069] Figure 3 This embodiment defines the logic diagram for determining the type of event of interest.
[0070] Figure 4 This is a schematic diagram of the deep learning-accelerated video acquisition system in this embodiment. Detailed Implementation
[0071] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0072] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0073] Please see Figure 1 As shown, it is a flowchart of the video acquisition method based on deep learning acceleration in this embodiment;
[0074] On the one hand, this embodiment provides a video acquisition method based on deep learning acceleration, including:
[0075] Real-time data collection of key leg coordinates of dribbling players during football matches, football speed, shooting distance, angle of movement, goalkeeper's distance from the goal line, movement speed, audience voice frequency, and commentator's pitch;
[0076] Based on the aforementioned key points on the leg and the shooting distance, it is determined that there is an intention to shoot;
[0077] The type of event of interest is determined based on the shooting distance, preset goal width, movement speed, movement angle, movement speed, goalkeeping distance, and preset comprehensive judgment threshold.
[0078] Several videos of interest are collected based on the deep learning acceleration model and the types of events of interest.
[0079] Based on the sound frequency, the pitch, and a preset relevance threshold, several highlight videos are selected from all the videos of interest.
[0080] The preset relevance threshold is adjusted based on the number of all the videos under interest within the preset adjustment period, the number of all the highlight videos, and the comprehensive judgment factor to obtain the adjusted relevance threshold; or, the preset comprehensive judgment threshold is adjusted to obtain the adjusted comprehensive judgment threshold.
[0081] The highlights videos, re-determined based on the adjusted relevance threshold or the adjusted comprehensive judgment threshold, are stored in the video database.
[0082] In this embodiment, the coordinates of several key leg points of the dribbling player refer to the position coordinates of several key parts of the dribbling player's legs (such as the hip joint, knee joint, ankle joint, etc.). Their specific movement patterns are often associated with shooting actions and can be accurately extracted using computer vision technology by deploying multiple high-definition cameras around the playing field. The speed of the football represents its movement speed on the field. An accelerometer can be embedded inside the football to directly measure and transmit the football's speed data. The shooting distance refers to the distance from the football to the goal when the player is about to shoot. This can be directly measured using a laser rangefinder. The motion angle refers to the direction of the football's movement relative to the goal line. This can be calculated by analyzing the trajectory of the football to determine its direction vector, simultaneously obtaining the direction vector of the goal line, and using the dot product formula to calculate the angle between the two vectors. The goalkeeping distance from the goalkeeper to the goal line refers to the vertical distance from the goalkeeper's position to the goal line when guarding the goal. High-precision cameras can be installed in the goal area to capture the goalkeeper's position in real time, and the distance between the goalkeeper's position and the goal line can be calculated using the field coordinate system. Movement speed refers to the goalkeeper's movement speed on the field, which can be measured in real time by wearing a GPS device on the goalkeeper. Audience sound frequency reflects the pitch and variation of shouts and cheers. Microphone arrays can be installed in different areas of the stadium to collect audience sound signals, and then Fourier transform can be used to convert the sound signals from the time domain to the frequency domain to obtain the sound frequency. Commentator pitch refers to the highness or lowness of the commentator's voice. By setting up professional microphone equipment in the commentator's work area, pitch detection algorithms in audio analysis tools can be used to determine the pitch value of the commentator's voice.
[0083] The preset comprehensive judgment threshold is a benchmark value used to determine the type of event of interest. It depends on the characteristics of a football match, the definition of a highlight moment, historical match data, and expert experience, and is usually set between 0.5 and 1. In this embodiment, it is set to 0.7, which ensures that key moments are not missed while avoiding excessive interference from irrelevant information.
[0084] Different video capture times are selected for different types of events of interest. For shooting events, a relatively long capture time may be needed to fully record the entire process from the player kicking the ball to the ball entering the goal or being saved. For defensive events, the capture time can be shorter, focusing on recording the key actions of the defensive player and the moment of interception. Then, based on the deep learning acceleration model, high-speed image sensors and professional capture equipment are used to achieve high frame rate and high refresh rate video capture, resulting in several videos of interest.
[0085] Deep learning-accelerated models record videos of events of interest by selecting different video capture times based on the type of event after the event occurs.
[0086] 1. Model
[0087] Convolutional Neural Networks (ResNet): Suitable for processing image and video data. It introduces residual blocks to address the vanishing gradient problem during deep network training, enabling the construction of very deep network structures. This allows for better extraction of complex features, resulting in excellent performance in tasks such as image classification and object detection. It can also be used for video analysis tasks, taking video frames as input and extracting features from the frames to identify the types of events of interest.
[0088] 2. Initial model parameters
[0089] 2.1 Weights: The He initialization method is usually used. For example, in a convolutional layer in a residual block, assuming a layer has 64 input channels and 128 output channels, the initial weights of that layer will follow a random normal distribution with a mean of 0 and a standard deviation of 64 + 1282. This ensures the effective propagation of signals in the neural network and avoids gradient vanishing or exploding problems.
[0090] 2.2 Bias: 0.
[0091] 3. Training process
[0092] 3.1 Data Preparation: A large amount of football match video data was collected and labeled, such as labeling video clips containing shooting actions as "shooting events" and video clips containing defensive actions as "defensive events," etc. Then, the video data was preprocessed, including cropping and scaling, to meet the size requirements of the model input, uniformly cropping video frames to a size of 224×224 pixels.
[0093] 3.2 Building the training environment: Divide the dataset into a training set (for model training, accounting for, for example, 70%) and a validation set (for evaluating model performance, accounting for, for example, 30%). Determine the loss function as the cross-entropy loss function, use the Adam optimizer, and set the initial learning rate to 0.001.
[0094] 3.3 Training Steps: The training set data is input into the ResNet model. The model first performs forward propagation to calculate the predicted probability for each event type, and then calculates the cross-entropy loss between the predicted result and the true label. Using the backpropagation algorithm, the model's weights and bias parameters are updated based on the gradient of the loss function, adjusting the parameters in each iteration in a direction that reduces the loss function. This process is repeated continuously. After multiple epochs (training cycles), the model gradually learns how to distinguish different types of events of interest.
[0095] 4. The trained model
[0096] The trained model exhibits high performance metrics on the validation set, such as classification accuracy exceeding 90%, accurately classifying input video frames into their corresponding event types. The model's parameters have been optimized, with weight and bias values enabling the network to effectively extract and classify features from the video, thus achieving accurate identification of relevant event types in football match videos.
[0097] The preset relevance threshold is a standard used to filter out highlight videos from the videos you are interested in. It depends on the relevance of sound frequency and pitch to the highlight moments and is usually set between 0.5 and 1. In this embodiment, it is set to 0.6, which can ensure that most highlight moments are filtered out while avoiding over-filtering due to the randomness of sound factors.
[0098] By collecting multi-dimensional key data from real-time football matches, the system determines the presence of a shooting intention based on collected leg key points and shooting distance. It then determines the type of event of interest based on shooting distance, preset goal width, movement speed, movement angle, movement speed, goalkeeping distance, and a preset comprehensive judgment threshold. A deep learning-accelerated model and the identified event types are used to collect several videos of interest. Combining sound frequency, pitch, and a preset relevance threshold, several highlight videos are selected from all videos of interest. The preset relevance threshold or the preset comprehensive judgment threshold is adjusted based on the number of videos of interest, the number of highlight videos, and a comprehensive judgment factor within a preset adjustment period. Finally, the highlight videos, re-determined based on the adjusted relevance threshold or comprehensive judgment threshold, are stored in a video database.
[0099] By collecting and analyzing multi-dimensional key data in real time, it is possible to accurately determine shooting intentions and the types of events of interest, and to filter out highlight videos by combining sound-related information. The system can determine a player's shooting intention based on key leg points and shooting distance, and identify the types of events of interest by combining data analysis of football and goalkeeper. The frequency of the audience's voice and the pitch of the commentator can help determine whether a video clip is exciting. By adjusting thresholds, the selection process for highlight videos can be dynamically optimized, making the highlight videos ultimately stored in the video database more consistent with the actual level of excitement. This effectively solves the problems of low collection efficiency and low accuracy caused by over-reliance on complex models and single data sources.
[0100] Specifically, the type of event of interest is determined based on the shooting distance, preset goal width, movement speed, movement angle, movement speed, goalkeeping distance, and a preset comprehensive judgment threshold, including:
[0101] Calculate the goal angle based on the shooting distance and the preset goal width;
[0102] The goal matching degree is calculated based on the movement speed, the movement angle, and the goal angle;
[0103] When the goal matching degree is greater than a preset matching degree threshold, a attention event is determined based on the shooting distance, the goal matching degree, the moving speed, the goalkeeping distance, and a preset comprehensive judgment threshold.
[0104] The type of event of interest is determined based on the goal matching degree and defensive capability factor.
[0105] The preset goal width refers to the actual width of a football goal, which follows strict and uniform rules. In this embodiment, the standard goal size of 7.32 meters for 11-a-side football is adopted.
[0106] The preset matching threshold is a critical value used to determine whether the goal matching degree meets the criteria for determining a noteworthy event. It depends on the intensity of the match and the player's skill level, and is usually set between 0.6 and 0.9. In this embodiment, it is set to 0.7, which ensures that shooting events with a high probability of scoring are filtered out without being too strict.
[0107] The goal angle is calculated using the shooting distance and a preset goal width. Then, the goal matching degree is calculated by combining the movement speed, movement angle, and goal angle. If the goal matching degree is greater than a preset matching degree threshold, then the shooting distance, goal matching degree, movement speed, goalkeeping distance, and a preset comprehensive judgment threshold are used to determine whether a noteworthy event has occurred. Once a noteworthy event is determined, the specific type of the noteworthy event is determined based on the goal matching degree and the defensive ability factor.
[0108] By comprehensively considering various factors, the types of events of interest can be determined in a comprehensive and accurate manner, avoiding the one-sidedness caused by judging a single factor. This is because the shooting distance and goal width are directly related to the goal angle, while the speed of movement, the angle of movement, and the goal angle all affect the goal matching degree. When the goal matching degree exceeds the preset matching degree threshold, the event of interest is comprehensively judged. This fully considers factors such as the difficulty of the shot, the movement trend of the ball, and the goalkeeper's defensive situation. The type of event of interest is determined based on the goal matching degree and defensive ability factor, which further refines the classification of different types of shooting events. This facilitates the subsequent targeted collection of videos of interest and the selection of highlights, improving the targeting and effectiveness of video selection.
[0109] Please see Figure 2 As shown, it is the logic diagram for determining the occurrence of an event of interest in this embodiment;
[0110] The occurrence of a noteworthy event is determined based on the shooting distance, the goal matching degree, the movement speed, the goalkeeping distance, and a preset comprehensive judgment threshold, including:
[0111] The defensive capability factor is calculated based on the moving speed and the goalkeeping distance.
[0112] A comprehensive judgment factor is calculated based on the goal matching degree, the defensive ability factor, and the shooting distance;
[0113] When the comprehensive judgment factor is greater than the preset comprehensive judgment threshold, a matter of concern is determined to have occurred.
[0114] The defensive capability factor is calculated based on movement speed and goal distance. Then, the goal matching degree, defensive capability factor and shooting distance are combined to further calculate the comprehensive judgment factor. This comprehensive judgment factor is then compared with the preset comprehensive judgment threshold. If the comprehensive judgment factor exceeds the preset threshold, the event of interest is determined to have occurred.
[0115] By comprehensively considering movement speed and goalkeeping distance to calculate the defensive capability factor, and then combining goal matching degree and shooting distance to calculate the comprehensive judgment factor, and comparing it with the preset comprehensive judgment threshold to determine the events of interest, this method closely links the defensive team's ability and the offensive team's shooting conditions in the game. It comprehensively evaluates the potential threat level and excitement level of the current shooting scene from the perspectives of both the offensive and defensive sides. The comprehensive judgment factor comprehensively reflects the probability of a successful shot and its entertainment value. Comparing it with the preset comprehensive judgment threshold can effectively filter out noteworthy moments in the game, ensuring that the collected video clips are more valuable and representative.
[0116] Please continue reading. Figure 3 As shown, this is the logic diagram for determining the type of event of interest in this embodiment;
[0117] The types of events of interest are determined based on the goal matching degree and defensive capability factors, including:
[0118] The matching degree of the goal is normalized to obtain a normalized matching degree value;
[0119] The defense capability factor is normalized to obtain the defense normalization value;
[0120] When the matching degree normalization value is greater than the defense normalization value, the type of the event of interest is determined to be a shooting event;
[0121] When the matching degree normalization value is less than or equal to the defense normalization value, the type of the event of interest is determined to be a defense event.
[0122] By normalizing the goal matching degree and defensive capability factor separately, we obtain the matching degree normalized value and the defensive normalized value. This scales the goal matching degree and defensive capability factor to between 0 and 1, eliminating differences in dimensions and orders of magnitude between different data points. This is existing technology and will not be elaborated further here. Then, we determine the type of event of interest by comparing these two normalized values. Specifically, if the matching degree normalized value is greater than the defensive normalized value, the event of interest is determined to be a shooting event; conversely, if the matching degree normalized value is less than or equal to the defensive normalized value, the event of interest is determined to be a defensive event.
[0123] By normalizing the goal matching degree and defensive capability factor separately, the differences in their dimensions and orders of magnitude can be eliminated, making different parameters comparable. Comparing the normalized matching degree value with the normalized defensive value to determine the type of event of interest effectively considers both the probability of a shot and the intensity of the defense. When the normalized matching degree value is greater than the normalized defensive value, it indicates that the conditions for a shot are more favorable, and the probability of scoring is relatively high; therefore, it is classified as a shooting event. Conversely, when the normalized matching degree value is less than or equal to the normalized defensive value, it means that the defending side's blocking ability is more advantageous in the current situation, making it difficult to create an effective shooting threat; therefore, it is classified as a defensive event.
[0124] Specifically, based on the sound frequency, the pitch, and a preset relevance threshold, a number of highlight videos are selected from all the videos of interest, including:
[0125] Calculate the standard deviation of the sound frequency at each time point from the initial time to the preset screening time to obtain several frequency fluctuation values;
[0126] Calculate the average value of all the frequency fluctuation values to obtain the frequency mean.
[0127] When the average frequency is greater than a preset average frequency threshold, the standard deviation of the pitch is calculated for each time point from the initial time to the preset screening time, and several pitch fluctuation values are obtained.
[0128] Based on all the frequency fluctuation values, all the pitch fluctuation values, and a preset relevance threshold, a number of highlight videos are selected from all the videos of interest.
[0129] The preset filtering duration is the time length used to calculate sound frequency fluctuations and pitch fluctuations. It depends on the average duration of exciting moments in a football match and the system's real-time requirements. It is usually set between 5 and 10 seconds. In this embodiment, it is set to 7 seconds, which can cover the key moments of exciting moments while avoiding excessive time.
[0130] The preset average frequency threshold is a critical value used to reflect the excitement level of the live performance. It depends on the change pattern of sound frequency when the audience is emotionally excited and the typical intensity of football matches, and is usually set between 50Hz and 200Hz. In this embodiment, it is set to 120Hz, which can effectively distinguish between ordinary moments and moments when the audience is more emotionally excited, thereby more accurately filtering out potentially exciting video moments.
[0131] By calculating the standard deviation of sound frequencies at each moment from the initial time to the preset filtering time, multiple frequency fluctuation values are obtained. Then, the average of all frequency fluctuation values is calculated to obtain the frequency mean. When the frequency mean is greater than a preset frequency mean threshold, the system further calculates the standard deviation of pitch at each moment from the initial time to the preset filtering time, obtaining multiple pitch fluctuation values. Finally, based on all frequency fluctuation values, pitch fluctuation values, and a preset relevance threshold, a number of highlight videos are selected from all the videos of interest.
[0132] By calculating frequency fluctuation values and further calculating the frequency mean, we can initially determine the emotional fluctuations of the live audience. This is because fluctuations in sound frequency reflect the audience's level of reaction; when the frequency mean exceeds a preset threshold, it indicates that the audience is highly excited, potentially indicating exciting moments. Next, we calculate pitch fluctuation values, which reflect the commentator's emotional changes. Commentators often exhibit significant pitch changes during exciting moments. Finally, we combine frequency fluctuation values, pitch fluctuation values, and a preset correlation threshold to filter out videos of exciting moments, reflecting the excitement of the competition from different angles and more comprehensively capturing the climaxes of the match.
[0133] Specifically, based on all the frequency fluctuation values, all the pitch fluctuation values, and a preset relevance threshold, a number of highlight videos are selected from all the videos of interest, including:
[0134] Normalize all frequency fluctuation values within the preset filtering time to obtain a frequency normalized dataset;
[0135] Normalize all the pitch fluctuation values within the preset screening time to obtain a pitch normalized dataset;
[0136] Calculate the correlation coefficient between the frequency-normalized dataset and the pitch-normalized dataset to obtain the correlation degree of change;
[0137] When the change relevance is greater than the preset relevance threshold, the video of interest is determined to be the highlight video, so as to filter out a number of highlight videos.
[0138] This involves scaling the goal matching degree and defensive capability factor to between 0 and 1 respectively, eliminating the differences in dimensions and orders of magnitude between different data. This is an existing technology and will not be elaborated on here.
[0139] By normalizing all frequency fluctuation values and all pitch fluctuation values within a preset filtering period, frequency-normalized datasets and pitch-normalized datasets are obtained. This involves scaling the frequency and pitch fluctuation values to between 0 and 1, eliminating differences in dimensions and orders of magnitude between different data points—a technique already in use and will not be elaborated upon here. Next, the correlation coefficient between the frequency-normalized dataset and the pitch-normalized dataset is calculated to obtain the correlation degree. If this correlation degree is greater than a preset correlation threshold, the corresponding video is determined to be a highlight video, thus selecting several highlight videos.
[0140] By normalizing the data, frequency-normalized datasets and pitch-normalized datasets are obtained. Then, the correlation coefficient between these two normalized datasets is calculated to obtain the correlation degree of change. When the correlation degree of change is greater than the preset correlation threshold, the corresponding video of interest is determined to be a highlight video. This is because the fluctuations in sound frequency and pitch are closely related to the excitement of the competition. Frequency fluctuations reflect the audience's reaction, and pitch fluctuations reflect the commentator's emotional changes. The correlation degree of change between the two can comprehensively reflect the intensity of the competition and the atmosphere of the scene.
[0141] Specifically, the preset relevance threshold is adjusted based on the number of all the videos under interest within a preset adjustment period, the number of all the highlight videos, and a comprehensive judgment factor to obtain an adjusted relevance threshold; or, the preset comprehensive judgment threshold is adjusted to obtain an adjusted comprehensive judgment threshold, including:
[0142] The proportion of highlight videos among all the videos under watch within the preset adjustment period is calculated to obtain the video accuracy rate;
[0143] When the video accuracy is greater than the maximum value of the preset accuracy range, the preset correlation threshold is increased based on the relative deviation between the video accuracy and the maximum value of the preset accuracy range and the preset correlation adjustment coefficient, to obtain the adjusted correlation threshold.
[0144] When the video accuracy is less than the minimum value of the preset accuracy range, the preset comprehensive judgment threshold is adjusted according to the comprehensive judgment factor of all the highlight videos to obtain the adjusted comprehensive judgment threshold.
[0145] The preset adjustment duration refers to the time window used to evaluate and adjust the relevance threshold and the comprehensive judgment threshold. It depends on the real-time requirements of the video acquisition system for event processing and the need for data statistical stability, and is usually set between 1 hour and 3 hours. In this embodiment, it is set to 2 hours to ensure that the data sample size has a certain representativeness and to avoid excessive data fluctuations due to a too short duration.
[0146] The preset accuracy range is a range used to measure whether the proportion of highlight videos among the videos watched is reasonable. It depends on the requirements for video capture quality and the definition of a highlight moment in the game, and is usually set between 50% and 90%. In this embodiment, it is set to [60%, 80%], which ensures that the selected highlight videos are of high quality, while avoiding overly strict screening that might miss some potential highlights.
[0147] The preset relevance adjustment coefficient determines the magnitude of the threshold adjustment. It depends on the sensitivity to changes in video accuracy and the required stability of the filtering results, and is typically set between 0.1 and 0.5. In this embodiment, it is set to 0.3 to ensure the system has sufficient responsiveness to changes in video accuracy while avoiding instability caused by excessively large threshold adjustments due to an overly large coefficient.
[0148] The video accuracy is calculated by determining the proportion of highlight videos among all watched videos. Then, if the video accuracy is greater than or equal to the maximum value of a preset accuracy range, the preset relevance threshold is increased based on the relative deviation between the video accuracy and this maximum value, as well as a preset relevance adjustment coefficient, thus obtaining an adjusted relevance threshold. If the video accuracy is less than the minimum value of the preset accuracy range, the preset comprehensive judgment threshold is adjusted based on the comprehensive judgment factor of all highlight videos, thus obtaining an adjusted comprehensive judgment threshold.
[0149] By calculating the video accuracy rate, the effectiveness of video filtering under the current threshold setting can be intuitively reflected. When the video accuracy rate is higher than or equal to the maximum value of the preset accuracy rate range, it indicates that the proportion of selected highlight videos is too high, which may lead to over-filtering and misjudging some general videos as highlight videos. In this case, increasing the preset relevance threshold can improve the filtering standard and reduce misjudgments. Conversely, when the video accuracy rate is lower than the minimum value of the preset accuracy rate range, it means that a large number of highlight videos are being missed. In this case, the preset comprehensive judgment threshold should be adjusted based on the comprehensive judgment factor of the highlight videos.
[0150] Specifically, the preset comprehensive judgment threshold is adjusted based on the comprehensive judgment factor of all the aforementioned highlight video clips, resulting in the adjusted comprehensive judgment threshold, including:
[0151] Calculate the average value of the comprehensive judgment factors for all the aforementioned highlight video clips to obtain the comprehensive mean;
[0152] The relative deviation between the overall mean and the preset overall judgment threshold is calculated to obtain the overall deviation;
[0153] The preset comprehensive judgment threshold is increased by increasing the preset comprehensive judgment threshold based on the comprehensive deviation and the preset comprehensive factor adjustment coefficient to obtain the adjusted comprehensive judgment threshold.
[0154] The preset comprehensive factor adjustment coefficient is used to adjust the preset comprehensive judgment threshold based on the comprehensive deviation. It depends on the fluctuation range of the comprehensive judgment factor, the rate of change of the competition data, and the required sensitivity of the threshold adjustment, and is typically set between 0.1 and 0.5. In this embodiment, it is set to 0.3, which allows for timely adaptation to data changes without causing threshold instability due to excessive adjustment.
[0155] The overall mean is obtained by calculating the average of the comprehensive judgment factors for all highlight videos. Next, the relative deviation between the overall mean and the preset comprehensive judgment threshold is calculated to obtain the overall deviation. Then, based on the overall deviation and the preset comprehensive factor adjustment coefficient, the preset comprehensive judgment threshold is increased to obtain the adjusted comprehensive judgment threshold.
[0156] The overall mean is calculated by averaging the comprehensive judgment factors of all featured video clips. This average reflects the overall level of the currently selected featured videos in terms of these factors. Next, the relative deviation between the overall mean and the preset comprehensive judgment threshold is calculated to obtain the overall deviation. This deviation reflects the difference between the threshold and the actual comprehensive judgment factors of the featured videos. Finally, by increasing the preset comprehensive judgment threshold based on the overall deviation and the preset comprehensive factor adjustment coefficient, the screening criteria are dynamically improved, making subsequent video screening more rigorous and accurate.
[0157] Specifically, determining the presence of a shooting intent based on the aforementioned key leg points and the shooting distance includes:
[0158] When the shooting distance is less than a preset distance threshold, the inverse cosine function of the key points is calculated based on the coordinates of the key points on the leg to obtain the leg swing amplitude;
[0159] When the leg swing amplitude exceeds a preset swing threshold, it is determined that there is an intention to shoot.
[0160] The preset distance threshold is a critical value used to determine whether the shooting distance is close enough to potentially indicate a shooting intention. It depends on the player's shooting range and the goalkeeper's defensive range, and is usually set between 5 and 15 meters. In this embodiment, it is set to 10 meters, which can cover most threatening shooting areas, while avoiding misjudging situations where the distance is too far and the probability of a shot is low as having a shooting intention.
[0161] The preset swing threshold is a standard value used to determine whether the leg swing amplitude is large enough to determine whether there is a shooting intention. It depends on the typical leg swing amplitude in a shooting action and the individual differences of players, and is usually set between 0.5 radians and 1.5 radians. In this embodiment, it is set to 0.8 radians, which can effectively distinguish between ordinary running or passing actions and leg swings with shooting intentions, thus improving the accuracy of the judgment.
[0162] When the shooting distance is less than a preset distance threshold, the arccosine function of the key points on the leg is calculated based on their coordinates to obtain the leg swing amplitude. Then, the calculated leg swing amplitude is compared with a preset swing threshold. If the leg swing amplitude is greater than the preset swing threshold, a shooting intent is ultimately determined.
[0163] By determining whether the shooting distance is less than a preset distance threshold, we can initially filter out players who are in a potentially shooting position. This is because the success rate and probability of a shot only significantly increase when a player is close to the goal. Next, we calculate the inverse cosine function of the key leg coordinates to obtain the leg swing amplitude. This operation quantifies the player's leg swing, and the leg swing amplitude is one of the key characteristics of the shooting action, directly related to the shooting intention. Finally, we compare the leg swing amplitude with a preset swing threshold; if it exceeds the threshold, we determine that there is a shooting intention.
[0164] Please continue reading. Figure 4 As shown, it is a schematic diagram of the video acquisition system based on deep learning acceleration in this embodiment;
[0165] On the other hand, this embodiment provides a video acquisition system based on deep learning acceleration, including:
[0166] The data acquisition module is used to collect in real time the coordinates of several key points on the legs of the dribbling player, the speed of the ball, the shooting distance, the angle of movement, the goalkeeper's distance from the goal line, the movement speed, the frequency of the audience's voice, and the pitch of the commentator during a football match.
[0167] The first determination module is connected to the data acquisition module and is used to determine whether there is a shooting intention based on the key points of the leg and the shooting distance.
[0168] The second determination module is connected to the data acquisition module and the first determination module respectively, and is used to determine the type of event of interest based on the shooting distance, the preset goal width, the movement speed, the movement angle, the movement speed, the goalkeeping distance and the preset comprehensive determination threshold.
[0169] A video acquisition module, which is connected to the second determination module, is used to acquire several videos of interest based on the deep learning acceleration model and the type of the interest event.
[0170] A filtering module, which is connected to the data acquisition module and the video acquisition module respectively, is used to filter out a number of exciting moments from all the videos of interest based on the sound frequency, the pitch and a preset relevance threshold.
[0171] An adjustment module, which is connected to the video acquisition module and the filtering module respectively, is used to adjust a preset relevance threshold based on the number of all the videos of interest, the number of all the videos of highlights, and a comprehensive judgment factor within a preset adjustment period, to obtain an adjustment relevance threshold, or to adjust a preset comprehensive judgment threshold to obtain an adjustment comprehensive judgment threshold.
[0172] A storage module, connected to the filtering module, stores the highlight videos, which are re-determined based on the adjusted relevance threshold or the adjusted comprehensive judgment threshold, into the video database.
[0173] The system collects real-time data on the coordinates of several key leg points of the dribbling player, the speed of the ball, the shooting distance, the angle of motion, the goalkeeper's distance from the goal line, movement speed, the frequency of the audience's voice, and the pitch of the commentator. The system determines the presence of a shooting intention based on the key leg points and shooting distance. It then determines the type of event of interest based on the shooting distance, preset goal width, ball speed, angle of motion, movement speed, goalkeeper's distance, and a preset comprehensive judgment threshold. Several videos of interest are collected using a deep learning acceleration model and the type of event. Highlight videos are selected from these videos using the audience's voice frequency, the commentator's pitch, and a preset relevance threshold. The preset relevance threshold or preset comprehensive judgment threshold is adjusted based on the number of videos of interest within a preset adjustment period, the number of highlight videos, and a comprehensive judgment factor, resulting in a new adjusted relevance threshold or adjusted comprehensive judgment threshold. Finally, the highlight videos re-determined based on the adjusted relevance threshold or comprehensive judgment threshold are stored in the video database.
[0174] By collecting real-time data such as the coordinates of key points on the player's legs and combining this with the shooting distance to determine the shooting intention, a precise basis is provided for subsequent determination of the type of event of interest. Next, the type of event of interest is determined by integrating multi-dimensional data such as shooting distance, goal width, and movement speed. These data collectively determine whether a shooting scene has the potential for excitement. Subsequently, relevant videos are collected based on a deep learning acceleration model and the type of event of interest. The deep learning model can quickly identify and extract images related to these events, ensuring that the collected videos are targeted and timely. Then, videos of exciting moments are filtered by audience voice frequency and commentator pitch. This sound data reflects the atmosphere and emotional intensity of the scene, which is closely related to exciting moments. Finally, preset relevance thresholds and comprehensive judgment thresholds are dynamically adjusted. Based on the number of videos of interest and exciting moments within a preset adjustment period and the comprehensive judgment factor, the system can adaptively optimize the selection criteria, ensuring that the video database stores both rich and accurate videos of exciting moments. This effectively solves the problems of low collection efficiency and low accuracy caused by over-reliance on complex models and single data sources.
[0175] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A video acquisition method based on deep learning acceleration, characterized in that, include: Real-time data collection of key leg coordinates of dribbling players during football matches, football speed, shooting distance, angle of movement, goalkeeper's distance from the goal line, movement speed, audience voice frequency, and commentator's pitch; Based on the aforementioned key points on the leg and the shooting distance, it is determined that there is an intention to shoot; The type of event of interest is determined based on the shooting distance, preset goal width, movement speed, movement angle, movement speed, goalkeeping distance, and preset comprehensive judgment threshold. Several videos of interest are collected based on the deep learning acceleration model and the types of events of interest. Based on the sound frequency, the pitch, and a preset relevance threshold, several highlight videos are selected from all the videos of interest. The preset relevance threshold is adjusted based on the number of all the videos under interest within the preset adjustment period, the number of all the highlight videos, and the comprehensive judgment factor to obtain the adjusted relevance threshold; or, the preset comprehensive judgment threshold is adjusted to obtain the adjusted comprehensive judgment threshold. The highlights videos, re-determined based on the adjusted relevance threshold or the adjusted comprehensive judgment threshold, are stored in the video database.
2. The video acquisition method based on deep learning acceleration according to claim 1, characterized in that, The types of events of interest are determined based on the shooting distance, preset goal width, movement speed, movement angle, moving speed, goalkeeping distance, and a preset comprehensive judgment threshold, including: Calculate the goal angle based on the shooting distance and the preset goal width; The goal matching degree is calculated based on the movement speed, the movement angle, and the goal angle; When the goal matching degree is greater than a preset matching degree threshold, a attention event is determined based on the shooting distance, the goal matching degree, the moving speed, the goalkeeping distance, and a preset comprehensive judgment threshold. The type of event of interest is determined based on the goal matching degree and defensive capability factor.
3. The video acquisition method based on deep learning acceleration according to claim 2, characterized in that, The occurrence of a noteworthy event is determined based on the shooting distance, the goal matching degree, the movement speed, the goalkeeping distance, and a preset comprehensive judgment threshold, including: The defensive capability factor is calculated based on the moving speed and the goalkeeping distance. A comprehensive judgment factor is calculated based on the goal matching degree, the defensive ability factor, and the shooting distance; When the comprehensive judgment factor is greater than the preset comprehensive judgment threshold, a matter of concern is determined to have occurred.
4. The video acquisition method based on deep learning acceleration according to claim 3, characterized in that, The types of events of interest are determined based on the goal matching degree and defensive capability factors, including: The matching degree of the goal is normalized to obtain a normalized matching degree value; The defense capability factor is normalized to obtain the defense normalization value; When the matching degree normalization value is greater than the defense normalization value, the type of the event of interest is determined to be a shooting event; When the matching degree normalization value is less than or equal to the defense normalization value, the type of the event of interest is determined to be a defense event.
5. The video acquisition method based on deep learning acceleration according to claim 4, characterized in that, Based on the sound frequency, the pitch, and a preset relevance threshold, a number of highlight videos are selected from all the videos of interest, including: Calculate the standard deviation of the sound frequency at each time point from the initial time to the preset screening time to obtain several frequency fluctuation values; Calculate the average value of all the frequency fluctuation values to obtain the frequency mean. When the average frequency is greater than a preset average frequency threshold, the standard deviation of the pitch is calculated for each time point from the initial time to the preset screening time, and several pitch fluctuation values are obtained. Based on all the frequency fluctuation values, all the pitch fluctuation values, and a preset relevance threshold, a number of highlight videos are selected from all the videos of interest.
6. The video acquisition method based on deep learning acceleration according to claim 5, characterized in that, Based on all the frequency fluctuation values, all the pitch fluctuation values, and a preset relevance threshold, a number of highlight videos are selected from all the videos of interest, including: Normalize all frequency fluctuation values within the preset filtering time to obtain a frequency normalized dataset; Normalize all the pitch fluctuation values within the preset screening time to obtain a pitch normalized dataset; Calculate the correlation coefficient between the frequency-normalized dataset and the pitch-normalized dataset to obtain the correlation degree of change; When the change relevance is greater than the preset relevance threshold, the video of interest is determined to be the highlight video, so as to filter out a number of highlight videos.
7. The video acquisition method based on deep learning acceleration according to claim 6, characterized in that, The preset relevance threshold is adjusted based on the number of all the videos under interest within the preset adjustment period, the number of all the highlight videos, and a comprehensive judgment factor to obtain an adjusted relevance threshold; or, the preset comprehensive judgment threshold is adjusted to obtain an adjusted comprehensive judgment threshold, including: The proportion of highlight videos among all the videos under watch within the preset adjustment period is calculated to obtain the video accuracy rate; When the video accuracy is greater than the maximum value of the preset accuracy range, the preset correlation threshold is increased based on the relative deviation between the video accuracy and the maximum value of the preset accuracy range and the preset correlation adjustment coefficient, to obtain the adjusted correlation threshold. When the video accuracy is less than the minimum value of the preset accuracy range, the preset comprehensive judgment threshold is adjusted according to the comprehensive judgment factor of all the highlight videos to obtain the adjusted comprehensive judgment threshold.
8. The video acquisition method based on deep learning acceleration according to claim 7, characterized in that, The preset comprehensive judgment threshold is adjusted based on the comprehensive judgment factor of all the aforementioned highlight video clips to obtain the adjusted comprehensive judgment threshold, including: Calculate the average value of the comprehensive judgment factors for all the aforementioned highlight video clips to obtain the comprehensive mean; The relative deviation between the overall mean and the preset overall judgment threshold is calculated to obtain the overall deviation; The preset comprehensive judgment threshold is increased by increasing the preset comprehensive judgment threshold based on the comprehensive deviation and the preset comprehensive factor adjustment coefficient to obtain the adjusted comprehensive judgment threshold.
9. The video acquisition method based on deep learning acceleration according to claim 8, characterized in that, Determining the presence of a shooting intent based on the aforementioned key leg points and the shooting distance includes: When the shooting distance is less than a preset distance threshold, the inverse cosine function of the key points is calculated based on the coordinates of the key points on the leg to obtain the leg swing amplitude; When the leg swing amplitude exceeds a preset swing threshold, it is determined that there is an intention to shoot.
10. A video acquisition system accelerated by deep learning, based on the video acquisition method accelerated by deep learning according to any one of claims 1-9, characterized in that, include: The data acquisition module is used to collect in real time the coordinates of several key points on the legs of the dribbling player, the speed of the ball, the shooting distance, the angle of movement, the goalkeeper's distance from the goal line, the movement speed, the frequency of the audience's voice, and the pitch of the commentator during a football match. The first determination module is connected to the data acquisition module and is used to determine whether there is a shooting intention based on the key points of the leg and the shooting distance. The second determination module is connected to the data acquisition module and the first determination module respectively, and is used to determine the type of event of interest based on the shooting distance, the preset goal width, the movement speed, the movement angle, the movement speed, the goalkeeping distance and the preset comprehensive determination threshold. A video acquisition module, which is connected to the second determination module, is used to acquire several videos of interest based on the deep learning acceleration model and the type of the interest event. A filtering module, which is connected to the data acquisition module and the video acquisition module respectively, is used to filter out a number of exciting moments from all the videos of interest based on the sound frequency, the pitch and a preset relevance threshold. An adjustment module, which is connected to the video acquisition module and the filtering module respectively, is used to adjust a preset relevance threshold based on the number of all the videos of interest, the number of all the videos of highlights, and a comprehensive judgment factor within a preset adjustment period, to obtain an adjustment relevance threshold, or to adjust a preset comprehensive judgment threshold to obtain an adjustment comprehensive judgment threshold. A storage module, connected to the filtering module, stores the highlight videos, which are re-determined based on the adjusted relevance threshold or the adjusted comprehensive judgment threshold, into the video database.
Citation Information
Patent Citations
Football match behavior recognition method and device based on deep learning and terminal equipment
CN110378245A
Method for extracting wonderful clip of badminton event video based on machine learning
CN111291617A
Multi-channel power line carrier communication system
CN119582876A