Sports event collection generation method and system based on artificial intelligence
Through multimodal data processing and deep learning models, video, audio, sensors and statistical data are integrated, and the problem of missing detection of key events in the existing technology is solved, and more accurate and personalized sports event highlights are generated, improving audience experience and content production efficiency.
Patent Information
- Application Number
- CN202510732997.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing automatic highlight generation method lacks the collaborative utilization of multi-dimensional real-time data in sports events, resulting in missed detection of key events and the generated highlights are not comprehensive and accurate enough, especially in a complex and changeable event environment, which cannot judge high-value events based on audience reactions and athlete status.
Through multimodal data processing and deep learning models, video, audio, sensors and statistical data are integrated, machine learning-based detection classification models are established, event events are identified and scored, and personalized highlights are generated based on user preferences.
It achieves a more accurate, comprehensive and personalized event highlight generation, improves the audience experience, reduces the complexity and time cost of manual operations, and ensures the effective presentation and smooth viewing experience of the climax part.
Smart Images

Figure CN120264103A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method and system for generating sports event highlights based on artificial intelligence. Background Art
[0002] With the wide application of artificial intelligence technology in the production of sports event programs, the automatic highlight generation technology has gradually become an important means to improve the production efficiency and quality of programs. However, the existing automatic highlight generation methods still have some deficiencies in intelligent analysis, data processing, and user experience, which affect the accuracy and viewing pleasure of the highlights.
[0003] After retrieval, "An Automatic Highlight System and Method for Event Programs" with publication number CN110012348B was published on September 10, 2019. This patent discloses an automatic highlight system and method for event programs, which relates to the field of sports event production technology. The system includes a data aggregation module, an intelligent analysis module, and an automatic highlight module. The intelligent analysis module extracts event tags of event data using intelligent algorithms based on the constructed knowledge graph, generates feature tags and scoring results, and the automatic highlight module intercepts event segments from high to low according to the scoring results until the preset highlight time length and number of events are reached, thereby realizing automatic highlights.
[0004] Existing methods rely on fixed-perspective video streams and commentary audio, lacking the collaborative utilization of multi-dimensional real-time data. Single-modal models are difficult to associate the intensity of the audience's cheers with the goal event, or combine the sudden increase in the athlete's heart rate to judge the critical confrontation moment, resulting in the omission of high-value events; In addition, in the existing methods, the intelligent analysis module mainly relies on the knowledge graph and preset scoring rules, and has limited ability to process and analyze event data. Especially in a complex and changeable event environment, key events will be missed. For example, in a football game, the existing technology cannot automatically identify the climax moment of the game (such as the goal moment) by real-time analyzing the decibel change of the audience's cheers and facial expressions (such as the smiles of the audience captured by the camera) and preferentially include it in the highlights; in a basketball game, the existing technology also cannot dynamically adjust the emotional tendency of the highlights by combining the excitement level of the commentator's tone and the sound field distribution of the audience seats, resulting in the generated highlights being incomplete and inaccurate. Summary of the Invention
[0005] Based on the above purpose, the present invention provides a method for generating sports event highlights based on artificial intelligence, including the following steps: S1: Use a video acquisition device to record a sports event, an audio acquisition device to collect on-site sounds, and at the same time obtain real-time data of the event through sensors and a statistical system; S2: Adjust the frame rate, reduce noise, and compress the video stream information, reduce noise and normalize the volume of the audio stream information, structure the event statistics information, and perform time alignment processing on the sensor data; S3: Establish a detection and classification model based on machine learning, use the processed video stream information, audio stream information, event statistics information, and sensor data in S2 as inputs, and output an event list including event type, occurrence time, event description, and event importance score; S4: According to the importance of the events and the highlight length set by the user, select the event segments with the highest importance scores, and for each selected event segment, generate an event segment containing video, audio, and text description, and generate a preliminary highlight containing multiple important event segments; S5: Edit the preliminary highlight, remove duplicate and redundant segments, adjust the segment order, add transition effects and background music; perform personalized editing according to the highlight parameters set by the user, such as highlight duration, event type, and emotional tendency.
[0006] Preferably, S1 specifically includes the following: S1.1: Use a video acquisition device to record a sports event, and use an audio acquisition device to collect on-site sounds, including the voices of athletes, the cheers of the audience, and the commentary of the event commentator; The video acquisition device includes multiple fixed cameras and mobile cameras. The fixed cameras are installed at key positions on the playing field, including both sides of the court and above the audience stands. The mobile cameras are installed on drones or track systems and can flexibly adjust the angle and position according to the dynamic changes of the event; The audio acquisition device includes multiple microphones, which are installed at different positions on the playing field, including the center of the court, the audience stands, and the commentary booth; S1.2: Obtain real-time data of the event through multiple sensors, including the physiological data, motion data, and environmental data of the athletes; The sensors include heart rate monitors, accelerometers, GPS positioning modules, and temperature and humidity sensors. The sensors are installed on the clothing, shoes, or equipment of the athletes, as well as at various key positions on the playing field; Obtain real-time statistical data of the event through a statistical system, including game time, score, number of fouls, and number of substitutions. The statistical system includes a handheld device used by the referee, a large display screen in the center of the playing field, and a back-end server.
[0007] Preferably, in S2, the steps for adjusting the frame rate, reducing noise, and compressing the video stream information specifically include the following: T1: Detect the original frame rate of the video stream using a video analysis algorithm. Dynamically adjust the frame rate of the video stream according to a preset target frame rate and the dynamic characteristics of the video content. The dynamic characteristics include the fast and slow changes of actions. During the frame rate adjustment process, interpolate the missing frames using an interpolation algorithm. T2: Use a time-domain noise reduction algorithm to remove noise in the time domain through inter-frame differences between multiple frames. Then, use a spatial-domain noise reduction algorithm to remove noise in the spatial domain through a spatial filter. The spatial filter includes a Gaussian filter and a median filter. Finally, use a deblocking effect algorithm to remove the block effect caused by compression. T3: Encode the video stream using the H.265 encoding standard. Then, further reduce the amount of encoded data through entropy coding technology. The entropy coding technology is specifically CABAC. Finally, dynamically adjust the compression parameters according to the network bandwidth and storage space through a bitrate control algorithm.
[0008] Preferably, S3 specifically includes the following steps: S3.1: Establish a detection and classification model based on deep learning, specifically a multi-modal Transformer model. S3.2: The input data of the model includes the processed video stream information, audio stream information, event statistics information, and sensor data. The video stream information is input into the visual branch of the model, and a CNN is used to extract image features. The audio stream information is input into the audio branch of the model, and a CNN is used to extract audio features. The event statistics information and sensor data are input into the text and sensor branches of the model, and an LSTM network is used to extract time series features. S3.3: In the middle layer of the model, use a multi-modal attention mechanism to fuse features of different modalities to generate a comprehensive feature vector. The multi-modal attention mechanism calculates the similarity between features of different modalities and selects the most relevant features for fusion. S3.4: At the output layer of the model, use a classifier to perform event detection on the comprehensive feature vector to generate an event list including event type, occurrence time, event description, and event importance score. The classifier uses the Softmax function to output the probability distribution of each event type and selects the event with the highest probability as the detection result. S3.5: Train the model using a large-scale labeled dataset to ensure the robustness and accuracy of the model in different event environments. During the training process, use the Adam optimizer, set the learning rate to 0.001, the batch size to 32, and the number of training epochs to 50. Use the cross-entropy loss function as the loss function to ensure the training effect and convergence speed of the model.
[0009] Preferably, the event types output by the model in S3 include goals, fouls, substitutions, and offsides. Each event type corresponds to a preset label. The model outputs the occurrence time of each event in seconds, accurate to milliseconds; The model outputs a description of each event, including detailed information about the event, such as the athlete's name, competition time, and scoring situation; The model outputs an importance score for each event, with the score ranging from 0 to 1. The higher the score, the more important the event. The scoring method uses a comprehensive scoring algorithm based on the event type, time, location, and athlete importance.
[0010] Preferably, in S3.2, the processed video stream information is input into the visual branch of the model, and CNN is used to extract image features. The video stream information is processed through multiple convolutional layers and pooling layers to extract high-level features for each frame. Each convolutional layer contains multiple convolutional kernels, and the size and stride of the convolutional kernels are adjusted according to the video resolution and feature extraction requirements. The pooling layer uses max pooling or average pooling to further reduce the dimension of the feature map. Finally, the extracted image features are flattened in the time and space dimensions to form a one-dimensional feature vector; The processed audio stream information is input into the audio branch of the model, and CNN is used to extract audio features. The audio stream information is processed through multiple convolutional layers and pooling layers to extract high-level features for each audio frame. Similar to the video branch, the settings of the audio convolutional layer and pooling layer are adjusted according to the frequency and time resolution of the audio. Finally, the extracted audio features are flattened in the time and frequency dimensions to form a one-dimensional feature vector; The processed event statistics information and sensor data are input into the text and sensor branches of the model, and the LSTM network is used to extract time series features. The text data is converted into high-dimensional vectors through the word embedding layer and then processed through multiple LSTM networks to extract time series features. The sensor data is directly processed through multiple LSTM networks to extract time series features. Finally, the extracted text and sensor features are flattened in the time dimension to form a one-dimensional feature vector.
[0011] Preferably, in S4, it specifically includes the following steps: S4.1: Extract the video segment, audio segment, and generated text description of the event segment output by S3, and perform event alignment processing; S4.2: Sort the event segments in descending order of importance score to generate a preliminary highlight reel containing multiple important event segments; S4.3: Dynamically adjust the playback speed and duration of the segment based on the emotional intensity and type of the event. For example, a high-importance goal segment is played in slow motion with multi-angle camera switching; a key foul event is inserted with real-time commentary audio to enhance the drama; S4.4: Adaptively adjust the priority of segment selection according to the highlight styles selected by the user (such as "passionate type" and "technical analysis type"). For example, the "passionate type" gives priority to events with high intensity of the audience's cheers, while the "technical type" focuses on the visual presentation of the athlete's sports data (such as speed and trajectory).
[0012] Preferably, in S5, it specifically includes the following steps: S5.1: Establish a similarity matrix of the feature vectors of all event segments in the preliminary highlight, calculate the similarity of each pair of event segments using cosine similarity, and mark the event segments with similarity exceeding the preset threshold as duplicate segments; S5.2: For the events marked as duplicate segments, screen them according to their importance scores, retain the event segment with the highest score, and delete other duplicate segments; S5.3: For event segments that are not completely duplicate but have similar content, use the methods of cutting and merging for processing; S5.4: Reorder the event segments according to the importance of the events and the user's preferences. During the sorting process, consider the time sequence and logical relationship of the events to ensure that the generated highlight conforms to the user's viewing habits; S5.5: Add transition effects between each event segment, such as fade-in / fade-out, overlay effects, and transition animations, to improve the viewing performance of the highlight. The selection of transition effects is optimized according to the type and emotional tendency of the events; S5.6: The rhythm and style of the background music match the emotional tendency of the events.
[0013] Correspondingly, an embodiment of the present invention further provides an artificial intelligence-based sports event highlight generation system for running any one of the artificial intelligence-based sports event highlight generation methods of the embodiments of the present invention, including a video acquisition device, an audio acquisition device, sensors, a statistical system, a data processing module, an intelligent analysis module, an automatic highlight generation module, a user setting module, a video editing module, and a personalized editing module. The video acquisition device, the audio acquisition device, the sensors, and the statistical system are respectively connected to the data processing module, the data processing module is connected to the intelligent analysis module, the intelligent analysis module is connected to the automatic highlight generation module, the automatic highlight generation module is connected to the video editing module, the video editing module is connected to the personalized editing module, and the user setting module is connected to the personalized editing module.
[0014] Preferably, the video acquisition device includes multiple fixed cameras and mobile cameras. The fixed cameras are installed at key positions in the stadium, including both sides of the court and above the spectator stands. The mobile cameras are installed on drones or rail systems and can flexibly adjust the angle and position according to the dynamic changes of the event to record the video information of the sports event. The audio acquisition device includes multiple microphones installed at different positions in the stadium, including the center of the court, the spectator stands, and the commentary booth, for collecting on-site sound information, including the voices of athletes, the cheers of the audience, and the commentary of the event commentator; The sensors include a heart rate monitor, an accelerometer, a GPS positioning module, and a temperature and humidity sensor, which are installed on the clothing, shoes, or equipment of athletes and at various key positions in the stadium to obtain the physiological data, motion data, and environmental data of the athletes. The statistical system includes a handheld device used by referees, a large display screen in the center of the stadium, and a back-end server for obtaining real-time statistical data of the event, including the game time, score, number of fouls, and number of substitutions.
[0015] Advantages of the present invention: 1. Through automated analysis and highlight generation, the complexity and time cost of manual operations are greatly reduced. The application of machine learning models makes event detection and classification more accurate and efficient, avoiding the long process of manual screening.
[0016] 2. By integrating the data of multiple acquisition devices and sensors, the system can provide more accurate, comprehensive, and personalized event highlights on the basis of intelligence, significantly improving the audience experience and meeting the dissemination requirements of sports events.
[0017] 3. Through the combination of this multi-modal data processing and deep learning models, this method can efficiently and intelligently identify and generate sports event highlights that meet the needs of the audience, improving the user experience and optimizing the content production efficiency.
[0018] 4. Through accurate timestamps and event sequencing, the video highlights can provide a smoother viewing experience, reduce the time gap between events, and ensure that the exciting parts of the event are effectively presented.
[0019] 5. Through the fine processing of video, audio, and sensor data by different branches, combining the advantages of LSTM and CNN, the accuracy, efficiency, and user experience of event highlight generation are effectively improved, and ultimately a richer and more personalized highlight viewing experience is provided for the audience. Description of the Drawings
[0020] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only for the present invention to more specifically describe the embodiments, and are not intended to specifically limit the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0021] Figure 1 It is a flowchart of the steps of the method of the present invention; Figure 2 It is a flowchart of the steps of S2 of the method of the present invention; Figure 3 It is a flowchart of the steps of S4 of the method of the present invention. Detailed implementation manners
[0022] The following will describe the present invention in detail with reference to the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; and the drawings are only for more specifically describing the embodiments, and are not intended to specifically limit the present invention.
[0023] Please refer to Figures 1-3 , an embodiment of the present invention provides a method for generating sports event highlights based on artificial intelligence, including the following steps: S1: Use a video capture device to record a sports event, an audio capture device to collect the on-site sound, and at the same time obtain the real-time data of the event through sensors and a statistical system; S2: Adjust the frame rate, reduce noise and compress the video stream information, reduce noise and normalize the volume of the audio stream information, structure the event statistical information, and align the sensor data in time; S3: Establish a detection and classification model based on machine learning, use the video stream information, audio stream information, event statistical information and sensor data processed in S2 as inputs, and output an event list including event type, occurrence time, event description and event importance score; S4: According to the importance of the event and the highlight length set by the user, select the event segments with the highest importance score, and for each selected event segment, generate an event segment including video, audio and text description, and generate a preliminary highlight including multiple important event segments; S5: Edit the preliminary highlight, remove duplicate and redundant segments, adjust the segment order, add transition effects and background music; perform personalized editing according to the highlight parameters set by the user, such as highlight duration, event type, emotional tendency.
[0024] In a possible implementation, first, the system uses multiple video capture devices (such as fixed cameras and mobile cameras) to record the event, ensuring that the panoramic view and critical moments of the game can be captured; at the same time, the audio capture device records the on-site sound in real time, including the cheers of the audience, the conversations of the athletes, and the voices of the event commentators. Through sensors and statistical systems, the system can obtain the physiological data, motion data (such as heart rate, position, acceleration) of the athletes, and the real-time statistical data of the game (such as scores, foul counts, substitution records, etc.). These data provide comprehensive inputs for subsequent analysis and highlight generation.
[0025] Furthermore, in step S2, various types of collected data are preprocessed. The video is processed through frame rate adjustment, noise reduction, and compression to ensure smooth display on different playback devices while ensuring that key event content is not lost. The audio stream is processed through noise reduction and volume normalization to eliminate background noise and ensure stable audio quality. The event statistical information and sensor data are processed through structuring and time alignment to ensure that this information is synchronized with the video and audio streams, facilitating subsequent intelligent analysis.
[0026] Furthermore, in step S3, the system constructs an event detection and classification model based on machine learning. This model can identify specific events from the processed multi-modal data (video, audio, event statistics, and sensor data) and classify them. The model outputs a list of events, which includes the type of the event (such as goal, foul, substitution, etc.), the time when the event occurred, the description of the event, and its importance score. Through this intelligent analysis process, the system can identify which events are the most watchable and worthy of being presented in the highlight.
[0027] Furthermore, in step S4, according to the importance score of each event and the highlight length set by the user, the system selects the event segments with the highest scores from the event list to generate event segments containing video, audio, and text descriptions. These event segments will form the preliminary highlight. Through this step, the system can ensure the high quality of the highlight content and the viewing interest of the audience.
[0028] Finally, the system edits and personalizes the preliminary highlight. By removing duplicate and redundant segments and adjusting the segment order, the highlight content is ensured to be compact, coherent, and non-repetitive. In addition, the system also makes personalized adjustments according to the user's settings (such as highlight duration, event type, and emotional tendency, etc.). The addition of transition effects and background music enhances the visual and auditory effects of the highlight and improves the overall viewing experience of the audience.
[0029] In the embodiment of the present invention, S1 specifically includes the following: S1.1: Use a video capture device to record a sports event, and use an audio capture device to collect on-site sounds, including the voices of athletes, the cheers of the audience, and the commentary of the event commentator. The video capture device includes multiple fixed cameras and mobile cameras. The fixed cameras are installed at key positions on the playing field, including both sides of the court and above the spectator stands. The mobile cameras are installed on drones or rail systems and can flexibly adjust the angle and position according to the dynamic changes of the event. The audio capture device includes multiple microphones, which are installed at different positions on the playing field, including the center of the court, the spectator stands, and the commentary booth. S1.2: Obtain real-time data of the event through a variety of sensors, including the physiological data, motion data, and environmental data of the athletes. The sensors include heart rate monitors, accelerometers, GPS positioning modules, and temperature and humidity sensors. The sensors are installed on the clothing, shoes, or equipment of the athletes, as well as at various key positions on the playing field. Obtain real-time statistical data of the event through a statistical system, including the game time, score, number of fouls, and number of substitutions. The statistical system includes a handheld device used by the referee, a large display screen in the center of the playing field, and a back-end server.
[0030] In a possible implementation, first, a sports event is recorded in its entirety through a video capture device. These devices include multiple fixed cameras and mobile cameras. The fixed cameras are installed at key positions on the playing field, such as both sides of the court and above the spectator stands, to capture the panoramic view and static scenes of the game. The mobile cameras are installed on drones or rail systems and can flexibly adjust the viewing angle and position to respond to the dynamic changes of the game, ensuring that the movement trajectories of the athletes, key tactical combinations, and special events can be captured in real time. This setup can provide multiple perspectives, which helps to accurately judge critical moments of the event later.
[0031] At the same time, the audio capture device, including multiple microphones, is arranged at different positions on the playing field, such as the center of the court, the spectator stands, and the commentary booth. The audio device is set up to ensure that the voices of the athletes, the cheers of the audience, and the event commentator can be captured. By accurately collecting these sound data, not only can the restoration of the event atmosphere be enhanced, but also multi-dimensional audio signals (such as the reaction of the audience and the change of the game rhythm) can be provided for subsequent intelligent analysis.
[0032] Furthermore, sensors are used to monitor the real-time data of athletes, providing physiological and motion state data for highlight generation. These sensors include heart rate monitors, accelerometers, GPS positioning modules, and temperature and humidity sensors. The sensors are usually installed on the athletes' clothing, shoes, or related sports equipment, enabling the real-time acquisition of athletes' physiological data (such as heart rate and breathing rate) and motion data (such as acceleration, speed, and position). These data help evaluate the athletes' physical condition and sports performance, and thus provide important basis for the intelligent algorithm to determine critical moments in the event (such as emergencies and high-intensity exercise moments of athletes).
[0033] In addition, environmental sensors (such as temperature and humidity sensors) are installed at key positions in the stadium to monitor the stadium environment in real time, ensuring the comprehensiveness of data and preventing external environmental factors from affecting the game rhythm and athletes' performance.
[0034] To further enrich the highlight content, the system obtains the real-time statistical data of the event through the statistical system, including game time, score, number of fouls, number of substitutions, etc. These statistical data are recorded and updated in real time through the devices in the referees' hands, the large display screens in the center of the stadium, and the back-end servers. The system can generate timestamps of events based on these data, facilitating the subsequent synchronization of video, audio, and sensor data.
[0035] The close connection of the above steps and the diversification of data sources form a complete event information collection system, ensuring the high quality and accuracy of highlight generation. The combination of video and audio data can provide vivid event scenes and atmospheres, while sensor data can provide accurate physiological response and motion trajectory analysis through monitoring the athletes' states. The real-time statistical data further enhances the timeliness of the highlights, enabling the audience to understand the key changes in the game in a short time.
[0036] This processing of data interconnection and synchronization can greatly improve the accuracy and efficiency of intelligent analysis, helping to accurately identify critical moments and exciting segments in the event. For example, by combining the sudden change in the athlete's heart rate with the reaction of the audience in the audio, the system can determine whether there is a major turning point in the game. At the same time, the synchronization of the real-time score and substitution records can ensure that the content shown in the highlights has a complete time sequence, avoiding the risk of missing important information.
[0037] In the embodiment of the present invention, in S2, the specific steps for frame rate adjustment, noise reduction, and compression processing of the video stream information are as follows: T1: Use video analysis algorithms to detect the original frame rate of the video stream, and dynamically adjust the frame rate of the video stream according to the preset target frame rate and the dynamic characteristics of the video content. The dynamic characteristics include the fast and slow changes of actions. During the frame rate adjustment process, interpolation algorithms are used to fill in the missing frames; T2: Use the time-domain noise reduction algorithm to remove the noise in the time domain through the inter-frame difference between multiple frames. Then, use the spatial-domain noise reduction algorithm to remove the noise in the spatial domain through a spatial filter, where the spatial filter includes a Gaussian filter and a median filter. Finally, use the deblocking algorithm to remove the block effect caused by compression. T3: Use the H.265 encoding standard to encode the video stream. Then, through entropy encoding technology, further reduce the amount of encoded data, where the entropy encoding technology is specifically CABAC. Finally, through the bitrate control algorithm, dynamically adjust the compression parameters according to the network bandwidth and storage space.
[0038] In a possible implementation, first use a video analysis algorithm to detect the original frame rate of the video stream. According to the dynamic characteristics of the video (such as the speed of action, scene changes, etc.), the system dynamically adjusts the frame rate of the video stream. The key here is to optimize the frame rate according to the actual scene content, so as to avoid irrelevant redundant frames and reduce unnecessary data processing. Specifically, when the athletes in the game move slowly or the scene is relatively static, the frame rate can be appropriately reduced; while in scenes of intense movement or rapid changes, the frame rate is increased to capture more details. During the frame rate adjustment process, the interpolation algorithm fills in the frames lost due to the reduction or adjustment of the frame rate, ensuring the coherence of the video stream and avoiding problems such as frame skipping or unevenness caused by frame loss. This process not only improves the dynamic expressiveness of the video but also effectively reduces the amount of video data.
[0039] Furthermore, noise reduction is to improve the clarity and viewing experience of the video image. In step T2, the time-domain noise reduction algorithm is first applied to remove the noise in the time domain through the inter-frame difference between multiple frames. The inter-frame difference can effectively distinguish the motion changes and noise in the video, so as to remove the unnecessary noise that interferes with the picture while maintaining the important details. Next, use the spatial-domain noise reduction algorithm to process the spatial noise in the image. Spatial filters such as Gaussian filters and median filters can smooth the image, reduce graininess, and make the picture clearer and more natural. Finally, to remove the block effect (such as mosaic or image block distortion) caused by video compression, the deblocking algorithm is applied to eliminate the image quality problems caused by low-bitrate compression and ensure that the video can still maintain a high picture quality after compression. This series of noise reduction processing steps can greatly improve the visual effect of the video highlights and avoid noise and distortion from affecting the viewing experience.
[0040] Finally, in the compression stage, the H.265 coding standard is used to encode the video stream. Compared with the old H.264 coding standard, H.265 has higher compression efficiency, which can significantly reduce the size of the video file under the same quality, adapting to various network environments and storage requirements. Subsequently, through entropy coding technology, the amount of encoded data is further reduced. CABAC (Context-based Adaptive Binary Arithmetic Coding) is an efficient entropy coding method that further compresses data and reduces the bandwidth requirements for storage and transmission by adaptively selecting the coding method. Finally, combined with the bitrate control algorithm, the system dynamically adjusts the compression parameters of the video stream according to changes in network bandwidth and storage space. This process ensures that the video will not freeze or experience delays due to insufficient bandwidth during transmission, and at the same time, as much video quality as possible is retained when storage space is limited.
[0041] The three steps of T1, T2, and T3 are closely connected and jointly act on the optimization and compression of the video stream. The frame rate adjustment in the T1 stage lays the foundation for subsequent noise reduction and encoding processing, ensuring that key frames are not lost in terms of content. The noise reduction and deblocking effect processing in the T2 stage eliminates the degradation of image quality and maintains the clarity and visual effect of the picture. The encoding and compression technology in the T3 stage ultimately ensures the efficient storage and transmission of the video, reducing bandwidth consumption and storage pressure. The effective connection of these three steps enables the generated video highlights to have both high-quality pictures and meet low storage requirements, providing an optimized solution for the production, playback, and sharing of subsequent sports event highlights.
[0042] In the embodiment of the present invention, S3 specifically includes the following steps: S3.1: Establish a detection and classification model based on deep learning, specifically a multi-modal Transformer model; S3.2: The input data of the model includes processed video stream information, audio stream information, event statistics information, and sensor data. The video stream information is input into the visual branch of the model, and a CNN (Convolutional Neural Network) is used to extract image features; the audio stream information is input into the audio branch of the model, and a CNN is used to extract audio features; the event statistics information and sensor data are input into the text and sensor branches of the model, and an LSTM (Long Short-Term Memory) network is used to extract time series features; S3.3: In the middle layer of the model, a multi-modal attention mechanism (Multi-Modal Attention Mechanism) is used to fuse features of different modalities to generate a comprehensive feature vector. The multi-modal attention mechanism selects the most relevant features for fusion by calculating the similarity between features of different modalities; S3.4: At the output layer of the model, use a classifier to perform event detection on the comprehensive feature vector, generating an event list including event type, occurrence time, event description, and event importance score. The classifier uses the Softmax function to output the probability distribution of each event type, and selects the event with the highest probability as the detection result; S3.5: Use a large-scale labeled dataset to train the model to ensure the robustness and accuracy of the model in different competition environments. During the training process, use the Adam optimizer, set the learning rate to 0.001, the batch size to 32, and the number of training epochs to 50. Use the cross-entropy loss function as the loss function to ensure the training effect and convergence speed of the model.
[0043] In a possible implementation, first construct a multi-modal Transformer model through deep learning techniques. The Transformer model has a powerful ability to process sequence data and can process multiple information sources of the input through the self-attention mechanism. In this method, the model not only processes video stream information but also fuses multiple modal information such as audio stream, event statistics information, and sensor data, providing rich context for subsequent event detection.
[0044] Alternatively, the multi-modal Transformer model can be replaced by a multi-modal fusion model based on CLIP (Contrastive Language-Image Pretraining); Furthermore, the input data includes video stream, audio stream, event statistics information, and sensor data. The video stream extracts visual features through a CNN (Convolutional Neural Network), and the audio stream also extracts audio features through a CNN. The event statistics information and sensor data are input into the text and sensor branches of the model, and an LSTM (Long Short-Term Memory) network is used to extract time series features. This can fully explore the connections between various modalities and lay a foundation for subsequent multi-modal fusion and event detection.
[0045] Furthermore, at the middle layer of the model, fuse the features of different modalities through a multi-modal attention mechanism. The core of the multi-modal attention mechanism is to calculate the similarity between the features of different modalities and select the most relevant features for fusion. This step can effectively enhance the selectivity of the information flow, avoid the interference of redundant information, and at the same time strengthen the features related to the current event, enabling the model to more accurately detect important events in complex scenarios.
[0046] Furthermore, at the output layer of the model, a classifier is used to process the comprehensive feature vector to generate information such as event type, occurrence time, event description, and event importance score. The classifier calculates the probability distribution of each event type through the Softmax function and selects the event with the highest probability as the final detection result. This step ensures the accuracy and real-time performance of event detection, especially in dynamic and complex sports events.
[0047] Furthermore, during the training process, a large-scale labeled dataset is used to train the model to ensure its robustness and accuracy in different event environments. The Adam optimizer is adopted during training, with a learning rate of 0.001, a batch size of 32, and 50 training epochs. The cross-entropy loss function is used as the loss function. Through these training settings, the model can fully learn the patterns in the data and achieve fast convergence.
[0048] In the embodiment of the present invention, the event types output by the model in S3 include goal, foul, substitution, and offside, and each event type corresponds to a preset label; The model outputs the occurrence time of each event in seconds, accurate to milliseconds; The model outputs the description of each event, including detailed information about the event, such as the name of the athlete, the game time, and the scoring situation; The model outputs the importance score of each event, with a scoring range of 0 to 1. The higher the score, the more important the event. The scoring calculation method uses a comprehensive scoring algorithm based on event type, time, location, and athlete importance.
[0049] In a possible implementation manner, in the S3 stage, the model classifies key events in the event video through deep learning analysis of multi-modal data. First, the model identifies significant events in the event through the fusion of visual and audio features, such as goals, fouls, substitutions, offsides, etc. In this step, the model uses a classification algorithm (such as Softmax or other deep neural network classifiers) to classify each event into a preset label (such as goal, foul, substitution, offside, etc.). The accurate prediction of these event types is the basis for generating highlights, ensuring that the content shown in the video highlights meets the audience's expectations of the event climax.
[0050] Furthermore, the occurrence time of the event is crucial for the timing and logic of generating video highlights. The model accurately calculates the occurrence moment of each event by analyzing the key frames in the video stream using a timing model (such as LSTM or Transformer). The output time accuracy reaches the second level and is accurate to milliseconds, which means that the generated video highlights can show the accurate time points of each event and can be smoothly connected with other events during playback, avoiding time deviation from affecting the viewing effect of the highlights.
[0051] Furthermore, after the event occurrence time and type are determined, the model needs to output a detailed event description. This description includes athlete information related to the event (such as the name of the goal-scoring athlete), time nodes of the game (such as the time in the first half or the second half), and score situation (such as the change in score after a goal). To generate this part of the information, the model comprehensively utilizes event statistics data (such as player lists, score statistics, game progress) and video analysis results, and matches the detailed information with each event. This process not only enhances the content richness of the highlights but also enables the audience to better understand the background of each event.
[0052] Furthermore, the importance scoring of events is a key decision point in generating video highlights, directly affecting the video editing and display priorities. The model uses a comprehensive scoring algorithm to score the importance of each event based on multiple factors (such as event type, time of event occurrence, location of event occurrence, and influence of relevant athletes, etc.). For example, a goal event is usually considered the most important, while fouls or substitutions may be scored relatively low according to the context and rhythm of the event. This scoring system helps the automated decision-making system screen out the events with the most viewing value by considering the context, importance of the event, and the overall situation of the game.
[0053] In the embodiment of the present invention, in S3.2, the processed video stream information is input into the visual branch of the model, and CNN is used to extract image features. The video stream information is processed through multiple convolutional layers and pooling layers to extract high-level features of each frame. Each convolutional layer contains multiple convolutional kernels, and the size and stride of the convolutional kernels are adjusted according to the video resolution and feature extraction requirements. The pooling layer uses max pooling or average pooling to further reduce the dimension of the feature map. Finally, the extracted image features are flattened in the time and space dimensions to form a one-dimensional feature vector; The processed audio stream information is input into the audio branch of the model, and CNN is used to extract audio features. The audio stream information is processed through multiple convolutional layers and pooling layers to extract high-level features of each audio frame. Similar to the video branch, the settings of the audio convolutional layer and pooling layer are adjusted according to the frequency and time resolution of the audio. Finally, the extracted audio features are flattened in the time and frequency dimensions to form a one-dimensional feature vector; The processed event statistics information and sensor data are input into the text and sensor branch of the model, and the LSTM network is used to extract time series features. The text data is converted into high-dimensional vectors through the word embedding layer and then processed through multiple LSTM networks to extract time series features. The sensor data is directly processed through multiple LSTM networks to extract time series features. Finally, the extracted text and sensor features are flattened in the time dimension to form a one-dimensional feature vector.
[0054] In a possible implementation, after the video stream is input into the model, it first passes through multiple convolutional layers and pooling layers of a convolutional neural network (CNN). Each convolutional layer extracts different image features through multiple convolutional kernels, and the size and stride of the convolutional kernels are adjusted according to the resolution of the video and the requirements of feature extraction. The pooling layer uses max pooling or average pooling to further reduce the dimension of the feature map, compress redundant information, help improve computational efficiency, and avoid overfitting. Finally, the extracted image features are flattened in the temporal and spatial dimensions to form a one-dimensional feature vector. This process helps the model capture key visual information in the video frames, such as the actions of athletes, scene changes, and the occurrence of events, etc.
[0055] Furthermore, the audio stream information is input into the audio branch of the model and processed through multiple convolutional layers and pooling layers. The audio convolutional layer extracts different frequency and temporal features in the audio signal through multiple convolutional kernels, and the pooling layer is adjusted according to the frequency and temporal resolution of the audio to further reduce the dimension of the audio features. Finally, the audio features are flattened in the temporal and frequency dimensions to form a one-dimensional feature vector. The purpose of this step is to capture audio features corresponding to video events, such as on-site sound effects, the reactions of the audience, the voices of commentators, etc. These information are crucial for judging the background and importance of events.
[0056] Furthermore, the event statistics information and sensor data are input into the text and sensor branches. The text data is first transformed into high-dimensional vectors through the word embedding layer and then processed through multiple long short-term memory (LSTM) networks to extract time series features. The role of the LSTM is to capture the temporal information in the text data, especially the temporal changes in event descriptions (such as goals, fouls, substitutions, etc.) in the event. The sensor data is directly processed through multiple LSTM networks to extract time series features, helping the model understand the physiological data of athletes or the changes in the venue during the game (such as speed, acceleration, venue temperature, etc.). These data can provide more background information for the occurrence of events.
[0057] Video, audio, and text / sensor data extract features in different branches respectively. Image and audio features are extracted through CNN, and time series features are extracted through LSTM, finally forming one-dimensional feature vectors. This design of multi-modal feature fusion enables the model to more comprehensively understand and capture key events in the event, rather than relying solely on a single information source, thus improving the accuracy and robustness of event recognition.
[0058] The visual branch extracts image features from the video, the audio branch extracts changes in the audio signal, and the text / sensor branch further provides time series features of the event process and the athlete's state. This enables the system to not only accurately capture the moments of each event but also understand the context and importance of each event. For example, through the cheers of the audience and the emotional changes of the commentator in the audio stream, the model can judge the influence of the goal event and the audience reaction, which plays a key role in sorting the important events in the highlights.
[0059] The feature extraction process of each branch undergoes dimensionality reduction processing by the pooling layer, which effectively reduces the feature dimension, improves the calculation efficiency, and at the same time maintains the effectiveness of the information. By flattening the feature vectors, the model can integrate complex multi-modal data into a unified representation, facilitating subsequent decision-making and highlight generation. In addition, using the LSTM network to process time series features can capture the temporal information before and after the occurrence of events, thereby more accurately extracting highly watchable events.
[0060] In the embodiment of the present invention, in S4, it specifically includes the following steps: S4.1: Extract the video segment, audio segment, and generated text description of the event segment output in S3, and perform event alignment processing; S4.2: Sort the event segments in descending order according to the importance score to generate a preliminary highlight containing multiple important event segments.
[0061] S4.3: Dynamically adjust the playback speed and duration of the segment based on the emotional intensity and type of the event. For example, a high-importance goal segment is played in slow motion with multi-angle camera switching; for a key foul event, real-time commentary audio is inserted to enhance the drama; S4.4: Adaptively adjust the segment selection priority according to the highlight style selected by the user (such as "passionate type" or "technical analysis type"). For example, the "passionate type" gives priority to events with a high intensity of audience cheers, and the "technical type" focuses on the visualization of the athlete's movement data (such as speed and trajectory).
[0062] In a possible implementation manner, in the S4.1 stage, the model first extracts the video segment, audio segment, and generated text description corresponding to each event segment from the S3 stage. The key task in this stage is to perform event alignment processing to ensure the consistency between each modality (video, audio, text) of each event. The specific steps are as follows: First, through the visual features extracted in the S3 stage, the video segment has been captured at the critical moment. The model extracts the corresponding period video segment from the original event video stream according to the event timestamp, ensuring that this segment can accurately reflect the visual features of the event itself.
[0063] Furthermore, for audio segment extraction, similarly, key signals in the audio stream (such as audience reactions, changes in the commentator's tone, venue sounds, etc.) will also correspond to the moments of the video segments to generate audio segments. Audio segment extraction must not only ensure time alignment, but also ensure that the emotions, sound intensity and other information in the audio are relevant to the event content.
[0064] Furthermore, for text description extraction, during the processing of event statistics and sensor data, the model generates a text description of the event, including a description of the specific event and other possible event details. These text descriptions are combined with video and audio clips through alignment processing to ensure the consistency of the text content with the audiovisual content.
[0065] Through the alignment processing in this stage, the system can establish a clear connection between the three modalities of video, audio and text, thereby providing high-quality multimodal data support for the next step of highlight generation.
[0066] In stage S4.2, the system scores the importance of each extracted event segment and sorts the event segments from high to low according to the score, thereby generating a preliminary collection containing multiple important event segments.
[0067] First, the importance of the event is scored. The score of each event clip is based on a variety of factors, including but not limited to the viewing experience of the event, the intensity of the game, goals, key fouls, thrilling moments, etc. The score may be based on the output of the machine learning model, taking into account the features in vision, audio, and text. For example, the audience reaction to the goal event, the emotional changes of the commentator, and the key moments during the game (such as the exciting action that turned the situation around) will increase the score of the event.
[0068] Furthermore, based on the scores, the event clips are sorted from high to low. The sorting criteria include not only the direct importance of the event (such as the goal event), but also the emotional intensity of the event, the plot development in the event, etc. The sorting results ensure that the highlights can highlight the climax of the game and the content that the audience is most interested in.
[0069] Furthermore, based on the sorted event fragments, the model concatenates these fragments into a preliminary collection. The preliminary collection usually contains multiple types of high-scoring events, and ensures that the connection between events is natural and smooth, avoiding incoherent jumps in the plot.
[0070] In the embodiment of the present invention, in S5, the following steps are specifically included: S5.1: Establish a similarity matrix of the feature vectors of all event segments in the preliminary collection, use cosine similarity to calculate the similarity of each pair of event segments, and mark the event segments whose similarity exceeds a preset threshold as duplicate segments; S5.2: For the events marked as repeated segments, screen them according to their importance scores, retain the event segment with the highest score, and delete other repeated segments; S5.3: For event segments that are not completely repeated but have similar content, use the method of cutting and merging for processing; S5.4: Reorder the event segments according to the importance of the events and the user's preference settings. During the sorting process, consider the chronological order and logical relationship of the events to ensure that the generated highlights conform to the user's viewing habits; S5.5: Add transition effects between each event segment, such as fade-in / fade-out, overlay effects, and transition animations, to improve the viewing performance of the highlights. The selection of transition effects is optimized according to the type and emotional tendency of the events; S5.6: The rhythm and style of the background music match the emotional tendency of the events.
[0071] In a possible implementation, first, use the cosine similarity calculation method to calculate the similarity for each pair of event segments in the preliminary highlights. The similarity matrix can reflect the similarity degree of the visual and emotional content between the segments. By setting a preset threshold, the system can screen out the event segments whose similarity exceeds this value and mark them as repeated segments. This process ensures that there is no redundant content in the highlights, avoids presenting the same type of events multiple times, and improves the compactness and diversity of the highlights.
[0072] Once the repeated segments are identified, the system will screen them according to the importance score of each segment. The segments with higher scores will be retained, and other low-score repeated segments will be deleted. Through this screening process, it can be ensured that each segment in the highlights has sufficient viewing value and avoid affecting the user's viewing experience due to excessive repeated content.
[0073] In sports events, some event segments have similar content but are not completely repeated. They may be different perspectives of the same type of event or slightly different action displays. For these segments, the system uses the method of cutting and merging to adjust the transition between the segments, so that it can not only maintain emotional coherence but also avoid long repetitions, ensuring the compactness and richness of the highlights.
[0074] Furthermore, reorder the segments by considering the importance of the events, the user's preferences, and the chronological order and logical relationship between the events. The sorting not only needs to highlight the climax part of the game but also ensure that the content presentation conforms to the user's viewing habits. For example, the tense part of the game may be placed in the first half of the highlights, while the relaxing or celebrating moments can be appropriately postponed. Through this sorting, the highlights can tell the game story more smoothly and enhance the user's sense of immersion.
[0075] Furthermore, to enhance the viewing experience of the highlights, the transition effects become a crucial aspect. The system will add appropriate transition effects between each event clip, such as fade-in / fade-out, overlay effects, and transition animations. These transition effects not only make the transition between clips more natural but also can be optimized according to the emotional tendency of the event. For example, for intense game moments, using fast transition effects can enhance the intensity of the emotion; while for celebration moments, gentle transitions can be used to improve the emotional fluency and comfort.
[0076] Furthermore, in the audio design of the highlights, the rhythm and style of the background music need to match the emotional tendency of the event. This means that for intense moments, fast-paced and exciting music can be selected, while for celebration moments, gentle and cheerful background music can be chosen. The background music can not only enhance the emotional tension of the viewing experience but also help users better perceive the atmosphere of the event.
[0077] Correspondingly, the embodiment of the present invention also provides an artificial intelligence-based sports event highlights generation system for running any one of the artificial intelligence-based sports event highlights generation methods in this embodiment, including a video acquisition device, an audio acquisition device, sensors, a statistics system, a data processing module, an intelligent analysis module, an automatic highlights generation module, a user setting module, a clip module, and a personalized editing module. The video acquisition device, the audio acquisition device, the sensors, and the statistics system are respectively connected to the data processing module. The data processing module is connected to the intelligent analysis module. The intelligent analysis module is connected to the automatic highlights generation module. The automatic highlights generation module is connected to the clip module. The clip module is connected to the personalized editing module. The user setting module is connected to the personalized editing module.
[0078] In the embodiment of the present invention, the video acquisition device includes multiple fixed cameras and mobile cameras. The fixed cameras are installed at key positions on the playing field, including both sides of the court and above the spectator stands. The mobile cameras are installed on drones or track systems and can flexibly adjust the angle and position according to the dynamic changes of the event to record the video information of the sports event. The audio acquisition device includes multiple microphones installed at different positions on the playing field, including the center of the court, the spectator stands, and the commentary booth, for collecting on-site sound information, including the voices of athletes, the cheers of the audience, and the commentary of the sports commentator. The sensors include a heart rate monitor, an accelerometer, a GPS positioning module, and a temperature and humidity sensor, which are installed on the athletes' clothing, shoes, or equipment, as well as at various key positions on the playing field, for obtaining the physiological data, motion data, and environmental data of the athletes. The statistics system includes a handheld device used by the referee, a large display screen in the center of the playing field, and a back-end server, for obtaining the real-time statistical data of the event, including the game time, score, number of fouls, and number of substitutions.
[0079] The present invention covers any alternatives, modifications, equivalent methods and solutions made to the essence and scope of the present invention. For the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention, and those skilled in the art can fully understand the present invention without the description of these details. In addition, well-known methods, processes, procedures, components and circuits are not described in detail to avoid unnecessary confusion to the essence of the present invention.
[0080] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for generating sports event highlights based on artificial intelligence, characterized in that, It includes the following steps: S1: Use a video capture device to record a sports event, an audio capture device to collect the on-site sound, and at the same time obtain the real-time data of the event through sensors and a statistical system; S2: Adjust the frame rate, reduce noise and compress the video stream information, reduce noise and normalize the volume of the audio stream information, structure the event statistical information, and perform time alignment processing on the sensor data; S3: Establish a detection and classification model based on machine learning, use the video stream information, audio stream information, event statistical information and sensor data processed in S2 as inputs, and output an event list including event type, occurrence time, event description and event importance score; The specific steps of S3 are as follows: S3.1: Establish a detection and classification model based on deep learning, specifically a multi-modal Transformer model; S3.2: Extract the image features in the video stream information and the audio features in the audio stream information through CNN, and use the LSTM network to extract the time series features in the event statistical information and sensor data; S3.3: The detection and classification model also includes a multi-modal attention mechanism to fuse features of different modalities; S4: According to the importance of the event and the highlight length set by the user, select the event segments with the highest importance score, and for each selected event segment, generate an event segment containing video, audio and text description, and generate a preliminary highlight containing multiple important event segments; S5: Edit the preliminary highlight, remove duplicate and redundant segments, adjust the segment order, add transition effects and background music; perform personalized editing according to the highlight parameters set by the user, including highlight duration, event type, and emotional tendency.
2. The method for generating a sports event highlights based on artificial intelligence according to claim 1, wherein, Specifically, S1 includes the following: S1.1: Use a video capture device to record a sports event, and use an audio capture device to collect the on-site sound, including the voices of athletes, the cheers of the audience, and the commentary of the event commentator; The video capture device includes multiple fixed cameras and mobile cameras. The fixed cameras are installed at key positions on the playing field, including both sides of the field and above the audience stands. The mobile cameras are installed on drones or track systems and flexibly adjust the angle and position according to the dynamic changes of the event; The audio capture device includes multiple microphones, which are installed at different positions on the playing field, including the center of the field, the audience stands, and the commentary booth; S1.2: Obtain the real-time data of the event through a variety of sensors, including the physiological data, movement data, and environmental data of the athletes; The sensors include heart rate monitors, accelerometers, GPS positioning modules, temperature and humidity sensors, and the sensors are installed on the clothing, shoes or equipment of the athletes, as well as at various key positions on the playing field; Obtain the real-time statistical data of the event through a statistical system, including game time, score, number of fouls, number of substitutions. The statistical system includes a handheld device used by the referee, a large display screen in the center of the playing field, and a back-end server.
3. A method for generating sports event highlights based on artificial intelligence according to claim 1, characterized in that, In S2, the specific steps for adjusting the frame rate, reducing noise and compressing the video stream information are as follows: T1: Detect the original frame rate of the video stream using a video analysis algorithm. Dynamically adjust the frame rate of the video stream according to the preset target frame rate and the dynamic characteristics of the video content. The dynamic characteristics include the fast and slow changes of actions. During the frame rate adjustment process, use an interpolation algorithm to fill in the missing frames. T2: Use a time-domain noise reduction algorithm to remove noise in the time domain through inter-frame differences between multiple frames. Then, use a spatial-domain noise reduction algorithm to remove noise in the spatial domain through a spatial filter. The spatial filter includes a Gaussian filter and a median filter. Finally, use a deblocking algorithm to remove the block effect caused by compression. T3: Use the H.265 coding standard to encode the video stream. Then, through entropy coding technology, further reduce the amount of encoded data. The entropy coding technology is specifically CABAC. Finally, through a bitrate control algorithm, dynamically adjust the compression parameters according to the network bandwidth and storage space.
4. A method for generating highlights of sports events based on artificial intelligence according to claim 1, characterized in that, Specifically, S3 also includes: S3.2: The input data of the model includes the processed video stream information, audio stream information, event statistics information, and sensor data. The video stream information is input into the visual branch of the model, and a CNN is used to extract image features. The audio stream information is input into the audio branch of the model, and a CNN is used to extract audio features. The event statistics information and sensor data are input into the text and sensor branches of the model, and an LSTM network is used to extract time series features. S3.3: In the middle layer of the model, use a multi-modal attention mechanism to fuse features of different modalities to generate a comprehensive feature vector. The multi-modal attention mechanism calculates the similarity between features of different modalities and selects the most relevant features for fusion. S3.4: In the output layer of the model, use a classifier to perform event detection on the comprehensive feature vector, generating an event list including the event type, occurrence time, event description, and event importance score. The classifier uses the Softmax function to output the probability distribution of each event type, and selects the event with the highest probability as the detection result. S3.5: Use a large-scale labeled dataset to train the model to ensure the robustness and accuracy of the model in different event environments. During the training process, use the Adam optimizer, set the learning rate to 0.001, the batch size to 32, and the number of training epochs to 50. Use the cross-entropy loss function as the loss function to ensure the training effect and convergence speed of the model.
5. The method for generating a sports event highlights based on artificial intelligence according to claim 4, wherein The event types output by the model in S3 include goals, fouls, substitutions, and offsides. Each event type corresponds to a preset label. The model outputs the occurrence time of each event in seconds, accurate to milliseconds. The model outputs the description of each event, including the detailed information of the event, including the athlete's name, game time, and scoring situation. The model outputs the importance score of each event. The scoring range is from 0 to 1. The higher the score, the more important the event. The scoring calculation method uses a comprehensive scoring algorithm based on the event type, time, location, and athlete importance.
6. The method for generating a sports event highlights based on artificial intelligence according to claim 5, wherein, In S3.2, the processed video stream information is input into the visual branch of the model. The CNN is used to extract image features. The video stream information is processed through multiple convolutional layers and pooling layers to extract high-level features for each frame. Each convolutional layer contains multiple convolutional kernels, and the size and stride of the convolutional kernels are adjusted according to the video resolution and feature extraction requirements. The pooling layer uses max pooling or average pooling to further reduce the dimension of the feature map. Finally, the extracted image features are flattened in the time and space dimensions to form a one-dimensional feature vector; The processed audio stream information is input into the audio branch of the model. The CNN is used to extract audio features. The audio stream information is processed through multiple convolutional layers and pooling layers to extract high-level features for each audio frame. Similar to the video branch, the settings of the audio convolutional layer and pooling layer are adjusted according to the frequency and time resolution of the audio. Finally, the extracted audio features are flattened in the time and frequency dimensions to form a one-dimensional feature vector; The processed event statistics information and sensor data are input into the text and sensor branches of the model. The LSTM network is used to extract time series features. The text data is converted into a high-dimensional vector through the word embedding layer and then processed through multiple LSTM networks to extract time series features. The sensor data is directly processed through multiple LSTM networks to extract time series features. Finally, the extracted text and sensor features are flattened in the time dimension to form a one-dimensional feature vector.
7. A method for generating sports event highlights based on artificial intelligence according to claim 5, characterized in that In S4, the following steps are specifically included: S4.1: Extract the video segment, audio segment, and generated text description of the event segment output in S3, and perform event alignment processing; S4.2: Sort the event segments in descending order according to the importance score to generate a preliminary highlight containing multiple important event segments; S4.3: Dynamically adjust the playback speed and duration of the segments based on the emotional intensity and type of the event; S4.4: Adaptively adjust the segment selection priority according to the highlight style selected by the user.
8. The method for generating a sports event highlights based on artificial intelligence according to claim 1, wherein, In S5, the following steps are specifically included: S5.1: Establish a similarity matrix for the feature vectors of all event segments in the preliminary highlight. Use cosine similarity to calculate the similarity between each pair of event segments. For event segments with a similarity exceeding the preset threshold, mark them as duplicate segments; S5.2: For the events marked as duplicate segments, screen them according to their importance scores, retain the event segment with the highest score, and delete other duplicate segments; S5.3: For event segments that are not completely duplicate but have similar content, use the methods of cutting and merging for processing; S5.4: Reorder the event segments according to the importance of the events and the user's preference settings. During the sorting process, consider the time order and logical relationship of the events to ensure that the generated highlight conforms to the user's viewing habits; S5.5: Add transition effects between each event segment, including fade-in / fade-out, overlay effects, and transition animations, to improve the viewing performance of the highlight. The selection of the transition effect is optimized according to the type and emotional tendency of the event; S5.6: The rhythm and style of the background music match the emotional tendency of the event.
9. An artificial intelligence-based sports event highlights generation system for running an artificial intelligence-based sports event highlights generation method according to any one of claims 1-8, characterized in that, It includes a video acquisition device, an audio acquisition device, sensors, a statistics system, a data processing module, an intelligent analysis module, an automatic highlights generation module, a user setting module, a clip module, and a personalized editing module. The video acquisition device, the audio acquisition device, the sensors, and the statistics system are respectively connected to the data processing module. The data processing module is connected to the intelligent analysis module. The intelligent analysis module is connected to the automatic highlights generation module. The automatic highlights generation module is connected to the clip module. The clip module is connected to the personalized editing module. The user setting module is connected to the personalized editing module.
10. A system for generating sports event highlights based on artificial intelligence according to claim 9, characterized in that, The video acquisition device includes multiple fixed cameras and mobile cameras. The fixed cameras are installed at key positions in the stadium, including both sides of the court and above the spectator stands. The mobile cameras are installed on drones or track systems and can flexibly adjust the angle and position according to the dynamic changes of the event, and are used to record the video information of the sports event. The audio acquisition device includes multiple microphones, which are installed at different positions in the stadium, including the center of the court, the spectator stands, and the commentary booth, and are used to collect the sound information on the site, including the voices of athletes, the cheers of the audience, and the commentary of the event commentator. The sensors include a heart rate monitor, an accelerometer, a GPS positioning module, and a temperature and humidity sensor, which are installed on the clothing, shoes or equipment of athletes and at various key positions in the stadium, and are used to obtain the physiological data, motion data and environmental data of athletes. The statistics system includes a handheld device used by referees, a large display screen in the center of the stadium, and a back-end server, and is used to obtain the real-time statistical data of the event, including the game time, score, number of fouls, and number of substitutions.
Citation Information
Patent Citations
Basketball game semantic event recognition method based on domain knowledge and deep multi-stage feature
CN108681712A
Automatic collection system and method of competition programs
CN110012348A
Audiovisual event positioning method fusing self-supervised multi-modal features
CN115393968A
Key scene automatic segmentation system and method based on video and audio characteristics
CN117609845A
Football match comprehensive sports performance evaluation method and system
CN118552088A
Cited By
Real-time wonderful picture snapshot method based on artificial intelligence
CN120730176A
An artificial intelligence-based highlight picture real-time capturing method
CN120730176B