Sports video commentary evaluation method and system based on multi-modal large language model

By building a multimodal large language model, obtaining and classifying sports video and text explanation data, and performing semantic label weighted scores, the balance of real-time and depth in sports event commentary is solved, and high-quality commentary content generation and evaluation is achieved.

CN120495958APending Publication Date: 2025-08-15BEIJING QIJI TECHNOLOGY CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510597489.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, it is difficult to achieve a balance between real-time and depth of commentary, manual commentary has a risk of lag, machine commentary lacks detailed description and tactical analysis, multimodal large language models lack the diversity of sports commentary content and text density, and cannot provide detailed information and tactical decomposition.

Method used

By building a multimodal large language model, obtaining the data set of sports video and text explanations, performing semantic classification, training a multimodal large language model, generating a sports video explanation model, and achieving cross-modal feature alignment and evaluation through semantic label weighted scores.

Benefits of technology

It realizes the automated generation and accurate evaluation of sports video commentary, improves the quality of real-time event detection, technical analysis and emotional expression, overcomes the technical bottlenecks of traditional methods in multimodal interaction processing, and provides detailed explanation content and tactical analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495958A_ABST
    Figure CN120495958A_ABST
Patent Text Reader

Abstract

The invention discloses a sports video commentary evaluation method and system based on a multi-mode large language model, and the method comprises the steps: obtaining a data set which comprises a data pair composed of a video segment of sports commentary and text commentary; semantic classification is carried out on the data set, semantic tags are determined, and the semantic tags divide the data into at least one of key event description, technical detail analysis, background information interpretation, tactical analysis, competition condition interpretation and emotion expression; constructing a multi-modal large language model, calling the data set to train the multi-modal large language model, and determining a sports video explanation model; and scoring the sports video explanation model, and determining an evaluation result. According to the method, the performance of the model in a sports explanation task can be more comprehensively reflected through a multi-dimensional evaluation method, and the limitation that fine-grained professional details, time dynamics and human emotion cannot be captured by traditional indexes is overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence and sports video processing technology, and more specifically, to a sports video commentary evaluation method and system based on a multimodal large language model. Background Art

[0002] In the related art, in sports events, platforms are usually equipped with one or more commentators to comment on the game videos. During the commentary process, the commentators are usually required to observe the key content in the video screen and then comment.

[0003] However, manual commentary requires extensive commentary experience and the ability to provide real-time commentary that captures the audience's emotions. Otherwise, it's difficult to capture key content in the video, resulting in delayed commentary and a significant impact on the user viewing experience. Furthermore, machine commentary struggles to capture the in-depth details of events and the interplay between multiple modalities, leading to a lack of specificity and nuance in the commentary and a poor viewing experience. Furthermore, current large-scale multimodal language models lack the diversity and text density of sports commentary, and therefore lack comprehensive coverage of sports video information. Consequently, they fail to achieve the same level of description, analysis, and insight provided by human commentators, providing viewers with detailed information, tactical breakdowns, and highlights. Summary of the Invention

[0004] The purpose of this application is to provide a sports video commentary evaluation method and system based on a multimodal large language model to solve at least one of the above technical problems.

[0005] The present application provides a sports video commentary evaluation method based on a multimodal large language model, comprising: obtaining a data set, the data set comprising data pairs consisting of video clips of sports commentary and textual commentary; semantically classifying the data set to determine semantic labels, wherein the semantic labels classify the data pairs into at least one of the following: description of key events, analysis of technical details, explanation of background information, tactical analysis, explanation of game situations, and emotional expression; constructing a multimodal large language model, calling the data set to train the multimodal large language model, and determining a sports video commentary model; and scoring the sports video commentary model to determine an evaluation result.

[0006] Furthermore, obtaining a data set includes: obtaining sports videos of the same type of events; extracting audio files corresponding to the sports videos, and converting the audio files into text-based commentary texts; editing the sports videos and commentary texts according to timestamps to form data pairs with one-to-one correspondence between video clips and text commentary, and determining them as the data set required for the target sports event.

[0007] Furthermore, the audio file is converted into a commentary text in text form, including: determining the first timestamp of each sentence in the commentary text corresponding to the current video clip; fusing the first timestamp of each sentence to determine the timestamp of the current commentary text; encoding each sentence to determine the character encoding; performing similarity calculation on the character encoding corresponding to each sentence, and based on the similarity result, simultaneously determining the semantic features of the character encoding corresponding to each sentence; determining whether the sentences corresponding to each video clip are merged based on the similarity result and the semantic features, and determining whether they are continuous clips of the same theme; dividing the audio file according to the timestamp of the current commentary text and the continuous clips of the same theme, and determining the commentary text corresponding to the same theme.

[0008] Further topics include fouls, goals and timeouts.

[0009] Furthermore, the sports video and the commentary text are edited according to the timestamps to form a one-to-one corresponding data pair between the video clip and the text commentary, including: fusing the first timestamps of each sentence to determine the timestamp of the current commentary text after fusion; and editing the complete sports video and the complete commentary text according to the fused timestamp to form a one-to-one corresponding data pair between the video clip and the text commentary.

[0010] Furthermore, determining the sports video commentary model includes: processing data pairs consisting of video clips and text commentary through a deformable attention mechanism to obtain the temporal offset between the modalities; performing feature fusion based on the bidirectional attention interaction and the temporal offset to determine the fusion features; training the multimodal large language model based on the fusion features in the data set until the convergence conditions are reached, and then determining the sports video commentary model that outputs text, wherein the text version of the commentary text is converted into audio through the speech synthesis capability.

[0011] Furthermore, key event description is used to characterize the ability to accurately detect and describe events, reflecting the proficiency in real-time event detection and textual representation; technical detail analysis is used to characterize the ability to analyze and interpret athlete movements, reflecting the ability of fine-grained visual understanding and contextual knowledge integration; background information interpretation is used to integrate visual and textual information about the history, characteristics and performance of athletes and teams, demonstrating the depth of the model in contextual knowledge integration and in-depth understanding of visual information; tactical analysis is used to interpret and explain the team's strategies, formations and tactical adjustments during the game, analyzing and conveying these insights through visual input, highlighting the model's temporal understanding of the dynamics of the game; game situation interpretation is used to interpret the current state of the game through visual clues, predict its possible progress and provide dynamic situational insights; emotion expression is used to detect emotional elements from visual data and effectively express emotions in text to attract the audience.

[0012] Furthermore, the sports video commentary model is scored to determine the evaluation results, including: dividing a test set from the data set; inputting the video clips corresponding to the test set into the sports video commentary model for interpretation to obtain the interpretation results of each test video clip; performing weighted calculation on the interpretation results of each test video clip according to the semantic label to determine the scoring value; and determining the evaluation results of the sports video commentary model based on the comparison of the scoring value with the preset scoring level.

[0013] Furthermore, the present application also includes: if the evaluation result is lower than the first threshold, it is determined that the sports video commentary model evaluation is unqualified, and a warning icon and / or voice alarm is issued to encourage the sports video commentary model to continue training; if the evaluation result is lower than the second threshold, it is determined that the sports video commentary model evaluation is qualified, and the test passing score is displayed, wherein the first threshold is less than the second threshold; if the evaluation result is higher than the second threshold, it is determined that the sports video commentary model evaluation is excellent, and the test excellent score is displayed.

[0014] The present application also provides a sports video commentary evaluation system based on a multimodal large language model, comprising: a data acquisition module for acquiring a data set, the data set comprising data pairs consisting of video clips of sports commentary and textual commentary; a label classification module for semantically classifying the data set and determining semantic labels, the semantic labels dividing the data pairs into at least one of the following: description of key events, analysis of technical details, explanation of background information, tactical analysis, explanation of game situations, and emotional expression; a model construction module for constructing a multimodal large language model, calling the data set to train the multimodal large language model, and determining a sports video commentary model; and a model evaluation module for scoring the sports video commentary model according to the semantic labels and determining an evaluation result.

[0015] From the above, it can be seen that the present application provides a sports video commentary evaluation method and system based on a multimodal large language model. Through video-text pair training, it can simultaneously capture visual actions (such as player movement and shooting trajectory) and text semantics (such as the tactical term "offside trap"), and realize cross-modal reasoning with spatiotemporal alignment; the multimodal model can improve the accuracy of key event recognition tasks compared with the unimodal model (text or video only), especially in high-speed sports scenes. The commentary data pairs are divided into 6 categories of semantic labels, so that the model has domain knowledge perception capabilities. At the same time, joint training enables the model to achieve a double breakthrough in fluency and professionalism.

[0016] In summary, by acquiring and classifying multimodal datasets, training a large language model with time offset and feature fusion capabilities, and scoring based on semantic labels, we have solved the problems of delayed manual commentary, lack of details in machine commentary, and multimodal interaction. The multidimensional evaluation method can more comprehensively reflect the performance of the model in sports commentary tasks, overcoming the limitations of traditional indicators that cannot capture fine-grained professional details, temporal dynamics, and human emotions. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0018] Figure 1 A flowchart of a sports video commentary evaluation method based on a multimodal large language model provided in an embodiment of the present application;

[0019] Figure 2 A schematic diagram of a principle flow chart of a sports video commentary evaluation method based on a multimodal large language model provided in an embodiment of the present application;

[0020] Figure 3 A schematic diagram of the structure of a sports video commentary evaluation system based on a multimodal large language model provided in an embodiment of the present application;

[0021] Figure 4 A schematic structural diagram of an electronic device in one embodiment of the present application;

[0022] Figure 5 Another structural diagram of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0023] The technical solutions in this application will be clearly and completely described below in conjunction with the drawings in this application. Obviously, the described embodiments are only some embodiments of this application, rather than all embodiments. The components of this application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the application provided in the drawings is not intended to limit the scope of the application for which protection is claimed, but merely represents selected embodiments of the application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of this application.

[0024] It should be noted that similar reference numerals and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures. At the same time, in the description of this application, the terms "first" and "second" are used only to distinguish the description and should not be understood as indicating or implying relative importance.

[0025] In the existing technology, sports video commentary and evaluation technology has undergone a development process from single-modal analysis to simple multimodal fusion. Early systems mainly relied on visual feature extraction of video frames, and used pre-trained models to identify key actions and evaluate basic commentary text. With the advancement of technology, some methods have tried to introduce audio waveform analysis to capture on-site sound characteristics, and combined with game metadata to enhance content accuracy. However, these methods still have significant defects in processing multimodal information: the simple splicing of visual and audio features leads to semantic fragmentation, metadata is only superimposed as text labels and cannot achieve deep integration, and the evaluation content often lacks event details or is out of context. Taking the shooting scene in a football match as an example, the existing system may fail to effectively associate the player actions, audience cheers and real-time score data in the video, resulting in a lack of emotional rendering and tactical analysis depth in the evaluation commentary.

[0026] To address these issues, the inventors first observed that traditional multimodal fusion methods are unable to effectively capture cross-modal semantic associations, such as the mapping between a shooting action in a video and the corresponding commentary term "long shot." Analysis revealed that existing feature encoding spaces lack a cross-modal alignment mechanism, making it difficult to form a unified semantic expression of visual, auditory, and textual features. Further research revealed that simple retrieval enhancement methods face a contradiction between computational efficiency and information richness in real-time scenarios. Based on these findings, the inventors proposed constructing a joint feature space to achieve deep cross-modal alignment and optimizing the retrieval process through sparse regularization. At the same time, to address the problem of one-sided evaluation indicators, the inventors innovatively used differences in the distribution of professional terms as a quality assessment criterion to form a closed-loop optimization mechanism.

[0027] Existing sports commentary typically relies on human commentators observing the video footage in real time to provide dynamic commentary. This approach requires extensive experience to accurately capture key content, but manual commentary carries the risk of lag, impacting the audience experience. While machine commentary can avoid manual fatigue, traditional algorithms struggle to interpret the multimodal information interactions within videos, resulting in a lack of detailed descriptions and tactical analysis, failing to meet user demands for real-time and comprehensive commentary.

[0028] To address these issues, the inventors discovered that the core conflict lies in balancing real-time performance and commentary depth. Traditional methods separate visual and textual information, failing to capture spatiotemporal correlations. Analyzing sports commentary scenarios revealed that commentary content possesses structured semantic features, such as the temporal correlation between key events and tactical adjustments. Based on this, they proposed constructing a semantic classification system to guide multimodal fusion and utilizing a large language model to handle cross-modal feature alignment, forming an end-to-end commentary generation and evaluation framework.

[0029] Therefore, this application proposes a sports video commentary evaluation method based on a multimodal large language model. By obtaining a data set containing video clips and corresponding text commentary, the data pairs are classified according to semantic labels into key event descriptions, technical detail analysis, background information explanation, tactical analysis, game situation explanation and emotional expression, and a multimodal large language model is constructed and trained to generate a sports video commentary model. Finally, the model evaluation is achieved through semantic label weighted scoring.

[0030] See also Figure 1 , is a flow chart of a sports video commentary evaluation method based on a multimodal large language model provided in an embodiment of the present application, which is described in detail as follows:

[0031] Step S101, obtaining a data set, the data set including data pairs consisting of video clips and textual descriptions of sports commentary;

[0032] Step S102 , semantically classifying the data set to determine semantic labels, where the semantic labels classify the data pairs into at least one of the following: description of key events, analysis of technical details, explanation of background information, tactical analysis, explanation of game situations, and emotional expression;

[0033] Step S103: constructing a multimodal large language model, calling a data set to train the multimodal large language model, and determining a sports video commentary model;

[0034] Step S104: scoring the sports video commentary model to determine an evaluation result.

[0035] Among them, the data set refers to a structured data set containing video clips and corresponding text commentary. Specifically, live sports event recordings and their synchronized commentary texts can be used, and training samples are formed after timestamp alignment. Semantic labeling refers to a classification system established based on the characteristics of sports commentary. Specifically, natural language processing technology can be used to identify the intent of the commentary text, and the data pairs are divided into six categories of semantic dimensions to guide the model to learn the characteristic expressions of different commentary scenarios. The multimodal large language model refers to a neural network architecture that integrates visual and text feature processing. Specifically, a Transformer-based cross-modal encoder can be used to achieve spatiotemporal alignment of video frame sequences and commentary texts through an attention mechanism. The scoring method refers to a quantitative evaluation mechanism based on semantic classification. Specifically, a weighted average algorithm can be used to assign weight coefficients according to the importance of different semantic labels to calculate a comprehensive score.

[0036] Specifically, the method first constructs a structured dataset and ensures accurate correspondence between video clips and commentary text through timestamp alignment. After semantic classification of the dataset, the constructed multimodal large language model learns the mapping relationship between visual features and text features through joint training. During model training, video frame sequences are passed through a convolutional network to extract spatial features, while text embedding vectors obtain semantic representations through word vector encoding. The cross-modal attention mechanism dynamically establishes associations between visual elements and text vocabulary to generate commentary content that includes tactical analysis and emotional expression. In the evaluation phase, the model performance is verified through a test set, and a weighted score is assigned based on the accuracy and completeness of the semantic labels to form a multi-dimensional evaluation result.

[0037] Compared with existing technologies, traditional machine interpretation methods focus solely on processing features from a single modality, such as analyzing video frames individually or generating independent text interpretations, resulting in information fragmentation between modalities. This solution establishes a multi-dimensional evaluation system through semantic classification, simultaneously optimizing visual understanding and text generation capabilities during model training. Existing technologies lack a dynamic scoring mechanism, making it impossible to quantitatively evaluate the quality of interpretation content. The weighted scoring method designed in this solution can be adaptively adjusted for different interpretation scenarios, effectively improving the objectivity of the evaluation results.

[0038] Through the above technical solution, this application realizes the automatic generation and accurate evaluation of sports video commentary. This method solves the problem of lack of detail and depth in traditional machine commentary through multimodal feature fusion, and uses a semantic classification system to ensure the integrity and diversity of the commentary content. The dynamic scoring mechanism can accurately reflect the performance differences of the model in dimensions such as real-time event detection and technical analysis, providing a clear direction for model optimization. Compared with manual commentary, this solution significantly improves the quality of tactical analysis and emotional expression while maintaining real-time performance, overcoming the technical bottleneck of traditional algorithms in multimodal interaction processing.

[0039] Optionally, in some embodiments, obtaining a data set includes:

[0040] Get sports videos of the same type of events;

[0041] Extract the audio files corresponding to sports videos and convert the audio files into text-based commentary texts;

[0042] Sports videos and commentary texts are edited according to timestamps to form data pairs with one-to-one correspondence between video clips and text commentary, which are determined as the data set required for the target sports event.

[0043] Among them, the same type of events refers to events with the same sports type or competition rules, such as football matches or basketball matches. They can be specifically filtered by event classification labels to ensure the uniformity and pertinence of the data set content.

[0044] The conversion of audio files into textual commentary refers to converting audio information into text through speech recognition technology. This can be achieved by using Alibaba Cloud Intelligent Voice Interaction or Tencent Cloud Voice Recognition API, and processing audio waveform data through acoustic models and language models to generate text.

[0045] The timestamp refers to the alignment mark of the video and audio on the timeline. Specifically, a time series analysis tool can be used to synchronize the video frames and audio waveforms. For example, the FFmpeg tool can be used to extract the time metadata of the audio and video.

[0046] Specifically, the raw audio stream of a sports video is converted into initial text with timestamps through a speech recognition interface, such as the start and end time of each sentence. The video is then segmented based on the timestamps, for example, binding the video clip from the 10th to the 15th second segment to the corresponding time interval of the commentary text to form precisely aligned data pairs. During the editing process, the video decoder and the text time tags are coordinated using a timecode synchronization protocol, such as the SMPTE timecode standard for frame-level alignment.

[0047] Compared with existing technologies, existing methods usually rely on manual annotation of the correspondence between video clips and commentary text, which is time-consuming and has subjective errors. However, this solution uses automated timestamp synchronization technology to achieve precise matching of audio and video data with text commentary, avoiding the problem of clip misalignment caused by manual operation.

[0048] Through the above technical solution, this application achieves automated pairing of sports event videos and commentary text, effectively solving the problems of inefficient traditional manual editing and inaccurate machine commentary, and providing high-quality multimodal alignment data for subsequent model training. For example, in a football game scene, a video clip of a shot can be accurately associated with the commentator's description of the action, ensuring a high degree of consistency between visual information and semantic expression during the model learning process.

[0049] Optionally, in some embodiments, converting the audio file into a text-based commentary text includes:

[0050] Determine the first timestamp of each sentence in the commentary text corresponding to the current video clip;

[0051] The first timestamps of each sentence are merged to determine the timestamp of the current commentary text;

[0052] Encode each statement and determine the character encoding;

[0053] Calculate the similarity of the character codes corresponding to each sentence, and determine the semantic features of the character codes corresponding to each sentence based on the similarity results;

[0054] Determine whether the sentences corresponding to the video clips should be merged based on the similarity results and semantic features, and determine whether they are continuous clips of the same theme;

[0055] The audio file is divided according to the timestamp of the current commentary text and the continuous segments of the same theme, and the commentary text corresponding to the same theme is determined.

[0056] Among them, the first timestamp refers to the start and end time marks corresponding to each sentence in the audio file, which can be implemented by the time alignment function of the speech recognition tool to locate the position of the sentence in the audio. Character encoding refers to the process of converting text sentences into numerical vector representations, which can be implemented by word embedding models or Transformer encoders to capture the semantic information of sentences. Similarity calculation refers to measuring the degree of association between different sentence encoding vectors, which can be implemented by cosine similarity or Euclidean distance algorithms to determine whether sentences belong to the same semantic topic. Continuous segments refer to a set of multiple sentences with the same topic and adjacent in time. It can be implemented by similarity threshold judgment and temporal continuity analysis to merge semantically related sentences.

[0057] Specifically, during the audio-to-text conversion process, a speech recognition tool is first used to extract the timestamp of each sentence. Sentences with adjacent timestamps are then fused together to form a text paragraph that matches the video clip. Simultaneously, an encoding model is used to convert text sentences into vector form, calculate the similarity of adjacent sentences, and combine semantic feature analysis to determine whether multiple sentences need to be merged into a continuous description of the same topic. For example, if multiple sentences describe the same foul, they are merged into a single topic segment. Finally, based on the fused timestamps and topic division results, the audio is segmented into commentary text paragraphs that strictly correspond to the video clips.

[0058] Compared to existing technologies, traditional methods only segment audio into text segments based on fixed time windows, resulting in temporal misalignment or semantic disconnection between the commentary and the video content. This solution dynamically analyzes the semantic similarity and temporal continuity of sentences, combining multiple sentences on the same topic into coherent paragraphs. This ensures that the generated commentary accurately matches the start and end times of key events in the video, while also improving the text's semantic integrity and contextual relevance.

[0059] Through the above technical solution, this application solves the problem of fragmented commentary content caused by mechanical cutting during the audio-to-text conversion process, ensuring that the description of the same event remains coherent at the timestamp and semantic levels, thereby improving the accuracy of subsequent model training and commentary generation. For example, in a football match, consecutive sentences such as "pass", "shoot", and "goal" are merged into the same topic segment to avoid the model misjudging them as independent events due to cutting, and ultimately making the generated commentary text more consistent with the actual progress in the video.

[0060] Optionally, in the above embodiment, the topics include fouls, goals and timeouts.

[0061] This application further proposes a sports video commentary evaluation method based on a multimodal large language model, where topics include fouls, goals and timeouts.

[0062] A foul refers to an athlete's violation of the rules during a match. This can be identified through visual or auditory signals such as physical contact, referee gestures, or whistles in video clips, and used to segment commentary segments involving disputed rules. A goal refers to a scoring event during a match. This can be detected through the visual trajectory of the ball entering the goal frame or keywords such as "score" or "break" in the commentary text, and used to categorize commentary content related to scoring. A pause refers to an interruption in the game due to a substitution, injury, or tactical adjustment. This can be identified through visual features such as a timer stop, players gathering, or coach gestures, and used to segment commentary segments that indicate changes in the game's tempo. These topic classifications provide clear event type anchors for the semantic label segmentation, thereby enhancing the model's understanding of the semantic associations between commentary text and video clips.

[0063] Specifically, in the process of clipping sports videos and commentary texts into data pairs according to timestamps, topic classification ensures that the same continuous segment only contains descriptions of a single event type by limiting the semantic boundaries of the merge process. For example, when keywords such as "yellow card" and "foul" are detected in the commentary text, and the video clip shows the referee showing a yellow card, the data pair will be labeled as the "foul" topic to avoid confusion with "goal" or "timeout" events. This classification mechanism enables the model to construct differentiated feature extraction paths for different topics during training. For example, for foul events, it focuses on learning the relationship between referee gestures and rule provisions, while for goal events, it strengthens the mapping relationship between ball trajectory and scoring terms.

[0064] Compared to existing technologies, traditional methods typically only mechanically align video and text in chronological order, failing to consider the semantic differences between event types. For example, when a commentator continuously describes a foul and its penalty, existing technologies may split it into multiple independent segments, preventing the model from fully understanding the causal relationship between the events. However, by limiting the core topics to three categories: fouls, goals, and timeouts, this solution further establishes event semantic units based on timestamp alignment, allowing the model to maintain the integrity and logical coherence of the event description when generating commentary.

[0065] Through the above technical solution, this application solves the semantic confusion problem caused by the intersection of multiple events in sports commentary and improves the model's recognition accuracy for key game nodes. For example, in a football match, when the video simultaneously shows a player falling to the ground and the referee blowing the whistle, the topic classification mechanism can accurately distinguish whether this clip represents a foul call rather than a timeout, thereby generating commentary content that includes the basis for the penalty and possible consequences, avoiding the incorrect description of "a player's injury caused the game to be suspended."

[0066] Optionally, in some embodiments, the sports video and the commentary text are clipped according to the timestamp to form a data pair with a one-to-one correspondence between the video clip and the text commentary, including:

[0067] Fusing the first timestamps of each sentence to determine the timestamp of the current narration text after fusion;

[0068] The complete sports video and the complete commentary text are edited according to the fused timestamps to form a one-to-one corresponding data pair between the video clip and the text commentary.

[0069] The first timestamp refers to the start and end time of each sentence in the original commentary text converted from the audio file. This can be achieved by using speech recognition technology combined with timeline marking, for example, by using speech endpoint detection to determine sentence boundaries and record time information. The fused timestamp refers to the continuous time period formed by merging the time intervals of multiple related sentences. This can be achieved by using semantic similarity analysis combined with time overlap detection, for example, by using natural language processing technology to identify consecutive sentences on the same topic and merge their time intervals.

[0070] Specifically, the audio of sports videos is first converted into time-stamped text commentary using speech recognition technology, with each sentence corresponding to an independent start and end time. Then, a text similarity calculation method, such as cosine similarity based on word vectors or semantic embeddings, is used to determine whether consecutive sentences belong to the same topic. When the similarity of adjacent sentences exceeds a set threshold and the time intervals overlap or are adjacent, the timestamps of multiple sentences are merged into a unified time period. Finally, the original sports video is precisely edited based on the merged time period, so that the generated video clips are completely aligned with the merged commentary text in the time dimension, forming a one-to-one corresponding training data pair.

[0071] Compared with existing technologies, traditional video editing methods often directly use the original sentence timestamps to segment the clips, resulting in the same event being cut into multiple incoherent segments. However, this application uses timestamp fusion processing to merge multiple commentary sentences around the same topic into a complete event description, achieving precise matching of video clips and commentary text in the time dimension, and solving the data alignment problem caused by scattered timestamps.

[0072] Through the above technical solution, this application effectively improves the temporal consistency and semantic integrity of training data pairs, ensuring that the multimodal large language model can accurately learn the dynamic relationship between video content and commentary text during training. This precise alignment of data construction significantly enhances the model's ability to understand the sequential events in sports events, laying the data foundation for generating coherent and contextually relevant commentary content.

[0073] Furthermore, the sports video commentary model is determined to include:

[0074] The deformable attention mechanism is used to process data pairs consisting of video clips and text descriptions to obtain the temporal offset between modalities.

[0075] Perform feature fusion based on bidirectional attention interaction and temporal offset to determine fusion features;

[0076] The multimodal large language model is trained based on the fusion features in the dataset until convergence conditions are reached, and then a sports video commentary model that outputs text is determined, wherein the text version of the commentary text is converted into audio output through speech synthesis capabilities.

[0077] Among them, the deformable attention mechanism refers to an attention calculation method that can dynamically adjust the focus area. It can be implemented using dynamic convolution kernels or deformable sampling grids. By capturing the dynamically changing spatiotemporal relationship between video and text, it solves the problem of temporal mismatch between modalities.

[0078] Among them, bidirectional attention interaction refers to the forward and reverse feature transfer of cross-modal information in the time dimension. It can be implemented using a cross-attention layer. By establishing a bidirectional association between video frames and text terms, the alignment capability of fine-grained features between modalities is enhanced.

[0079] Among them, timing offset refers to the asymmetric correspondence between the video frame sequence and the text word sequence on the time axis. It can be modeled through optical flow estimation or dynamic time warping algorithm to correct the difference in time steps between modalities.

[0080] Among them, speech synthesis capability refers to converting the generated text commentary into speech with natural rhythm. Specifically, it can be achieved using an end-to-end speech synthesis model, and outputting an audio stream that conforms to the commentary scenario through the acoustic feature prediction and waveform generation module.

[0081] Specifically, during the model training phase, video clips and corresponding textual descriptions are first input into the deformable attention module, which dynamically generates attention weights by analyzing the spatiotemporal relationship between video frames and textual terms, and automatically detects and compensates for the temporal offset between the two. Next, the bidirectional attention interaction module performs cross-modal feature fusion in the temporally aligned feature space, establishing contextual associations between video features and textual features through a cross-attention mechanism to form fused features containing multimodal semantic information. During training, the model is optimized by minimizing the difference between the generated text and the annotated description, and the model converges when the training loss stabilizes. The resulting sports video commentary model is able to automatically generate semantically logical textual descriptions based on the input video, and then convert the text into audio output with natural pauses and intonation through a speech synthesis module.

[0082] Compared to existing technologies, traditional multimodal models typically use a fixed-window attention mechanism, which cannot effectively handle the dynamic offset between video and text on the timeline. This results in overlooking the precise alignment of keyframes and corresponding terms during feature fusion. This solution, however, actively models temporal offsets through a deformable attention mechanism and combines it with bidirectional interaction to enhance cross-modal associations. This allows for more accurate capture of fine-grained correspondences, such as between a shot in a soccer match and the commentary of "wonderful volley," avoiding the synchronization issues between commentary and the video caused by temporal misalignment in traditional methods.

[0083] Through the above technical solution, this application achieves precise temporal alignment between video content and commentary text, effectively improving the quality of cross-modal feature fusion and enabling the generated commentary content to maintain a high degree of synchronization with the video when describing key events. At the same time, the audio commentary output using speech synthesis technology has a more natural rhythm and emotional expression, resolving the problems of stiff text and monotonous intonation in traditional machine commentary, thereby significantly enhancing the user's on-site experience when watching sports events.

[0084] Optionally, in some embodiments, key event description is used to characterize the ability to accurately detect and describe events, reflecting the proficiency in real-time event detection and textual representation; technical detail analysis is used to characterize the ability to analyze and interpret athlete movements, reflecting the ability of fine-grained visual understanding and contextual knowledge integration; background information interpretation is used to integrate visual and textual information about the history, characteristics and performance of athletes and teams, demonstrating the depth of the model in contextual knowledge integration and in-depth understanding of visual information; tactical analysis is used to interpret and explain the team's strategies, formations and tactical adjustments during the game, analyzing and conveying these insights through visual input, highlighting the model's temporal understanding of the dynamics of the game; game situation interpretation is used to interpret the current state of the game through visual clues, predict its possible progression and provide dynamic situational insights; emotion expression is used to detect emotional elements from visual data and effectively express emotions in text to attract the audience.

[0085] This application further proposes that in the process of sports video commentary evaluation, data pairs are divided into six categories of semantic labels: key event description, technical detail analysis, background information interpretation, tactical analysis, game situation interpretation and emotional expression, so as to improve the model's commentary ability through multi-dimensional evaluation.

[0086] Among them, key event description refers to the real-time capture of key actions or nodes in the video through event detection algorithms, such as using inter-frame difference method combined with target detection technology to identify the moment of goal or foul, in order to evaluate the accuracy of the model in real-time event capture and text generation.

[0087] Technical detail analysis refers to the use of fine-grained visual recognition technology to analyze the athlete's movement trajectory, such as extracting joint angle data through posture estimation algorithms, combining it with sports mechanics knowledge to generate technical explanations, reflecting the model's ability to understand micro-movements and integrate knowledge.

[0088] Contextual information interpretation refers to associating historical game data and real-time footage through multimodal knowledge graphs. For example, after identifying a specific player in a video clip, its career data can be automatically retrieved to evaluate the model's performance in contextual information fusion and deep reasoning.

[0089] Tactical analysis refers to tracking changes in player formations based on time-series modeling algorithms, such as using graph convolutional networks to analyze offensive and defensive positional relationships, generate tactical adjustment suggestions, and verify the model's level of dynamic scene understanding and strategy interpretation.

[0090] Match interpretation refers to the use of probability prediction models to calculate the direction of the game in combination with real-time scores and remaining time. For example, the Markov chain is used to predict the probability of winning or losing, and the reliability of the model is tested in situation analysis and forward-looking commentary.

[0091] Emotional expression refers to the use of emotion recognition models to extract features of audience cheers or athletes' expressions, such as generating contagious commentary through voiceprint analysis and facial emotion recognition, to test the effectiveness of the model in emotional transmission and audience interaction.

[0092] Specifically, during the model training phase, each semantic label corresponds to an independent quality assessment module. For example, in the key event description module, the model needs to locate the event trigger point in the video frame and generate synchronized commentary text. The system calculates the score by comparing the time difference between event detection and the coverage of the commentary text. In the tactical analysis module, the model needs to identify the formation change nodes in the video and output the tactical name and adjustment suggestions. The system scores based on the accuracy of tactical classification and the rationality of the strategy. The weighted scoring results of the six categories of labels together constitute a comprehensive evaluation index, ensuring that the model achieves balanced improvement in the six core capabilities of event capture, detail analysis, background integration, strategy interpretation, situation judgment, and emotional transmission.

[0093] Compared to existing technologies, traditional machine commentary relies solely on matching video content with fixed templates. For example, it mechanically outputs basic phrases like "goal" when a goal is scored, lacking tactical analysis and emotional expression. This solution establishes a hierarchical evaluation system using six categories of semantic tags. For example, during the testing phase, it can independently test whether the model can simultaneously explain the referee's decision when identifying offside, or generate context for red cards based on a player's historical foul data. This achieves comprehensive breakthroughs in the accuracy, depth, and appeal of commentary content.

[0094] Through the above-mentioned technical solutions, this application effectively addresses the issues of monotonous content and lack of multimodal interaction in machine commentary. For example, in a scenario analyzing technical details, the model can accurately explain the swing angle and force principle of a football player's shot, avoiding the traditional method of simply describing a "powerful shot" in a general sense. In a scenario expressing emotions, the model can automatically adjust the passion of the commentary based on the level of the audience's cheers, significantly enhancing the audience's immersive experience compared to existing synthesized speech with fixed emotional intonation.

[0095] Furthermore, the sports video commentary model is scored to determine the evaluation results, including:

[0096] Split the test set from the dataset;

[0097] Input the video clips corresponding to the test set into the sports video commentary model for interpretation, and obtain the interpretation results of each test video clip;

[0098] The explanation results of each test video clip are weighted according to the semantic tags to determine the score;

[0099] The evaluation results of the sports video commentary model are determined by comparing the score value with the preset score level.

[0100] Among them, test set partitioning refers to separating an independent subset from the original dataset for verifying model performance. This can be achieved by random sampling or stratified sampling. This partitioning ensures that the evaluation process is not interfered with by the training data and can objectively reflect the model's generalization ability on unknown data.

[0101] Among them, weighted calculation refers to assigning corresponding weights to the commentary results according to the differences in the importance of different semantic tags. Specifically, the hierarchical analysis method or expert scoring method can be used to determine the weight coefficient. This calculation method can perform differentiated evaluations based on the multi-dimensional characteristics of sports commentary, avoiding evaluation bias caused by a single indicator.

[0102] Among them, the preset scoring level refers to the pre-set model performance evaluation standard range, which can be implemented using a percentage system or a five-level classification system. This level provides a clear performance reference benchmark for model optimization, making the evaluation results explainable and operational.

[0103] Specifically, a test set is first divided proportionally from the trained dataset, for example, the test set accounts for 20% of the total data volume. The video clips in the test set are input into the trained sports video commentary model, and the model generates corresponding text commentary results based on the video content. The generated commentary text is then compared with the original standard text commentary of the test set, and the matching score under each label is calculated according to the semantic label category. For example, the weight of the key event description label is set to 0.3, and the weight of the technical details analysis label is set to 0.2. The comprehensive score value is obtained by weighted summation. Finally, the score value is compared with the preset three level intervals, for example, less than 60 points is unqualified, 60-80 points are qualified, and above 80 points are excellent, thereby outputting a qualitative evaluation conclusion of the model performance.

[0104] Compared to existing technologies, traditional methods rely on subjective human evaluation of model commentary quality, which is inefficient and inconsistent. This solution, through an automated weighted scoring mechanism, can simultaneously process large amounts of test data while ensuring consistent evaluation criteria. Compared to holistic scoring methods that don't distinguish between semantic dimensions, this solution quantifies multiple attributes of sports commentary in a layered manner, significantly improving the granularity and accuracy of the evaluation results.

[0105] Through the above technical solution, this application realizes the objective quantitative evaluation of the performance of sports video commentary models, solving the technical defects of low efficiency of manual evaluation and single dimension of machine evaluation. This solution accurately identifies the differences in the model's capabilities in different commentary dimensions through a semantic label weighting mechanism. For example, it can detect the model's weaknesses in tactical analysis, providing a clear direction for subsequent model optimization. At the same time, the preset scoring level provides an operational decision-making basis for model performance judgment. For example, when the score value falls below the first threshold, the retraining mechanism is automatically triggered, effectively improving the efficiency of model iteration.

[0106] Furthermore, this application also includes:

[0107] If the evaluation result is lower than the first threshold, it is determined that the sports video commentary model fails the evaluation, and a warning icon and / or voice alarm is issued to prompt the sports video commentary model to continue training;

[0108] If the evaluation result is lower than the second threshold, it is determined that the sports video commentary model has passed the evaluation and the test result is displayed, wherein the first threshold is lower than the second threshold;

[0109] If the evaluation result is higher than the second threshold, it is determined that the sports video commentary model is evaluated as excellent, and the excellent test result is displayed.

[0110] Among them, the first threshold refers to the pre-set minimum passing standard line, which can be implemented specifically by a dynamically adjusted percentage value. For example, a score lower than 60% of the total score is used as the first threshold, and its function is to quickly identify the deficiencies of the model in key semantic classification items. The second threshold refers to an advanced standard line higher than the first threshold, for example, it is set to 80% of the total score, which is used to divide the passing and excellent levels, and form a gradient evaluation system through a dual threshold mechanism. Warning icons or voice alarms refer to triggering methods that provide feedback through visual or audible signals, such as popping up a red warning sign or playing a prompt sound in the model training interface, and its function is to remind developers in real time to pay attention to the substandard status of the model. Displaying test scores means outputting the scoring results in text or graphical form, such as generating a bar chart containing the score proportions of each semantic label, which is used to intuitively present the performance distribution of the model in different evaluation dimensions.

[0111] Specifically, during the model training process, the test set video clips are processed by the model to generate explanatory text, which is weighted and scored according to the semantic label classification. When the score value is lower than the first threshold, the system automatically triggers a warning signal and forces the start of additional training processes, such as reloading the substandard semantic category data for targeted optimization. If the score is between the first and second thresholds, the system only outputs a qualified mark and records the current model parameter version, allowing developers to choose whether to continue optimization according to their needs. When the score exceeds the second threshold, the system generates an excellent mark and opens the model deployment permission, while providing traceable score details for horizontal comparative analysis. Through a hierarchical feedback mechanism, the model training process can dynamically adjust the optimization strategy according to the performance of different stages.

[0112] In some specific embodiments, the first threshold can be dynamically calibrated based on the average failure rate of historical training data. For example, in a football commentary model, if the tactical analysis tag scores below 50 points for three consecutive tests, the first threshold is automatically lowered from 60% to 55%. Voice alerts can use pre-recorded prompt audio or real-time synthesized voice content, such as playing "The model does not meet the requirements in the emotional expression dimension. Please check the coverage of the training data." The display interface for outstanding scores can integrate a multi-dimensional radar chart, while also noting the score fluctuation range of each semantic tag.

[0113] Compared to existing technologies, existing model evaluation methods typically use a single threshold judgment mechanism, distinguishing only between qualified and unqualified states and lacking gradient feedback for intermediate states. For example, traditional methods only generate error logs when the model score falls below the preset standard and fail to proactively trigger retraining, resulting in inefficient optimization. This solution, however, uses dual thresholds to divide the model into three processing modes, combining proactive alerts with visual feedback. This closes the loop between model performance evaluation and training adjustments, effectively shortening the iteration cycle.

[0114] Through the above technical solution, this application can dynamically adjust the model training strategy according to the scoring results to avoid the optimization direction deviation caused by a single evaluation standard. By triggering warnings, qualified labels and excellent labels in a hierarchical manner, developers can quickly locate the weak links of the model in key semantic tags. For example, when the score of the technical detail analysis dimension is too low, it is prioritized to supplement the athlete's movement decomposition data. At the same time, the multi-level feedback mechanism improves the degree of automation of the model training process, reduces the frequency of manual intervention, and ensures that the commentary model achieves balanced development in multimodal semantic understanding capabilities.

[0115] See also Figure 2 , is a schematic diagram of a principle flow chart of a sports video commentary evaluation method based on a multimodal large language model provided in an embodiment of the present application, which is described in detail as follows:

[0116] Data Collection:

[0117] First, we collected a large number of high-definition sports videos from different online sports events and activities, and included original English commentary. These videos and commentary are the basic resources for building the dataset.

[0118] Audio Extraction:

[0119] Extract English commentary audio from complete sports game videos and retain their corresponding timestamps. This step separates the video and audio information, laying the foundation for subsequent commentary text processing.

[0120] Textual Merge:

[0121] Since the extracted commentary sentences are fragmented, they are integrated into a continuous commentary text through the merging process. Specifically, if the previous and current commentaries are Com1 and Com2, and their timestamps are [b1, e1] and [b2, e2], respectively, the merged commentary and timestamps can be obtained by a specific formula:

[0122] New Com=Com1+Com2

[0123] New Time Stamp=[b1,e2]

[0124] Merging decisions involve two stages: text merging and GPT merging. The text merging stage uses a sentence encoder and cosine similarity to merge captions with high textual similarity, such as repeated cheers or consecutive sentences describing the same content. The GPT merging stage uses GPT-4o-mini to determine whether caption segments with consistent semantics should be merged. This is applicable to consecutive segments describing the same topic.

[0125] Video Slicing:

[0126] Before editing, the timestamps are manually optimized. Then, the complete sports game video is cut into video clip-commentary pairs according to the timestamps.

[0127] Six-Dimensional Label Generation:

[0128] Finally, GPT-4o-min is used to classify the semantic content of the explanations. Each explanation sentence is classified into one or more of the six dimensions mentioned in the metrics section.

[0129] Dataset Split:

[0130] The dataset is divided into training and testing sets to facilitate model training and evaluation.

[0131] Statistical Analysis & Comparison of Datasets:

[0132] Detailed statistics are presented for CommentarySet and compared with other existing sports video captioning datasets to demonstrate its advantages in diversity, commentary style, and data coverage. The six sports included in CommentarySet are carefully selected to ensure strong diversity. The six selected sports show significant diversity in game speed, player density, commentary length, and commentary frequency. For example, table tennis has the shortest average clip length of only 9.96 seconds, while track and field, gymnastics, and tennis clips are 1.5 to 2 times longer than table tennis, reflecting the faster game pace and shorter commentary content. In terms of commentary frequency, basketball has the highest frequency, while football and tennis have relatively low frequencies, which is closely related to the frequency of events occurring on the field.

[0133] Commentary style is another key aspect. The content and characteristics of commentary vary across different sports, which prompted the proposal of a six-dimensional index. By analyzing the distribution of the six labels in different sports, such as Figure 3 As shown, different commentary styles are observed. For example, in basketball, key event captions account for 44.33%, highlighting the successive on-court events; in gymnastics, 61.84% of commentaries include background information, as the commentator frequently introduces each athlete. The diversity of sports leads to a diversity of commentary styles, which is quantitatively reflected in the label distributions of different sports. CommentarySet contains commentary style data from different sports, providing a foundational reference for commentary generation tasks, which is a unique statistical concept compared to other datasets.

[0134] Existing benchmarks for Video Large Language Models (Video LLMs) primarily rely on evaluation metrics designed for question answering (QA) or multiple-choice questions, which are not well-suited for sports commentary tasks. Traditional metrics such as BLEU, ROUGE_L, METEOR, CIDEr, and SPICE primarily evaluate the lexical or structural similarity between generated and reference texts, focusing on aspects such as n-gram overlap, sequence matching, and semantic accuracy. However, these metrics do not capture the model's ability to understand fine-grained professional details, temporal dynamics, or human emotions in the commentary. Furthermore, compared to traditional GPT-based methods, the latter evaluates commentary directly based on ground truth, which not only provides a more detailed and fine-grained scoring criterion for GPT, but also allows for a more intuitive demonstration of the model's capabilities across different dimensions.

[0135] In order to better evaluate the multifaceted nature of sports commentary, a new six-dimensional index is proposed, including the following aspects:

[0136] Key Events Caption:

[0137] This dimension evaluates the model's ability to accurately detect and describe important events (such as scores, turnovers, fouls, and substitutions), reflecting the model's proficiency in real-time event detection and text representation.

[0138] Technical Detail Analysis:

[0139] This dimension evaluates the model's ability to analyze and interpret player actions (such as passing and shooting techniques), reflecting the model's capabilities in fine-grained visual understanding and integration of contextual knowledge.

[0140] Background Information Interpretation:

[0141] This dimension measures the model's ability to integrate visual and textual information about the history, characteristics, and performance of players and teams, demonstrating the depth of the model in integrating contextual knowledge and deeply understanding visual information (such as game intensity and event frequency).

[0142] Tactical Analysis:

[0143] This dimension assesses the model’s skill in deciphering and interpreting team strategy, formations, and tactical adjustments during matches, analyzing and communicating these insights through visual input, and highlighting the model’s temporal understanding of match dynamics.

[0144] Match Situation Interpretation:

[0145] This dimension assesses the model’s ability to interpret the current state of the game (e.g., score and momentum) through visual cues, predict its likely progression, and provide dynamic situational insights.

[0146] Emotional Expression:

[0147] This dimension measures the model’s ability to detect emotional elements (such as athlete reactions and audience energy) from visual data and effectively express these emotions in text to engage the audience.

[0148] In this embodiment, by proposing a method for constructing a fine-grained table tennis narrative (FTTN) dataset and an innovative sports video narrative evaluation method (SVNM), not only is the multimodal large model's in-depth understanding of complex sports event video content greatly enhanced, but also the quality and accuracy of intelligent commentary are significantly improved by comprehensively considering key dimensions such as relevance, consistency, fluency, and logic. In addition, the introduced KL divergence evaluation method provides a new perspective for measuring the difference between model output and human professional commentary, further ensuring the humanization and acceptability of the commentary content. The integration and application of these technologies not only optimizes the automated process of sports video subtitle generation, but also brings breakthrough progress to the field of intelligent sports commentary, providing viewers with a richer and more vivid viewing experience, while also laying a solid foundation for future research and application.

[0149] Experiments are conducted to verify the superiority of the newly proposed six-dimensional metric over traditional metrics. Traditional metrics such as BLEU, ROUGE_L, METEOR, CIDEr, and SPICE primarily evaluate the lexical or structural similarity between generated text and reference text. These metrics primarily focus on aspects such as n-gram overlap, sequence matching, and semantic accuracy when evaluating model performance. However, these metrics do not fully capture the model's ability to understand fine-grained professional details, temporal dynamics, or human emotions in sports commentary tasks. Therefore, a new six-dimensional metric is proposed, which includes key event descriptions, technical detail analysis, background information interpretation, tactical analysis, game situation interpretation, and emotional expression to more comprehensively evaluate the performance of sports video commentary.

[0150] Experimental results show that traditional metrics have low scores on all models (BLEU scores below 1 and CIDEr scores around 0.1), making it difficult to distinguish the performance of different models. In contrast, the six-dimensional metric provides a more nuanced evaluation and effectively distinguishes the capabilities of the models. For example, although InterVL-Chat-2 has the lowest BLEU score and a relatively low CIDEr score, the six-dimensional metric highlights its superior performance across all sports categories.

[0151] In order to better verify the advantages of the six-dimensional index, a total of 36 human evaluators were recruited to conduct user research. The evaluators were shown the ground truth explanation and the explanations generated by two different models (Explanation 1 and Explanation 2). The evaluators gave a choice for each sample, where A represents the preference for explanation 1, B represents the two are equal, and C represents the preference for explanation 2. The most frequent choice in each sample was taken as the final result. Similarly, for other indicators, the results were also classified into three choices: A, B and C. By comparing the scores,

[0152] Comparing human-generated results with the six-dimensional metric and three previous metrics (the original LLM, BLEU, and CIDEr), the results show a 60% overlap between the metric and human decisions, approximately 1.5 times higher than the previous best method, the original LLM. Furthermore, the original LLM's results closely overlap with CIDEr, suggesting that the LLM itself does not outperform traditional NLP metrics. It is the improvements brought to LLMs by the six-dimensional metric that enhance their performance in evaluating commentary.

[0153] See also Figure 3 , a sports video commentary evaluation system based on a multimodal large language model provided by this application, comprising:

[0154] A data acquisition module 301 is used to acquire a data set, the data set including data pairs consisting of video clips and textual descriptions of sports commentary;

[0155] The label classification module 302 is used to semantically classify the data set and determine semantic labels. The semantic labels classify the data pairs into at least one of the following: key event description, technical detail analysis, background information explanation, tactical analysis, game situation explanation and emotional expression;

[0156] The model construction module 303 is used to construct a multimodal large language model, call a data set to train the multimodal large language model, and determine a sports video commentary model;

[0157] The model evaluation module 304 is used to score the sports video commentary model according to the semantic tags and determine the evaluation result.

[0158] Among them, the data acquisition module refers to the data set that obtains video and text alignment through automatic collection and processing technology. Specifically, it can be implemented by video segmentation algorithms and speech-to-text tools, such as using timestamp alignment technology to match video clips with corresponding commentary text. The label classification module refers to the multi-dimensional classification of data content based on semantic parsing algorithms. Specifically, it can be implemented by using text classification models in natural language processing, such as automatically annotating category labels of commentary text through pre-trained semantic recognition models. The model construction module refers to the multimodal model training architecture that integrates visual and text features. Specifically, it can be implemented by using deformable attention mechanisms and bidirectional feature interaction technologies, such as enhancing the relevance of video and text through cross-modal temporal offset compensation algorithms. The model evaluation module refers to a weighted scoring mechanism based on classification labels. Specifically, it can be implemented by using sub-item indicator weight allocation and comprehensive scoring algorithms, such as setting scoring weights for different semantic categories and calculating the weighted total score.

[0159] Specifically, the system automatically constructs an aligned dataset of video and commentary through the data acquisition module and performs multi-dimensional semantic classification of the commentary content using the label classification module. When training a multimodal large language model, the model construction module simultaneously processes visual and text sequences through cross-modal feature fusion techniques. For example, a deformable attention mechanism dynamically adjusts the association weights between video frames and text tokens. The model evaluation module scores the generated commentary text based on the semantic classification results, for example, giving tactical analysis commentary a higher score than background information commentary. Finally, a weighted calculation yields a comprehensive evaluation result.

[0160] Compared to existing technologies, existing systems typically rely on single-modal data or manually annotated semantic tags, failing to accurately capture the multi-dimensional characteristics of video events. For example, traditional machine commentary systems generate fixed-template commentary text solely through video recognition, lacking in-depth analysis of technical details and tactics. This system, through multimodal semantic classification and a weighted scoring mechanism, can optimize specifically for different commentary dimensions. For example, if a model scores low on the emotional expression dimension during testing, the sampling ratio of relevant training data can be increased.

[0161] Through the above technical solution, this application realizes the refined evaluation and optimization guidance of sports video commentary models. Through automated semantic classification and weighted scoring mechanisms, the system can accurately identify the performance differences of the model in different commentary dimensions. For example, when it is found that the model has a temporal understanding deviation in the tactical analysis task, it can be improved by adjusting the parameters of the two-way attention interaction module. This solution effectively solves the problems of the monotony of existing machine commentary content and the lack of in-depth analysis. For example, in multiple basketball game tests, the commentary text generated by the system outperforms traditional methods in terms of key event detection accuracy and the frequency of technical terminology.

[0162] Regarding the specific limitations of the sports video commentary evaluation system based on a multimodal large language model, please refer to the limitations of the sports video commentary evaluation method based on a multimodal large language model above, which will not be repeated here. The various modules in the above-mentioned sports video commentary evaluation device based on a multimodal large language model can be implemented in whole or in part by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the electronic device in the form of hardware, or can be stored in the memory of the electronic device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0163] In this embodiment, the sports video commentary evaluation system based on a multimodal large language model is essentially equipped with multiple modules to execute the sports video commentary evaluation method based on a multimodal large language model in any of the above embodiments. The specific functions and technical effects can be referred to the above embodiments and will not be repeated here.

[0164] In one embodiment, an electronic device is provided. The electronic device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The electronic device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, the functions or steps on the server side of the above method are implemented.

[0165] In one embodiment, an electronic device is provided. The electronic device may be a client, and its internal structure diagram may be as follows: Figure 5 As shown. The electronic device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, the functions or steps of the client side of the above method are implemented.

[0166] In one embodiment, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0167] A data set is obtained, which includes data pairs consisting of video clips and textual commentary of sports commentary; the data set is semantically classified to determine semantic labels, and the semantic labels classify the data pairs into at least one of the following: description of key events, analysis of technical details, explanation of background information, tactical analysis, explanation of game situation, and emotional expression; a multimodal large language model is constructed, the data set is used to train the multimodal large language model, and a sports video commentary model is determined; the sports video commentary model is scored to determine an evaluation result.

[0168] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0169] A data set is obtained, which includes data pairs consisting of video clips and textual commentary of sports commentary; the data set is semantically classified to determine semantic labels, and the semantic labels classify the data pairs into at least one of the following: description of key events, analysis of technical details, explanation of background information, tactical analysis, explanation of game situation, and emotional expression; a multimodal large language model is constructed, the data set is used to train the multimodal large language model, and a sports video commentary model is determined; the sports video commentary model is scored to determine an evaluation result.

[0170] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or electronic device can be referred to the relevant descriptions on the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0171] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the above-mentioned computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0172] Those skilled in the art will clearly understand that for the sake of convenience and brevity in description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the above-mentioned device or system can be divided into different functional units or modules to complete all or part of the functions described above.

[0173] The embodiments provided above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A sports video commentary evaluation method based on a multimodal large language model, characterized in that: include: Obtaining a data set, the data set comprising data pairs consisting of video clips of sports commentary and textual commentary; Performing semantic classification on the data set to determine semantic labels, wherein the semantic labels classify the data pairs into at least one of the following: description of key events, analysis of technical details, explanation of background information, tactical analysis, explanation of game situations, and emotional expression; Constructing a multimodal large language model, calling the data set to train the multimodal large language model, and determining a sports video commentary model; The sports video commentary model is scored to determine an evaluation result.

2. The method according to claim 1, characterized in that Get the dataset, including: Get sports videos of the same type of events; Extracting the audio file corresponding to the sports video and converting the audio file into a text-based commentary text; The sports video and the commentary text are edited according to the timestamps to form a data pair with one-to-one correspondence between the video clip and the text commentary, which is determined as the data set required for the target sports event.

3. The method according to claim 2, characterized in that Convert the audio file into a text-based commentary, including: Determining a first timestamp of each sentence in the commentary text corresponding to the current video clip; Merging the first timestamps of the respective sentences to determine the timestamp of the current commentary text; Encoding each of the statements to determine a character encoding; Calculating similarity between the character codes corresponding to the sentences, and determining semantic features of the character codes corresponding to the sentences based on the similarity results; Determining whether the sentences corresponding to the video clips are merged based on the similarity results and the semantic features, and determining whether they are continuous clips of the same theme; The audio file is divided according to the timestamp of the current commentary text and the continuous segments of the same theme, and the commentary text corresponding to the same theme is determined.

4. The method according to claim 3, characterized in that Topics covered include fouls, goals, and timeouts.

5. The method according to claim 3, characterized in that The sports video and the commentary text are edited according to the timestamps to form a data pair with one-to-one correspondence between the video clip and the text commentary, including: Merging the first timestamps of the sentences to determine the merged timestamp of the current commentary text; The complete sports video and the complete commentary text are edited according to the fused timestamps to form a data pair with a one-to-one correspondence between the video clip and the text commentary.

6. The method according to claim 1, characterized in that Determine the sports video commentary model including: The deformable attention mechanism is used to process data pairs consisting of video clips and text descriptions to obtain the temporal offset between modalities. Perform feature fusion based on the bidirectional attention interaction and the temporal offset to determine a fusion feature; A multimodal large language model is trained based on the fusion features in the data set until convergence conditions are reached, and then a sports video commentary model that outputs text is determined, wherein the text version of the commentary text is converted into audio output through speech synthesis capabilities.

7. The method according to claim 1, characterized in that The key event description is used to characterize the ability to accurately detect and describe events, reflecting the proficiency in real-time event detection and textual representation; the technical detail analysis is used to characterize the ability to analyze and interpret athlete movements, reflecting the ability of fine-grained visual understanding and contextual knowledge integration; The background information interpretation is used to integrate visual and textual information about the history, characteristics and performance of athletes and teams, demonstrating the depth of the model's contextual knowledge integration and in-depth understanding of visual information; the tactical analysis is used to interpret and explain the team's strategy, formation and tactical adjustments during the game, analyzing and conveying these insights through visual input, highlighting the model's temporal understanding of the dynamics of the game; the game situation interpretation is used to interpret the current state of the game through visual clues, predict its possible progress and provide dynamic situational insights; the emotion expression is used to detect emotional elements from visual data and effectively express emotions in text to attract the audience.

8. The method according to claim 7, characterized in that Scoring the sports video commentary model to determine an evaluation result includes: Splitting a test set from the dataset; Inputting the video clips corresponding to the test set into the sports video interpretation model for interpretation, and obtaining interpretation results for each test video clip; Performing weighted calculation on the interpretation results of each test video clip according to the semantic tags to determine a scoring value; The evaluation result of the sports video commentary model is determined based on the comparison of the score value with the preset score level.

9. The method according to any one of claims 1 to 8, characterized in that: Also includes: If the evaluation result is lower than a first threshold, it is determined that the sports video commentary model fails the evaluation, and a warning icon and / or voice alarm is issued to urge the sports video commentary model to continue training; If the evaluation result is lower than a second threshold, it is determined that the sports video commentary model has passed the evaluation and a passing test result is displayed, wherein the first threshold is lower than the second threshold; If the evaluation result is higher than the second threshold, it is determined that the sports video commentary model is evaluated as excellent, and the excellent test result is displayed.

10. A sports video commentary evaluation system based on a multimodal large language model, characterized in that: include: A data acquisition module, configured to acquire a data set, wherein the data set includes data pairs consisting of video clips and textual descriptions of sports commentary; a label classification module, configured to perform semantic classification on the data set and determine semantic labels, wherein the semantic labels classify the data pairs into at least one of the following: description of key events, analysis of technical details, explanation of background information, tactical analysis, explanation of game situations, and emotional expression; A model construction module is used to construct a multimodal large language model, call the data set to train the multimodal large language model, and determine a sports video commentary model; The model evaluation module is used to score the sports video commentary model according to the semantic tags and determine the evaluation result.

Citation Information

Cited By

  • Real-time wonderful picture snapshot method based on artificial intelligence

    CN120730176A

  • Expressway abnormal event visual detection effect evaluation method and system

    CN121366332A

  • Optimization method for generating fine-grained video description and electronic equipment

    CN121459238A

  • Sports event commentary real-time generation method

    CN122113928A