Multi-modal large language model-based sports video commentary generation method and system

By constructing a multimodal large language model, the problems of insufficient integration of multimodal information and insufficient context utilization in sports video commentary are solved, and high-precision commentary generation is achieved to ensure that the commentary content is highly consistent with the live game.

CN120495957APending Publication Date: 2025-08-15BEIJING QIJI TECHNOLOGY CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510597487.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing sports video commentary generation technology lacks the ability to capture multimodal information, insufficient utilization of metadata and context information, and lacks targeted evaluation indicators, resulting in a lack of specificity, depth and semantic consistency of the commentary content.

Method used

Build a multimodal large language model, align sports video, audio and commentary text through timestamps, project it to a shared embedding space, and use multimodal clustering memory units and retrieval to enhance context learning mechanisms, optimize feature alignment and generate commentary.

Benefits of technology

The in-depth integration of multimodal information is achieved, the accuracy, scene correlation and semantic consistency of the commentary content are improved, and high-quality commentary that conforms to the actual competition situation is generated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495957A_ABST
    Figure CN120495957A_ABST
Patent Text Reader

Abstract

The invention discloses a sports video commentary generation method and system based on a multi-modal large language model, and the method comprises the steps: obtaining a multi-modal data set which comprises a sports video, and an audio and a commentary text corresponding to the sports video; constructing a multi-modal large language model, encoding the sports video, the audio and the explanation text so as to project corresponding video frames, audio waveforms and metadata to a shared embedding space, and determining a multi-modal embedding vector; a multi-modal clustering memory unit is set, multi-modal embedded vectors are grouped, and feature alignment between modals is optimized through comparative learning and information entropy regularization; based on a retrieval enhanced context learning mechanism, a historical instance is retrieved through sparse regularization distance measurement to serve as reference input of a current input multi-mode embedded vector; and jointly inputting the current multi-modal embedded vector and the reference input into the multi-modal large language model to obtain a sports video commentary. According to the method and the device, the problems of insufficient multi-modal information integration and insufficient context utilization are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence and sports video processing technology, and more specifically, to a method and system for generating sports video commentary based on a multimodal large language model. Background Art

[0002] Existing technologies for generating sports video commentary have significant shortcomings in several areas. First, they lack the ability to capture complex multimodal information. Existing methods typically employ single-modal features or simple modal fusion, making it difficult to capture the in-depth details of events and the interactions between multiple modalities. This results in commentary that lacks specificity and nuance. For example, in football matches, existing technologies are unable to accurately capture the connections between the details of players' movements, the audience's real-time reactions, and the context of the game. As a result, the generated commentary often appears hollow and lacks depth.

[0003] Secondly, there's insufficient utilization of metadata and contextual information. While some methods attempt to incorporate structured metadata to enhance commentary accuracy, the effective integration of this information with video, audio, and other modalities is still immature, resulting in insufficient scene relevance and stylistic diversity in the generated commentary. For example, in a football match, existing technologies may not fully utilize metadata such as player statistics, match time, and score, resulting in commentary that is disconnected from the actual game and lacks relevance and personalization.

[0004] Furthermore, evaluation metrics lack objectivity and comprehensiveness. Current generation quality assessments primarily rely on general language generation metrics, but lack specificity in measuring the commentary's semantic consistency, emotional expression, and actual audience experience. For example, in a tennis match, existing technology may be unable to accurately assess whether the commentary truly reflects key events, emotional shifts, and the audience's actual feelings, resulting in evaluation results that lack persuasiveness and practicality.

[0005] In summary, existing technologies are mostly based on a simple combination of pre-trained visual models and language models. This leads to significant deficiencies in feature alignment and semantic mapping between modalities, especially in scenarios requiring high-precision semantic differentiation. Furthermore, most existing methods neglect to fully optimize emotional expression and stylized commentary, resulting in mediocre emotional expression and stylization, failing to meet the diverse needs of audiences. Summary of the Invention

[0006] The purpose of this application is to provide a method and system for generating sports video commentary based on a multimodal large language model to solve at least one of the above technical problems.

[0007] The present application provides a method for generating sports video commentary based on a multimodal large language model, comprising: obtaining a multimodal dataset, the dataset comprising a sports video, and audio and commentary text corresponding to the sports video, wherein the sports video, audio, and commentary text are calibrated based on timestamps; constructing a multimodal large language model, encoding the sports video, audio, and commentary text so that the corresponding video frames, audio waveforms, and metadata are projected into a shared embedding space, and determining a multimodal embedding vector; setting a multimodal clustering memory unit, grouping the multimodal embedding vectors, and optimizing feature alignment between modalities through contrastive learning and information entropy regularization; based on a retrieval-enhanced contextual learning mechanism, retrieving historical instances as reference inputs for a current input multimodal embedding vector through a sparse regularized distance metric; and jointly inputting the current input multimodal embedding vector and the reference input into the multimodal large language model to obtain a sports video commentary.

[0008] Further obtaining a multimodal data set includes: preprocessing the data set, including cleaning, normalization and enhancement; dividing the sports video by events based on timestamps to determine video segments, where events include goals, fouls and timeouts; performing action recognition on the video segments to determine classification labels, where classification labels include shots, saves and passes; transcribing the audio verbatim, aligning the screen with the video segment based on timestamps to determine semantic consistency; and annotating the player information, venue information and event status in the commentary text to ensure that the sports video, audio and commentary text in the data set maintain semantic consistency based on timestamps.

[0009] Furthermore, the sports video, audio and commentary text are encoded so that the corresponding video frames, audio waveforms and metadata are projected into a shared embedding space to determine a multimodal embedding vector, including: using a pre-trained encoder to encode the sports video, audio and commentary text respectively to generate a first embedding vector, a second embedding vector and a third embedding vector respectively; the first embedding vector, the second embedding vector and the third embedding vector are projected into the shared embedding space through a fully connected layer to determine a joint multimodal embedding vector.

[0010] Furthermore, the multimodal clustering memory unit performs the following operations: establishes a cross-modal joint feature space to store multimodal embedding data; calculates the similarity weights between different modal features through a contrastive learning algorithm, and performs full-batch comparisons on positive and negative sample pairs to maximize training efficiency; imposes information entropy regularization constraints when dynamically updating cluster centers, and generates cluster centers through semantic grouping to maintain consistency in semantic relationships.

[0011] Furthermore, the multimodal clustering memory unit also performs the following operations: each unimodal embedding is moved closer to its corresponding mixed-modal cluster center through contrastive loss to achieve unimodal alignment, while moving away from other cluster centers; each unimodal embedding is reconstructed, and a reconstruction loss function is added to the loss function to reconstruct the constraints of the loss function.

[0012] Furthermore, the retrieval-enhanced context learning mechanism specifically includes: building a multimodal feature retrieval index library to store the multimodal embeddings of historical scenes and the corresponding commentary texts; designing a similarity measurement function based on sparse coding to calculate the correlation between the current scene and historical instances; selecting the commentary segments of the top K most relevant instances as contextual prompts and inputting them into the multimodal large language model.

[0013] Furthermore, the expressions corresponding to the retrieval enhanced context learning mechanism include:

[0014] d=||qk||1+β||k-μ||1

[0015] Where q is the feature of the query instance, k is the feature of the multimodal instance in the database, μ is the mean vector of the multimodal feature distribution, which is used for centering; ||·||1 represents the L1 norm, which increases sparsity; β is the weight that controls the sparsity regularization.

[0016] Furthermore, the present application also includes: a KL divergence-based professional terminology distribution comparative evaluation method, which comprehensively evaluates the relevance, consistency, fluency and logic of the content by quantifying the differences in the distribution of professional terminology in the interpretation text and video events.

[0017] Furthermore, this application also includes: constructing a professional terminology dictionary covering technical movements, tactical names and event-specific vocabulary for specific sports; counting the distribution frequency of terms corresponding to key frames in video events; calculating the KL divergence value of the commentary text and the video term distribution to quantify the degree of semantic matching between the two.

[0018] The present application also provides a sports video commentary generation system based on a multimodal large language model, including: a data set module, used to obtain a multimodal data set, the data set including sports videos, and audio and commentary text corresponding to the sports videos, and the sports videos, audio and commentary text are calibrated based on timestamps; a feature fusion module, used to construct a multimodal large language model, encode the sports videos, audio and commentary text so that the corresponding video frames, audio waveforms and metadata are projected into a shared embedding space, and a multimodal embedding vector is determined; a feature alignment module, used to set a multimodal clustering memory unit, group the multimodal embedding vectors, and optimize the feature alignment between modalities through contrast learning and information entropy regularization; a retrieval enhancement module, based on a retrieval enhanced context learning mechanism, retrieves historical instances as reference inputs of the current input multimodal embedding vector through a sparse regularized distance metric; and a commentary generation module, used to jointly input the current input multimodal embedding vector and the reference input into the multimodal large language model to obtain the sports video commentary.

[0019] As can be seen from the above, the present application provides a method and system for generating sports video commentary based on a multimodal large language model, which aligns sports videos, audio and commentary text by timestamps to ensure that cross-modal information is strictly synchronized in the time dimension. Video frames, audio waveforms and text metadata are projected into a shared embedding space to achieve alignment of multimodal features at the semantic level. For example, the goal action in the video, the cheers in the audio and the "Goal!" in the commentary text are mapped to the same semantic space, enhancing the model's understanding of multimodal associations. The embedding vectors are grouped by multimodal clustering memory units, and feature alignment is optimized using contrastive learning (such as CLIP's contrastive loss) and information entropy regularization. Contrastive learning forces different modal representations of the same event to be close, and representations of different events to be far away, thereby reducing the semantic gap between modalities. For example, the visual, audio and text features of the same goal are clustered into the same group, while the representations of different events are pulled apart. Using a sparse regularized distance metric, the model retrieves reference examples from historical examples that are most relevant to the current input. For example, when fed a football game offensive sequence, the model retrieves commentary text and audio features from historical data related to similar offensive scenarios as a reference for generating commentary. The model then combines the multimodal embedding vector of the current input with the retrieved reference input, dynamically incorporating contextual information using the Transformer's self-attention mechanism. For example, the model can incorporate expressions such as "Excellent teamwork!" from historical commentary to generate more contextually accurate current commentary. The retrieval enhancement mechanism incorporates external knowledge, reducing the model's reliance on large-scale parameters and lowering generation uncertainty. For example, the model can reference commentary styles of similar scenes in historical data to avoid generating repetitive or irrelevant content. The shared embedding space and clustered memory unit design support real-time data processing, enabling the model to quickly adapt to the needs of different sports and commentary styles. For example, the same model can generate commentary for different leagues, such as the Premier League and La Liga, without requiring separate training for each league.

[0020] In summary, by constructing a multimodal joint embedding space, dynamically clustering to optimize feature alignment, and enhancing the generation process through contextual retrieval, the problems of insufficient multimodal information integration, insufficient context utilization, and imperfect evaluation system in existing technologies are solved. This approach has the advantages of improving the accuracy of commentary content and enhancing scene relevance. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0022] Figure 1A flowchart of a method for generating sports video commentary based on a multimodal large language model provided in an embodiment of the present application;

[0023] Figure 2 A complete flowchart of a method for generating sports video commentary based on a multimodal large language model provided in an embodiment of the present application;

[0024] Figure 3 A structural diagram of a sports video commentary generation system based on a multimodal large language model provided in an embodiment of the present application;

[0025] Figure 4 A schematic diagram of the principle structure of a sports video commentary generation system based on a multimodal large language model provided in an embodiment of the present application;

[0026] Figure 5 A schematic structural diagram of a multimodal clustering memory unit provided in an embodiment of the present application;

[0027] Figure 6 A schematic structural diagram of an electronic device in one embodiment of the present application;

[0028] Figure 7 Another structural diagram of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0029] The technical solutions in this application will be clearly and completely described below in conjunction with the drawings in this application. Obviously, the described embodiments are only some embodiments of this application, rather than all embodiments. The components of this application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the application provided in the drawings is not intended to limit the scope of the application for which protection is claimed, but merely represents selected embodiments of the application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of this application.

[0030] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.

[0031] Existing technologies for generating sports video commentary primarily rely on single-modality or simple multimodal fusion methods, which struggle to effectively capture complex event details and cross-modal connections. For example, in football matches, existing solutions are unable to simultaneously analyze the interplay between player movements, audience reactions, and the context of the game, resulting in a lack of depth in the generated commentary. Furthermore, insufficient integration of metadata with video and audio results in a disconnected and monotonous commentary. Furthermore, common language evaluation metrics struggle to measure semantic consistency and emotional expression in commentary, hindering practical application effectiveness.

[0032] To address these issues, we must first address the gaps in interpretation caused by insufficient alignment of multimodal features. Analysis revealed semantic bias when unimodal features are projected into a shared space, necessitating the introduction of a cross-modal comparative learning mechanism. Secondly, the low utilization of historical contextual information can be addressed by building a search index to link similar scenarios. Finally, semantic consistency evaluation should be combined with quantitative analysis of the distribution of specialized terminology to ensure that the content aligns with the actual events.

[0033] See also Figure 1 , is a flow chart of a method for generating sports video commentary based on a multimodal large language model provided in an embodiment of the present application, which is described in detail as follows:

[0034] Step S101: obtaining a multimodal data set, the data set including sports videos, and audio and commentary text corresponding to the sports videos, wherein the sports videos, audio and commentary text are calibrated based on timestamps;

[0035] Step S102: construct a multimodal large language model, encode the sports video, audio, and commentary text, so that the corresponding video frames, audio waveforms, and metadata are projected into a shared embedding space, and a multimodal embedding vector is determined;

[0036] Step S103: Setting a multimodal clustering memory unit, grouping the multimodal embedding vectors, and optimizing the feature alignment between modalities through contrastive learning and information entropy regularization;

[0037] Step S104: Based on the retrieval-enhanced contextual learning mechanism, historical instances are retrieved through a sparse regularized distance metric as reference inputs for the current input multimodal embedding vector;

[0038] In step S105 , the current input multimodal embedding vector and the reference input are jointly input into a multimodal large language model to obtain sports video commentary.

[0039] Multimodal datasets are collections of sports videos, corresponding audio, and commentary text. Timestamp calibration can be used to achieve cross-modal alignment, for example, by segmenting video clips by goal events and synchronizing them with audio transcripts. This dataset provides a semantically consistent input foundation for subsequent feature fusion.

[0040] Projecting into a shared embedding space involves encoding video frames, audio waveforms, and text metadata into a unified vector representation. This can be achieved by using a pre-trained encoder to generate unimodal embeddings, which are then mapped through a fully connected layer. This technical feature eliminates modal differences and establishes a common representational foundation for cross-modal interaction.

[0041] The multimodal clustering memory unit is a module that semantically groups embedded vectors. Specifically, it uses a contrastive learning algorithm to calculate cross-modal similarity and dynamically updates cluster centers through information entropy regularization. This unit improves the accuracy of feature alignment between modalities and addresses the problem of unimodal semantic bias.

[0042] The retrieval-enhanced contextual learning mechanism selects relevant explanation segments from historical examples as references. Specifically, a sparse coding similarity metric is used to select the top K relevant examples. This mechanism enhances the model's understanding of scene context and improves the logical coherence of the explanation.

[0043] Specifically, the system first encodes video, audio, and text data with aligned timestamps into a unified vector, eliminating the modality gap. Clustered memory units are then used to optimize feature distribution, allowing similar event features from different modalities to cluster in the embedding space. During the generation phase, sparse similarity retrieval matches historical scenarios, combining the current input with the reference context into the model. This process enables deep interaction of multimodal information and reuse of historical experience, ensuring a high-precision semantic match between commentary and actual match action.

[0044] Compared to existing technologies, traditional approaches rely on single-modal features or simple splicing and fusion, failing to effectively align cross-modal semantics. For example, when capturing a shot, existing technologies may overlook the connection between the audience's cheers and slow-motion replays. However, this approach achieves collaborative optimization of visual and auditory features through clustered memory units. Furthermore, existing technologies lack a historical scene retrieval mechanism, making it difficult to maintain the consistency of commentary style. This approach, however, achieves context-adaptive enhancement through sparse coding metrics.

[0045] Through the above technical solutions, this application has achieved significant improvements in the accuracy of multimodal feature fusion. For example, in a football game scene, it can accurately associate a player's shooting action, the referee's whistle, and score change information. The cross-modal alignment mechanism ensures that the commentary content contains game details and background information, such as synchronously describing the goalkeeper's historical data when identifying a save. The retrieval enhancement mechanism enables the model to reuse the commentary logic of similar scenes, such as automatically introducing tactical analysis templates during a game pause. The semantic consistency evaluation module effectively quantifies the matching degree between the commentary content and the video event to avoid generating descriptions that are out of touch with reality.

[0046] In some embodiments, a multimodal large language model backbone is responsible for perceiving multimodal information from video, audio, and metadata, and generating corresponding commentary text;

[0047] Multimodal clustering memory units for storing and retrieving multimodal embeddings, building a cross-modal joint feature space through contrastive learning, and supporting retrieval-enhanced contextual learning processes;

[0048] Retrieval-enhanced context learning mechanism, which enhances the generated context through multimodal feature retrieval and introduces a distance metric based on sparsity regularization to optimize retrieval efficiency and context diversity;

[0049] The KL divergence-based comparative evaluation method for professional terminology distribution quantifies the differences in the distribution of professional terminology between interpretation text and video events, and comprehensively evaluates the relevance, consistency, fluency, and logic of the content.

[0050] Among them, the backbone of the multimodal large language model refers to a deep learning architecture that uses a cross-modal encoder. Specifically, it can be implemented by using a structure that connects a visual Transformer with an audio convolutional network in parallel, and multimodal feature fusion is achieved through a shared attention mechanism. The multimodal clustering memory unit refers to a storage module with a semantic grouping function. Specifically, the K-means algorithm can be used to cluster mixed modal embeddings, and the feature space distribution can be optimized through a contrast loss function. The retrieval-enhanced context learning mechanism refers to a method for dynamically acquiring relevant context. Specifically, it can be implemented using an approximate nearest neighbor search algorithm, and the retrieval dimension is reduced by introducing an L1 norm constraint. For example Figure 2 As shown in step S106, the professional term distribution comparison evaluation method based on KL divergence refers to a probability distribution difference quantification tool, which can be implemented by combining a pre-built domain term dictionary with a text analysis model to evaluate content quality by calculating the distribution difference between the generated text and the standard term set.

[0051] Specifically, video frames, audio waveforms, and structured metadata are fed into the multimodal large language model backbone and converted into feature vectors via the visual encoder and audio encoder, respectively. These features undergo cross-modal attention calculations in a shared embedding space, forming a joint representation that incorporates spatiotemporal correlations. A multimodal clustering memory unit stores the embedding vectors of historical scenes by semantic category. When a new scene is input, it retrieves semantically similar historical examples by comparing similarities. A retrieval-enhanced contextual learning mechanism concatenates the retrieved historical commentary segments with the current scene features and feeds this into the language decoder to generate the commentary text. During this process, a sparse regularized distance metric ensures that only critical contextual information is retained, avoiding interference from irrelevant content. The generated commentary text is then fed into an evaluation module, where KL divergence is calculated against a predefined football terminology library to quantify term coverage and distribution rationality.

[0052] Compared to existing technologies, traditional methods employ single-modal feature extraction and post-processing fusion strategies, making it difficult to coordinate visual action recognition with speech sentiment analysis. For example, in a basketball game scenario, existing technologies might independently process dunk video features and audience cheers, failing to establish a correlation between "dunk action, cheer intensity, and key moments in the game." This solution achieves cross-modal semantic alignment through a joint feature space, creating a unified event representation for the player's leaping motion and the cheering sound waves. Regarding real-time retrieval, traditional cosine similarity-based retrieval methods require traversing the entire historical data set. However, this solution utilizes a sparse regularization metric to significantly reduce computational complexity, enabling millisecond-level response times to retrieve valid context.

[0053] Through the above technical solutions, this application effectively improves the efficiency of integrating multimodal information, accurately associates video actions, on-site sounds and player data in football game scenes, and generates commentary content that includes tactical terms and emotional expressions. By constructing a cross-modal joint feature space, the semantic deviation problem caused by the separation of visual and auditory features in traditional methods is solved. The retrieval enhancement mechanism is used to dynamically introduce relevant historical scenes to enhance the contextual coherence and style diversity of the commentary. The evaluation method based on the distribution of professional terms breaks through the limitations of the traditional BLEU indicator, ensuring that the generated commentary content conforms to the domain knowledge specifications and covers key event elements.

[0054] Optionally, in some embodiments, obtaining a multimodal dataset includes:

[0055] Preprocess the data set, including cleaning, normalization and enhancement;

[0056] Divide sports videos into events based on timestamps to identify video segments, where events include goals, fouls, and timeouts;

[0057] Perform action recognition on the video clips and determine the classification labels, including shooting, saving and passing;

[0058] The audio is transcribed verbatim and aligned with the video clips according to the timestamps to ensure semantic consistency;

[0059] The player information, venue information and event status in the commentary text are annotated so that the sports videos, audios and commentary texts in the dataset maintain semantic consistency based on timestamps.

[0060] Cleaning refers to removing noisy data or invalid segments from a dataset. This can be achieved using a timestamp-based sliding window filtering method, where invalid segments shorter than a set duration are filtered out by setting a time window threshold. Normalization refers to standardizing data formats from different sources into standardized inputs. This can be achieved through a combination of frame sampling and feature scaling, for example, adjusting the video frame rate to a fixed value and the audio sampling rate to a specific frequency. Enhancement refers to increasing data diversity to improve model generalization. This can be achieved through spatiotemporal data augmentation methods, such as mirroring video segments or adding Gaussian noise. Event segmentation refers to identifying key nodes based on game rules. This can be achieved through a rule-based state machine algorithm, segmenting video segments by identifying triggers such as the referee's whistle or score changes. Action recognition refers to classifying and labeling motion behaviors in videos. This can be achieved using a 3D convolutional neural network model, such as a pre-trained SlowFast network to extract spatiotemporal features. Verbatim transcription refers to converting an audio stream into a text sequence. This can be achieved using an end-to-end speech recognition model, such as a Transformer architecture with a Connectionist Temporal Classification loss. Screen alignment ensures temporal synchronization between audio and video content. This can be achieved using a dynamic time warping algorithm, which calculates a similarity matrix between audio waveform features and video optical flow features to achieve cross-modal alignment. Player tagging involves identifying and recording player identifiers. This can be achieved by combining object detection with facial recognition, for example, using the YOLO model to locate players and then using FaceNet to extract facial features.

[0061] Specifically, during the data preprocessing phase, a sliding window filtering method is first used to remove data segments shorter than a set threshold, for example, filtering out invalid video segments shorter than a second. The retained video data is then framed, the frame rate uniformly adjusted to frames per second, and the audio waveform is resampled to hertz. Random Gaussian noise and mirror flipping operations are added to enhance the diversity of the video data. During the event segmentation phase, a state machine model is established by analyzing the rules of the game. When the referee's whistle frequency exceeds a set threshold or the score changes, the video segmentation operation is triggered to generate video segments containing the complete event. A pre-trained SlowFast network is used to extract spatiotemporal features for each segmented video segment, and a fully connected layer classifier outputs action labels such as shots and saves. After the audio data is converted into text sequences using an end-to-end speech recognition model, a dynamic time warping algorithm is used to calculate the similarity matrix between the audio MFCC features and the video optical flow features, achieving frame-by-frame synchronization of audio and video. Finally, an object detection model is used to locate the player's position, and identity recognition is achieved by combining it with a facial feature database. Player numbers, location information, and match timestamps are recorded as structured metadata.

[0062] Compared with existing technologies, traditional methods usually use fixed intervals to split video clips, which cannot accurately capture key points of the game. However, this solution uses a rule-driven event division mechanism to accurately capture video segments containing complete game events. In existing technologies, there is often a timing deviation between audio transcription and video content. This solution uses a dynamic time warping algorithm to achieve precise alignment of audio and video, eliminating the problem of time misalignment in cross-modal data. Traditional methods mostly rely on manual labeling of player information. This solution automatically generates structured metadata through target detection and facial recognition technology, significantly improving data labeling efficiency. Existing data sets generally have the problem of semantic inconsistency between modalities. This solution uses multi-stage alignment processing to ensure strict synchronization of video clips, audio text and text commentary in the time dimension.

[0063] Through the above technical solutions, this application effectively solves the problems of low quality and semantic inconsistency of multimodal data. Through event-driven video segmentation methods, it ensures that each data unit contains complete game events; combined with automated action recognition and audio-visual alignment technology, it eliminates the time deviation of cross-modal data; and uses structured metadata annotation to enhance the semantic richness of the data set. These improvements enable subsequent model training to obtain high-quality multimodal aligned data, laying the foundation for generating accurate and coherent sports commentary. For example, in a football game scenario, this solution can accurately associate video clips of shooting actions, corresponding cheers from the audience, and technical terms in the commentary text, significantly improving the semantic accuracy and scene adaptability of the generated commentary.

[0064] For example, data is collected in the following ways:

[0065] Data source: Collect high-quality sports event videos from television broadcasts and Internet platforms (such as live events and professional sports commentary channels), ensuring coverage of multiple sports (such as football, table tennis, etc.).

[0066] Video quality: All videos are screened to ensure a frame rate of no less than 25fps and a resolution of no less than 720p, ensuring clear details and laying the foundation for accurate labeling and modal information extraction.

[0067] For example, data annotation is performed in the following way

[0068] Key time point annotation: We annotate important events in the video (such as goals, fouls, and timeouts) with relevant timestamps, laying the foundation for the subsequent generation of commentary and metadata. We use the first s seconds of the timestamp as the start point of the video segment and the last s seconds as the end point (for example, in a football match, s is 15, forming a 30-second time window) to crop the video segment.

[0069] Action labeling: Through action recognition models and manual review, key actions appearing in video clips are classified and labeled (such as shooting, saving, passing, etc.).

[0070] Narration segment association: The narration content in the video clip is transcribed verbatim and aligned with the video event screen to ensure semantic consistency.

[0071] Background metadata annotation: including player personal information (such as name, position, past performance), event status (such as score, game time) and venue information.

[0072] Quality Assurance: A combination of automated annotation and manual verification is used. Initial annotations are generated using tools (such as speech recognition models and natural language processing algorithms) and reviewed by sports professionals to ensure that the annotations are highly relevant to the video events.

[0073] For example, data enhancement is performed in the following ways:

[0074] Combined with contextual scenarios, fine-grained meta-information is added to each video, such as the background of the event, player tactical analysis, etc., to enhance the richness and semantic depth of the dataset.

[0075] In this embodiment, compared with existing datasets (such as SoccerNet-V2 or MatchTime), this dataset integrates video, audio and rich background metadata, and provides multimodal fine-grained annotation. It pays more attention to the semantic connection between events and enhances fine-grained annotation. For example, in the same game, the timing and contextual information of adjacent events are annotated to highlight the details; the causal relationship between different actions (such as passing leads to goals) is also reflected in the annotation. By combining the semantic information of the context, it ensures that the video events are highly relevant to the commentary paragraphs, achieves high-precision matching, and avoids the situation where the annotation content is generalized or inconsistent with the events. The commentary content covers a variety of annotations of different commentary styles (such as passionate and analytical), ensuring that diversity and stylization lay the foundation for the stylized generation of the model.

[0076] In the process of building the dataset, historical scenarios similar to the current event (such as similar tactical actions, past performances of the same players, etc.) are extracted and stored, and historical scenario instances are used as the core resource of the Retrieval Augmentation Generation (RAG) module.

[0077] Historical scene instances are directly derived from the multimodal feature space annotated by fine-grained datasets. Using a clustering algorithm, similar scenes and their explanations are aggregated and stored to form a retrieval library for model learning.

[0078] Unique characteristics of annotations, with the following differences:

[0079] The annotation process incorporates temporal logic and context between events. For example, chain annotation is performed on the commentary and subsequent tactical analysis of the same event, enhancing the coherence of the generated commentary while also improving temporal order and contextual information.

[0080] For commentators with different styles (such as passionate commentary or technical analysis commentary), the language styles are classified and labeled, so that the model can generate personalized commentary content and improve the personalization and stylization of the commentary.

[0081] By aligning and reviewing the voice commentary text word by word, we ensure that the generated text accurately reflects the video events semantically, thus ensuring the semantic consistency of the model.

[0082] The application of the dataset in the model, for example, in actual use:

[0083] Through pre-trained visual, audio, and language encoders, the video frames, audio waveforms, and metadata in the dataset are projected into a shared embedding space, completing multimodal feature alignment and providing complete contextual information for the model.

[0084] Historical scene instances are used as a retrieval library, and the most similar multimodal fragments are passed to the language model as context through the RAG mechanism. Context-enhanced explanation generation is achieved through retrieval-enhanced learning.

[0085] In some embodiments, aligning video frames with commentary text enables the model to capture fine-grained actions and events (such as brief, coordinated passes or quick counterattacks), improving the model's ability to capture detailed information. Leveraging the commentator's emotional style annotations in metadata helps the model generate commentary that better aligns with the scene's emotions, helping to strengthen emotion and style learning. The historical scene retrieval function enables the model to generate stylized commentary consistent with similar past events, improving the coherence and diversity of the commentary.

[0086] In this embodiment, the application of this dataset enables the model to not only focus on the action itself when generating commentary, but also to combine the context of the game and historical data to generate richer and more detailed commentary content.

[0087] The dataset constructed in this way provides a solid data foundation for generating high-quality sports commentary, while significantly enhancing the model's ability to capture details, express emotions, and understand context.

[0088] Existing datasets typically focus on a single modality (such as video frames or text commentary) or simple modal fusion, while this dataset achieves deep integration of video, audio and metadata, and constructs richer and more fine-grained multimodal features.

[0089] Fine-grained annotation and semantic association, key time point annotation: Events in the video (such as goals, fouls, tactical adjustments) are accurately annotated and aligned with timestamps. Action annotation: By combining action recognition models with manual review, key actions (such as shots, passes, and saves) are annotated with detailed classification.

[0090] During the dataset construction process, historical scene instances were extracted and stored as core resources for the Retrieval Augmentation Generation (RAG) module. These instances consist of historical segments and their accompanying textual descriptions that are highly similar to the current event, forming a cross-modal embedding feature space. A clustering algorithm was used to categorize events (e.g., a pass leading to a shot), enabling the reuse and optimization of event templates.

[0091] Optionally, in some embodiments, encoding sports video, audio, and commentary text so that corresponding video frames, audio waveforms, and metadata are projected into a shared embedding space, and determining a multimodal embedding vector includes:

[0092] Use pre-trained encoders to encode sports videos, audio, and commentary text, respectively, to generate the first embedding vector, the second embedding vector, and the third embedding vector;

[0093] The first embedding vector, the second embedding vector, and the third embedding vector are projected into a shared embedding space through a fully connected layer to determine a joint multimodal embedding vector.

[0094] Among them, the pre-trained encoder refers to a feature extraction model trained based on large-scale data. Specifically, it can be implemented using the CLIP model in the visual field, the Wav2Vec model in the audio field, and the BERT model in the text field, which encode video frames, audio waveforms, and text sequences into high-dimensional vector representations respectively. The fully connected layer refers to a linear transformation layer with learnable weights, which can be implemented using a multi-layer perceptron structure. It is used to reduce the dimensionality of the embedding vectors of different modalities and fuse them into a unified dimensional space. The shared embedding space refers to a jointly optimized vector space, which can be achieved through a cross-modal comparative learning objective function, making the semantic information of different modalities comparable in geometric distance. The joint multimodal embedding vector refers to a unified representation that integrates video, audio, and text features. Specifically, it can be generated through vector splicing or weighted summation, providing a basis for subsequent feature alignment and semantic understanding.

[0095] Specifically, during the encoding phase, a pre-trained visual encoder extracts spatiotemporal features from the sports video frame sequence, forming a first embedding vector of dimension d_v. A pre-trained audio encoder extracts spectral features from the audio waveform, forming a second embedding vector of dimension d_a. A pre-trained language model extracts semantic features from the commentary text, forming a third embedding vector of dimension d_t. These three embedding vectors are then fed into their respective fully connected layers, where the dimensions are uniformly mapped to a shared space of d_e, where d_e can be set to, for example, 512 or 768. This shared space mapping ensures that the feature vectors of different modalities satisfy cosine similarity constraints on their geometric distribution, enabling cross-modal feature alignment.

[0096] Compared to existing technologies, traditional methods typically employ single-modal encoding or simple splicing fusion, resulting in significant differences in feature distribution between modalities and weak semantic correlation. This solution, through joint optimization of a shared embedding space, projects video action features, audio intonation features, and text semantic features into a unified metric space, overcoming the feature matching difficulties inherent in traditional methods due to modal heterogeneity. For example, in a football game scene, the visual features of a shot, the audio features of the crowd cheering, and the text description of a "wonderful shot" can all be represented by similar vectors in the shared space.

[0097] Through the above technical solution, this application effectively solves the technical bottleneck of heterogeneous multimodal feature fusion, achieving a deep correlation between video images, sound information and text semantics. Specifically, in sports commentary scenarios, it can accurately capture the visual details of players' shooting movements, the intensity changes of the audience's cheers, and the inherent connections between professional commentary terminology, thereby generating commentary content that is highly consistent with the progress of the game.

[0098] Furthermore, the multimodal clustering memory unit performs the following operations:

[0099] Establish a cross-modal joint feature space to store multimodal embedding data;

[0100] The similarity weights between features of different modalities are calculated through contrastive learning algorithms, and full batch comparisons are performed on positive and negative sample pairs to maximize training efficiency.

[0101] Information entropy regularization constraints are imposed when dynamically updating cluster centers, and cluster centers are generated through semantic grouping to maintain semantic relationship consistency.

[0102] Among them, the cross-modal joint feature space refers to mapping the embedded data of different modalities to a shared representation of the same dimensional space. Specifically, it can be implemented using a fully connected layer or an attention mechanism to eliminate heterogeneous differences between modalities. The contrastive learning algorithm refers to the use of similarity weights to screen sample pairs with high correlation between modalities. Specifically, cosine similarity or Euclidean distance can be used to measure feature similarity, and redundant calculations can be reduced by full-batch comparison. The information entropy regularization constraint refers to the introduction of an information entropy minimization objective when updating cluster centers. Specifically, the information entropy of the cluster center distribution can be calculated and added to the loss function to prevent the cluster center distribution from deviating from the semantic grouping structure.

[0103] Specifically, the cross-modal joint feature space projects video, audio, and text embeddings onto a shared dimension through a fully connected layer, eliminating modal differences. During training, the contrastive learning algorithm imposes similarity maximization constraints on positive sample pairs and separation constraints on negative sample pairs to improve feature alignment efficiency. When dynamically updating cluster centers, information entropy regularization constrains the distribution of cluster centers to make them closer to the semantic grouping structure. For example, different modal features of the same game event are assigned to the same cluster to avoid semantic drift. As a result, the multimodal cluster memory unit can optimize the accuracy of feature alignment while maintaining the consistency of cross-modal semantic relationships.

[0104] Compared to existing methods, which typically use static feature pooling or unimodal clustering, this approach results in insufficient cross-modal feature alignment and an inability to adapt to dynamic semantic changes. This approach, however, uses dynamically updated cross-modal cluster centers, combined with contrastive learning and information entropy constraints, to capture fine-grained associations between multimodal features while suppressing noise interference. For example, in a football match, the video, cheers, and commentary of a goal event can be assigned to the same dynamic cluster center, whereas existing methods may misalign features due to a lack of semantic grouping constraints.

[0105] Through the above technical solution, this application can effectively improve the accuracy of multimodal feature alignment and ensure that cross-modal semantic relationships remain stable during dynamic updates. For example, in a basketball game scene, the "three-point shot" action in the commentary text and the shooting action and audience cheers in the video can be accurately aligned at the same cluster center, avoiding the semantic fragmentation problem caused by modal differences in traditional methods, thereby improving the content relevance and context consistency of the generated commentary.

[0106] Furthermore, the multimodal clustering memory unit also performs the following operations:

[0107] Each unimodal embedding is moved closer to its corresponding mixed-modal cluster center through contrastive loss to achieve unimodal alignment while moving away from other cluster centers;

[0108] Each unimodal embedding is reconstructed and a reconstruction loss function is added to the loss function to constrain the reconstruction loss function.

[0109] Multimodal clustering memory units are used to store and retrieve multimodal embeddings, construct cross-modal joint feature spaces through contrastive learning, and support retrieval-enhanced contextual learning processes.

[0110] Among them, contrastive loss refers to a loss function that is optimized by comparing the similarity differences between positive and negative sample pairs. Specifically, cosine similarity can be used to calculate the degree of match between unimodal embeddings and mixed-modal cluster centers, and backpropagation can be used to adjust the distribution of the embedding space. Mixed-modal cluster centers refer to semantic aggregation points generated by multimodal joint embedding features through a clustering algorithm. Specifically, a dynamic k-means algorithm can be used to iteratively update according to multimodal semantic relationships to characterize common features across modalities. Unimodal alignment refers to optimizing contrastive loss to make the unimodal embedding vector close to the mixed-modal cluster area to which it belongs in the shared space. Specifically, this can be achieved by constraining the distance threshold between different modal embeddings under the same event, ensuring that visual, auditory, and textual modalities are mapped at the same semantic level. Reconstruction loss function refers to the constraints for decoding and restoring unimodal embedding vectors. Specifically, an autoencoder structure can be used to reversely reconstruct the embedding vector into the original modal features, and the feature expression ability is maintained by calculating the reconstruction error.

[0111] Specifically, during the training phase, the embedding vectors of the visual, auditory, and textual modalities are compared with the mixed-modal cluster centers. For the visual embedding vector, its similarity score with the correct cluster center is calculated as a positive sample pair. Cluster centers of other events are randomly selected as negative sample pairs. A contrastive loss function is optimized to cluster the visual features towards the correct cluster center. Simultaneously, the auditory and textual embedding vectors undergo the same operation to ensure that the features of each modality form compact semantic clusters in the shared space. During the feature alignment process, each unimodal embedding is required to maintain a minimum distance from the corresponding mixed-modal cluster center and a maximum distance from other irrelevant cluster centers. Furthermore, each unimodal embedding vector is reconstructed into the original modality features through a fully connected network. For example, the visual embedding is reconstructed into a video frame feature vector, and the reconstruction loss is calculated using the mean squared error. The contrastive loss and the reconstruction loss are linearly superimposed according to preset weights to form the final optimization target. The encoder and cluster center parameters are updated synchronously during the backpropagation process.

[0112] Compared with existing technologies, traditional methods only constrain the distribution of multimodal features through simple contrastive learning, fail to establish a bidirectional optimization mechanism between dynamic cluster centers and unimodal embeddings, and lack explicit constraints on feature reconstruction capabilities. This solution constructs cross-modal semantic anchors through mixed-modal cluster centers, utilizes contrastive loss to align unimodal features with the joint semantic space, and simultaneously maintains the representational power of unimodal features through reconstruction loss, forming a dual constraint mechanism. This design effectively solves the problem of semantic drift in unimodal features during cross-modal fusion in traditional methods, and avoids feature degradation caused by over-reliance on inter-modal contrastive learning.

[0113] Through the above technical solution, the present application solves the technical problem of inconsistency between single-modal features and mixed-modal semantic space in the process of multimodal alignment, and improves the matching accuracy of action recognition and language description when generating video commentary. In a football game scene, when a video clip contains a shooting action, the visual embedding can be accurately aligned to the center of the mixed cluster containing the shooting semantics, while the audio embedding is aligned with the corresponding cheering cluster center, ensuring that the generated commentary text covers both action details and scene atmosphere information. Furthermore, the reconstruction loss constraint enables each modality encoder to maintain the ability to express the original features during the cross-modal fusion process, avoiding the loss of modality-specific information due to over-optimization of contrast loss, thereby improving the information integrity and semantic coherence of the commentary content.

[0114] Among them, the multimodal clustering memory unit of the present application refers to a computing module that performs persistent storage of mixed modal embeddings and establishes a retrieval mechanism. Specifically, the K-means clustering algorithm can be used to group the embeddings to form semantically consistent feature clusters. Its function is to solve the problem that single-modal storage cannot reflect multimodal correlation.

[0115] Among them, the cross-modal joint feature space refers to the mathematical expression of mapping the embeddings of different modalities to the same semantic space. Specifically, a contrastive learning algorithm can be used to optimize the distribution of embeddings of different modalities, and forced alignment is achieved by calculating the similarity loss function between modalities. Its role is to eliminate the semantic mapping deviation caused by simple splicing between modalities in traditional methods.

[0116] Among them, contrastive learning refers to a method of training a model by optimizing the distance relationship between similar samples and dissimilar samples. Specifically, information entropy regularization can be used to constrain the distribution consistency of embeddings of different modalities. Its role is to ensure that video, audio and metadata embeddings are aligned within the same semantic cluster.

[0117] Specifically, the multimodal clustering memory unit projects video frames, audio waveforms, and metadata into a shared embedding space through an encoder to form a joint feature representation. In the storage phase, the K-means clustering algorithm is used to divide the mixed modal embeddings into semantically consistent clusters and store them persistently in the database. In the retrieval phase, similar instances are quickly matched by calculating the sparse regularized distance between the current scene embedding and the center of the historical cluster. In the feature alignment phase, a triplet loss function is constructed using a contrastive learning framework to force the embeddings of similar events in different modalities to cluster in the same cluster, while pushing away the embeddings of events of different categories. As a result, the retrieval-enhanced contextual learning mechanism can perform cross-modal association retrieval based on semantically aligned multimodal features, providing the generation process with historical commentary fragments that are highly relevant to the current scene.

[0118] Compared with existing technologies, existing methods usually use single-modal feature libraries to store visual or textual information and achieve multimodal fusion through simple splicing, resulting in insufficient cross-modal semantic relevance. For example, traditional commentary systems rely solely on text keyword matching when retrieving historical data, and are unable to capture the correlation pattern between video actions and audience cheers. However, this solution achieves deep integration of cross-modal information by constructing a joint feature space to align different modal data within the same semantic cluster. In addition, existing technologies lack regularization constraints on the embedding distribution, while this solution optimizes intra-cluster compactness and inter-cluster discrimination through contrastive learning, significantly improving the semantic consistency of feature retrieval.

[0119] Through the above technical solution, this application can accurately retrieve multimodal features of similar scenes in a historical database based on videos of shots outside the penalty area, crowd cheers, and player metadata in the context of football match commentary. For example, when Messi's long-range shot is detected, the system can cross-modally associate the term "long-range shot," the cheer waveform, and the player's historical data to generate commentary text that includes professional terms such as "outside the penalty area" and "threat" with accurate time information, effectively improving scene relevance and terminology accuracy.

[0120] Optionally, the retrieval enhanced context learning mechanism specifically includes:

[0121] Build a multimodal feature retrieval index library to store multimodal embeddings of historical scenes and corresponding explanatory texts;

[0122] Design a similarity measurement function based on sparse coding to calculate the correlation between the current scene and historical instances;

[0123] The explanation segments of the top K most relevant instances are selected as contextual prompts and input into the multimodal large language model.

[0124] Among them, the multimodal feature retrieval index library refers to a structured database that stores multimodal embeddings of video, audio, and text in historical scenes and their corresponding explanations. Specifically, it can be implemented using a hierarchical index structure or a graph database to achieve fast retrieval and matching of multimodal data. The similarity measurement function of sparse coding refers to the calculation of the correlation between features through a mathematical method of sparse regularization constraints. Specifically, it can be implemented using an optimization algorithm with L1 norm constraints, which improves the robustness of feature matching by reducing noise interference. The commentary segments of the top K most relevant instances refer to the historical commentary content that is closest to the semantics of the current input scene, selected based on the similarity measurement results. Specifically, it can be implemented using a maximum heap sorting algorithm, which provides accurate contextual references for the model by dynamically intercepting key information.

[0125] Specifically, the historical multimodal data is first encoded and stored in a feature retrieval index library to form a scalable retrieval dataset. When processing new input scenarios, the correlation score between the current multimodal embedding and all instances in the database is calculated using a sparsely coded similarity measurement function. This function introduces an L1 norm constraint, forcing the similarity calculation process to focus on key feature dimensions and suppressing the interference of redundant information. The top K most relevant instances are then selected based on the score ranking, and their corresponding commentary text fragments are extracted as contextual cues. Finally, the current input multimodal embedding and the retrieved contextual cues are jointly input into the large language model, and the generation of commentary is achieved through a cross-modal attention mechanism.

[0126] Compared with existing technologies, traditional methods typically use linear similarity calculations or single-modality retrieval, resulting in insufficient matching of retrieval results with multimodal scenarios and low computational efficiency. This solution, through sparse coding feature selection and a hierarchical index structure, significantly reduces computational complexity while maintaining retrieval accuracy. Furthermore, the retrieval mechanism based on multimodal joint embedding can capture cross-modal semantic associations, avoiding the contextual bias caused by modal fragmentation in traditional methods.

[0127] Through the above technical solution, this application effectively improves the accuracy and applicability of contextual references in the generation of sports video commentary. By retrieving historical commentary clips that are highly relevant to the current scene, the model can more accurately capture the semantic focus of game events. For example, in a penalty kick scene in a football match, it can automatically associate historical commentary clips of similar shooting actions to generate commentary content that includes descriptions of specific player technical actions, solving the problem of hollow commentary content caused by the lack of targeted context in traditional methods.

[0128] Optionally, the expression corresponding to the search enhanced context learning mechanism includes:

[0129] d=||qk||1+β||k-μ||1

[0130] Where q is the feature of the query instance, k is the feature of the multimodal instance in the database, μ is the mean vector of the multimodal feature distribution, which is used for centering; ||·||1 represents the L1 norm, which increases sparsity; β is the weight that controls the sparsity regularization.

[0131] Among them, the feature of the query instance refers to the multimodal embedding vector corresponding to the scene for which the explanation needs to be generated. It can be obtained by encoding video frames, audio waveforms and metadata, and is used to represent the semantic information of the current scene.

[0132] The mean vector of the multimodal feature distribution refers to the average feature vector calculated from the multimodal embedding data of the historical scene. It can be achieved through sliding window averaging or global statistical methods to eliminate feature distribution offset.

[0133] The L1 norm is a method for calculating the sum of the absolute values of each dimension of a feature vector. It can be applied to similarity metrics to filter out key feature dimensions through sparsity constraints. The sparsity regularization weight is a hyperparameter that balances the sparsity constraint with the original similarity in similarity calculations. It can be determined through grid search or validation set tuning to suppress interference from noisy features.

[0134] Specifically, the expression centralizes the query instance features and historical instance features to eliminate feature distribution differences, and then uses the L1 norm to calculate the sparse similarity between the two. During the calculation process, the strength of the sparse constraint is controlled by the β parameter, so that the similarity measure can highlight key feature dimensions and suppress redundant information. For example, in a football game scenario, when the input video contains a shooting action, the expression can filter out key modal features related to the shot in the historical database, such as the player's leg movement features, the shooting sound waveform features, and the shooting success rate metadata. After sorting by sparse similarity, the top K most relevant historical commentary clips are selected as contextual prompts and input into the multimodal large language model to generate the commentary of the current scene. Compared with existing technologies, traditional methods usually use Euclidean distance or cosine similarity for instance retrieval. Such methods are easily affected by the noise component in high-dimensional features, resulting in insufficient semantic matching between the retrieval results and the current scene. This solution, by introducing L1 norm sparsity constraints and centralization processing, can effectively filter non-critical feature components. For example, in a basketball game scenario, it can suppress the interference of the waveform characteristics of the audience's cheers on the generation of tactical action commentary, thereby improving the relevance of retrieval instances.

[0135] Through the above technical solution, the present application can significantly improve the semantic alignment accuracy between historical instances and current scenes, especially when processing sports videos containing complex multimodal interactions, such as high-speed sliding and collision actions in ice hockey games. It can accurately match commentary clips with similar motion trajectories, collision sounds and penalty metadata in the historical database, and then generate real-time commentary text with specific content and logical coherence.

[0136] Optionally, in some embodiments, Figure 2 As shown in step S106, by quantifying the difference in the distribution of professional terms in the interpretation text and the video event, the relevance, consistency, fluency and logic of the content are comprehensively evaluated.

[0137] Among them, KL divergence refers to an asymmetric indicator used to measure the difference between two probability distributions. Specifically, it can be achieved by calculating the relative entropy of the distribution of professional terms in the commentary text and the video event, and is used to quantify the degree of deviation in the semantic matching between the two. Professional term distribution refers to the statistical results of the frequency of occurrence of technical actions, tactical names and event-specific vocabulary related to specific sports in video key frames or commentary text. Specifically, it can be achieved by constructing a domain dictionary and counting the number of times the terms appear within a time window, which is used to reflect the semantic focus of different modal data. Term distribution frequency refers to the probability of occurrence of various professional terms in the key frames of video events or commentary text segments. Specifically, it can be achieved through word frequency statistics and normalization processing within a sliding time window, and is used to establish comparable distribution models.

[0138] Specifically, a pre-built dictionary of specialized terminology covers the technical movements, tactical names, and event-specific vocabulary of the target sport. The action recognition model extracts timestamps from key frames in video events, and the frequency of occurrence of each term in the time window before and after the corresponding moment is counted. After normalization, a video term distribution vector is formed. The generated commentary text is segmented, and then the vocabulary is matched according to the term dictionary, and the term distribution vector within the same time window is counted. The KL divergence value is obtained by calculating the relative entropy between the video term distribution and the commentary text distribution. The smaller the value, the higher the semantic match between the two. This indicator is used to optimize the loss function during model training, ensuring that the generated commentary text maintains semantic synchronization with the video events in terms of the use of specialized terminology.

[0139] Compared to existing technologies, traditional methods rely on general language generation metrics to evaluate commentary quality, but they are unable to effectively measure the accuracy of professional terminology and its relevance to the context. This solution, by introducing a domain-specific dictionary and calculating KL divergence, establishes a targeted evaluation mechanism for sports commentary. This can accurately detect missing, misused, or misplaced professional terminology in commentary content, improving the alignment of the evaluation system with domain requirements.

[0140] Through the above technical solution, this application achieves targeted optimization of the professionalism and scene fit of the generated commentary text, ensuring that the use of terminology accurately corresponds to the video event action in the time dimension, effectively avoiding the problem of empty commentary content caused by missing terminology or timing deviation, and at the same time enhancing the accuracy and credibility of the commentary text in professional dimensions such as tactical analysis and technical action description.

[0141] Optionally, the specific implementation method of step S106 further includes:

[0142] Build a dictionary of professional terms, covering the technical movements, tactical names and event-specific vocabulary of specific sports;

[0143] Count the term distribution frequencies corresponding to key frames in video events;

[0144] Calculate the KL divergence value of the generated commentary text and the video term distribution to quantify the degree of semantic matching between the two.

[0145] The professional terminology dictionary refers to a collection of technical and tactical terms organized according to the characteristics of the sport. This can be constructed by combining expert annotation with the competition rulebook, covering specific terms such as shooting, offside, and tactical fouls. Its purpose is to establish a unified terminology reference standard and ensure the integrity of professional expression in multimodal data.

[0146] Term distribution frequency refers to the frequency of occurrence of specialized terms in events corresponding to key frames in a video. This can be achieved by statistically matching event annotations in video clips with a term dictionary. This metric reflects the core semantic characteristics of the video content and provides a benchmark for subsequent distribution difference calculations. The KL divergence measures the information difference in term distribution between the generated text and the video content, specifically calculated using the asymmetry between probability distributions. This metric objectively quantifies the degree of semantic match between the interpretation text and the video events.

[0147] Specifically, first, a structured term dictionary is constructed based on the target sport field, such as words such as "corner kick", "penalty kick", and "yellow card" in football matches. After event parsing of the video key frames, the frequency of occurrence of each term on the video timeline is counted to form a normalized probability distribution. At the same time, term extraction and frequency statistics are performed on the generated commentary text to obtain the corresponding text distribution. By calculating the KL divergence value of the two distributions, it is possible to accurately measure whether the commentary content accurately covers the key technical points in the video events. For example, if the frequency of "shooting" events in the video is significantly higher than the corresponding term distribution in the commentary text, the KL divergence value will increase, indicating that the generated content has semantic deviations.

[0148] In some specific implementations, the terminology dictionary can be automatically expanded using natural language processing techniques. For example, by using historical match commentary corpora to perform word frequency statistics and keyword extraction, further supplementing specialized expressions for tactical names and player technical moves. Term statistics for key frames in a video can be combined with the classification labels output by the action recognition model to achieve automated annotation. Smoothing can be incorporated into the KL divergence calculation process to prevent zero-probability issues from affecting the evaluation results.

[0149] Compared to existing technologies, traditional methods rely on general language model metrics such as BLEU or ROUGE, which only measure surface-level similarity in text and fail to capture the semantic consistency of specialized terminology. This approach, by constructing a domain-specific terminology system and combining it with quantitative analysis of distribution differences, can effectively identify deviations between commentary content and video events at the technical level. For example, in a basketball game, existing technologies may fail to detect the omission of tactical terms such as "pick-and-roll" and "fast break," while this approach explicitly reveals these semantic omissions through KL divergence calculations.

[0150] Through the above technical solution, this application addresses the problem of traditional evaluation methods' inadequate measurement of professional semantic matching, improving the consistency of commentary content with the technical details of video events. By quantifying differences in term distribution, it is possible to accurately identify semantic deviations in generated commentary, such as incomplete descriptions of tactical actions or inappropriate use of event-specific vocabulary, thereby guiding model optimization and ensuring that the output content conforms to professional expression standards in the sports field.

[0151] See also Figure 4 , which is a schematic diagram of the overall architecture of the principle structure of the sports video commentary generation system based on the multimodal large language model provided in the embodiment of the present application, such as Figure 1 As shown, it mainly includes the following two core components:

[0152] The Multimodal Large Language Model (MLLM) backbone is responsible for perceiving multimodal information from video, audio, and metadata, and generating corresponding commentary text.

[0153] A multimodal clustering memory unit is used to store and retrieve multimodal embeddings, construct a cross-modal joint feature space through contrastive learning, and support the retrieval-augmented in-context learning (RA-ICL) process.

[0154] The above two modules are bridged by a retrieval engine to utilize historical scene instances to optimize the generation effect of the current task.

[0155] See Figure 1 As shown in the figure, 1. Embedding mapping layer: Multimodal inputs (video, audio, metadata) are generated through their respective pre-trained encoders (such as visual encoder ViT-L / 14 and audio encoder LanguageBind-Audio) to generate embedding vectors, which are projected into a shared joint feature space through a multi-layer perceptron (MLP).

[0156] Number of layers: The embedding map uses two fully connected layers, each containing 512 neurons. Activation function: The ReLU activation function is used to improve nonlinear expression capabilities. Regularization: A dropout layer (with a dropout rate of 0.3) is added to avoid overfitting.

[0157] 2. Contrastive Learning Module

[0158] Contrastive loss design: The Masked Margin Softmax (MMS) loss function is used to increase the similarity between the same instance modalities and reduce the similarity between different instances.

[0159] Loss calculation: Losses are calculated for all modality pairs (video-audio, audio-metadata, metadata-video) and full batch comparisons are performed on positive and negative samples to maximize training efficiency.

[0160] Hyperparameters: The value of the margin parameter δ is chosen to be 0.2 to ensure the effectiveness of contrastive learning.

[0161] 3. Clustering Module

[0162] The K-means clustering algorithm is used to group the mixed modal embeddings (weighted average of video, audio and metadata) into k clusters.

[0163] Initial center: Preset the initial cluster center based on the event categories (such as goals, fouls, and timeouts) in all video data to avoid falling into local optimality.

[0164] Update mechanism: Adopt an incremental update strategy to dynamically optimize the cluster center to make it closer to the actual distribution.

[0165] 4. Unimodal alignment, each unimodal embedding is moved closer to its corresponding mixed-modal cluster center through contrastive loss, while moving away from other cluster centers.

[0166] 5. Feature reconstruction: To enhance the generalization ability of the model, two linear layers are used to reconstruct the embedding of each modality, and a reconstruction loss (such as L2 norm loss) is added to the loss function. Regularized training: The reconstruction loss is used to constrain the model to ensure the stability of contrastive learning and clustering training.

[0167] 3. Total loss function

[0168] Combining the contrast loss and information entropy regularization term, the final multimodal alignment loss function is:

[0169] L total =L contrast +λL entropy

[0170] Among them, λ is a hyperparameter used to balance the weights of contrast loss and regularization term, L entropy is the loss function of the information entropy regularization term, L contrast is the contrast loss function.

[0171] 1.1 Contrastive Loss

[0172] The contrast loss is in the form of Masked Margin Softmax (MMS):

[0173]

[0174] Among them, Z vi,mi Embed v for video modal i and metadata embedded in m i Similarity in the joint feature space (such as dot product); δ is the margin hyperparameter used to control the difficulty of comparison; M vi,mj is a mask matrix used to distinguish positive and negative sample pairs, where positive sample pairs are marked as 0 and negative sample pairs are marked as , and B is the number of mask indicators.

[0175]

[0176] Among them, Z x,y is the similarity between sample feature x and sample feature y in the joint feature space.

[0177] 2. Information Entropy Regularization

[0178] In order to enhance the consistency of distribution between modalities, regularization based on information entropy is introduced to limit the consistency of the distribution of modal embedding and joint distribution. The goal is to make each modal distribution P x Keep it similar to the joint modal distribution P.

[0179] 2.1 Joint Distribution Definition

[0180] Assume that the embedding of the three modalities (video, audio, metadata) is represented as z v ,z a ,z m , and their joint distribution P is generated by a weighted combination of all modal embeddings:

[0181]

[0182] Among them, P v 、P a 、P m Represents the joint features of video, audio, and metadata respectively, and each modal distribution P x Generated from normalized embeddings.

[0183] 2.2 Information Entropy Regularization Term

[0184] Use cross entropy to measure the single mode distribution P x The similarity with the joint distribution P is as follows:

[0185]

[0186] Among them, v, a, and m are video, audio, and metadata respectively, and N is the total number of features.

[0187] 3. Total loss function

[0188] Combining the contrast loss and information entropy regularization term, the final multimodal alignment loss function is:

[0189] L total =L contrast +λL entropy

[0190] Among them, λ is a hyperparameter used to balance the weights of contrast loss and regularization term, L entropy is the loss function of the information entropy regularization term, L contrast is the contrast loss function.

[0191] See also Figure 5 , is a structural diagram of a multimodal clustering memory unit provided in an embodiment of the present application, and is described in detail as follows:

[0192] The multimodal clustering memory unit is the core module of this technical solution, which is used to optimize, store and retrieve the joint features of multimodal data (video, audio, metadata). The specific implementation steps are as follows:

[0193] Feature Mapping and Alignment

[0194] Embeddings from video, audio, and metadata are separately projected into a unified joint representation space.

[0195] The projection process is completed through a multi-layer perceptron (MLP), and the corresponding mapper aligns the features into a compatible format, ensuring that features from different modalities can be compared and fused.

[0196] Multimodal contrastive learning

[0197] Example of a joint representation space:

[0198] In the SoccerComment framework, the features of video, audio and metadata are encoded as embedding representations (z v 、z a 、z m ), and then optimize their similarity in the joint space via contrastive learning.

[0199] These modal features are not simply “stored in rows”, but are integrated through multimodal fusion (e.g., through z mix =(z v +z a +z m ) / 3 to form a unified embedding representation, which is then organized into semantically consistent clusters via K-means clustering.

[0200] The coupling between modes is expressed as:

[0201] Different modalities are coupled through a shared representation space. For example, contrastive loss can be used to shorten the distance between modalities of the same instance while increasing the distance between modalities of different instances.

[0202] This representation enables the model to retrieve and match in the joint space without relying on features of the original modalities.

[0203] Comparative learning strategies and time correlation

[0204] The Masked Margin Softmax (MMS) loss used in the paper combines all negative samples in the batch, making contrastive learning more efficient. Specifically:

[0205] Comparison between current task and historical examples: In SoccerComment, the system performs semantic matching by comparing the embedding of the current scene with the embedding of similar scenes in the historical database. This approach not only improves the retrieval ability of the current task but also enriches the contextual learning.

[0206] Time correlation: For example, by retrieving scenes of similar time segments (15 seconds before and after) to build historical context instances.

[0207] Connection to contextual learning

[0208] Retrieval-enhanced contextual learning is closely related to contrastive learning:

[0209] Contrastive learning provides an optimized representation space, making the retrieval of multimodal features more efficient;

[0210] Contextual learning (such as RA-ICL) integrates historical context into the generation process by retrieving similar instances, thereby improving the diversity and accuracy of generated content.

[0211] For example, semantic clustering

[0212] Use clustering algorithms (such as K-means) to semantically group multimodal features and generate cluster centers.

[0213] These cluster centers are used to represent the core features of similar scenes and provide support for the model's retrieval mechanism.

[0214] For example, feature reconstruction

[0215] In order to further stabilize the feature representation, the feature reconstruction module regenerates and optimizes the original features to ensure that the feature distribution of the model in the joint representation space is robust.

[0216] This approach can capture detailed features that may be missed in contrastive learning while avoiding feature conflicts between modalities.

[0217] For example, comprehensive loss optimization

[0218] The entire module is jointly optimized via contrastive learning loss (to enhance features between similar modalities) and reconstruction loss (to ensure stability).

[0219] The model is continuously adjusted during training, making the multimodal features in the joint representation space more compact and semantically rich.

[0220] Retrieval-augmented contextual learning (RA-ICL) extension

[0221] The core of RA-ICL is to enhance the generated context through multimodal feature retrieval. To optimize retrieval efficiency and context diversity, a distance metric based on sparse regularization is introduced:

[0222] d=||qk||1+β||k-μ||1

[0223] Where q is the feature of the query instance; k is the feature of the multimodal instance in the database; μ is the mean vector of the multimodal feature distribution, which is used for centering; ||·||1 both represent the L1 norm, which increases sparsity; β is the weight that controls the sparsity regularization.

[0224] The optimized retrieval mechanism ensures that the retrieved instances are not only similar to the query instances but also representative and diverse.

[0225] Clustering and multimodal reconstruction enhancement

[0226] In order to further improve the clustering effect and generation stability, an enhanced clustering loss based on target similarity is defined

[0227]

[0228] Among them, u c is the center of the cluster to which the current instance belongs; crth To control the degree of orthogonality between cluster centers and avoid overlapping clusters; (u i ·u j ) 2 is the inner product square between cluster centers, used to constrain orthogonality, Z i It represents any feature in the video, audio and metadata.

[0229] The multimodal reconstruction loss is expanded to:

[0230]

[0231] Where: z m It is a reconstruction feature; is the mean of all modal features and is used to constrain inter-modal consistency; is the reconstruction loss between the balanced intra-modality and inter-modality, M is the number of reconstructed features, and Z is a vector or matrix in the latent space, representing the abstract features of the input data.

[0232] Through the above technical solutions, this application solves the defects of existing evaluation methods in semantic consistency and domain knowledge adaptation. For example, in a football game scenario, if the commentary text does not accurately mention "long shots outside the penalty area" but only describes "wonderful shots", traditional indicators may give high scores due to language fluency, but this solution identifies the lack of key information through differences in term distribution and gives low scores in the relevance dimension, thereby promoting the generation model to optimize the use of professional terms. At the same time, the consistency dimension ensures the accurate embedding of metadata such as player names and game times, avoiding the common information dislocation problems in existing technologies. The logical dimension further verifies the causal relationship of events, such as the reasonable correlation between shooting actions and score changes, to enhance the authenticity of the commentary and the audience experience.

[0233] This application further proposes a multimodal clustering memory unit that uses the K-means clustering algorithm to group mixed modal embeddings to form semantically consistent clusters, optimizes feature alignment between modalities through contrastive learning and information entropy regularization, and ensures the consistency of semantic relationships between video, audio, and metadata; the RA-ICL mechanism retrieves historical instances related to the current task and inputs them as references into the MLLM, which is combined with the multimodal input of the current task to provide rich contextual support and generate commentary content with high relevance and diversity based on this contextual information.

[0234] Among them, the multimodal clustering memory unit refers to a storage module that groups mixed modal embeddings through a clustering algorithm. Specifically, the K-means algorithm can be used to divide the joint embedding space of video, audio and metadata into semantically consistent clusters. The similarity measure between samples of different modalities is optimized through comparative learning, and the compactness and diversity of the clustering process are constrained by information entropy regularization. This unit realizes the effective alignment of features between modalities by constructing a cross-modal joint feature space.

[0235] Among them, the RA-ICL mechanism refers to a retrieval-enhanced contextual learning process. Specifically, a sparse regularized distance metric model can be used to screen multimodal semantic fragments in the historical database. By calculating the semantic similarity between the current scene embedding and the historical instance, the most relevant commentary fragment is dynamically selected as the reference context of the generation process. This mechanism enhances the style change ability of the generated text by reusing the diverse expression patterns of historical commentaries.

[0236] Specifically, the multimodal clustering memory unit first maps video frames, audio waveforms, and metadata into a shared embedding space using a pre-trained encoder to form a mixed-modal feature vector. The K-means algorithm is then used to cluster the feature vectors, with each cluster representing a scene type with similar semantics. By contrasting the learning loss function, samples of the same category from different modalities are forced to be close in the embedding space. An information entropy regularization term is introduced to constrain the distribution balance of the clustering results and prevent overfitting of single-modal features. In the RA-ICL mechanism, the top K historical instances with the highest semantic relevance are selected by calculating the sparse regularized distance between the current scene embedding and the historical cluster center. Their corresponding commentary text is then input into the generative model as dynamic context. During the commentary generation phase, the MLLM backbone simultaneously processes the current multimodal input and the retrieved context fragments. Using an attention mechanism, it fuses real-time scene information with historical expression patterns to generate commentary that is both consistent with the details of the current event and exhibits linguistic style variations.

[0237] Compared with existing technologies, traditional methods typically use single-modal feature extraction or simple feature concatenation and fusion, failing to establish a cross-modal joint semantic space. This results in weak contextual relevance and a single expression pattern in the generated commentary. This solution constructs structured multimodal feature clusters, precisely aligns semantic information from different modalities, and utilizes a retrieval enhancement mechanism to incorporate diverse expressions into historical commentary, significantly improving the contextual relevance and linguistic richness of the commentary.

[0238] Through the above technical solution, this application effectively solves the problem of monotonous commentary content caused by insufficient cross-modal feature alignment. In a football match scenario, when a player's shot is detected, the system quickly matches historical shot event clusters through clustered memory units, retrieves commentary clips containing terms such as "long shot" and "threat", and combines them with the current game's real-time score and audience cheers to generate commentary text that accurately describes the shot details while incorporating diverse styles of expression. For example, while retaining the core message of "Messi's long shot is threatening", it dynamically selects different expressions such as "wonderful volley" and "death blow".

[0239] Specifically, the system first extracts video frames, audio waveforms, and metadata from video events. A multimodal feature encoder is then used to identify and count a set of sports-related terminology, such as "long shot" and "outside the penalty area" in football matches. Simultaneously, the generated commentary text undergoes semantic parsing to extract the distribution of terms within the same set of terms. The KL divergence between the two distributions is calculated to measure the difference in term coverage between the generated text and the video event. A smaller difference indicates greater semantic consistency. Furthermore, multimodal contextual association analysis is incorporated into the term distribution construction process. For example, a cross-modal mapping is established between the shooting action in the video and the term "long shot" in the commentary to assess the logical relevance of the commentary content to the scene. By jointly analyzing the completeness of term distribution coverage, the accuracy of semantic mapping, and textual coherence, a multi-dimensional evaluation is achieved, ranging from single-term matching to the overall semantic distribution.

[0240] Compared with existing technologies, traditional methods rely on manual rules or simple word frequency statistics to evaluate commentary quality, focusing only on word repetition rate while ignoring the contextual relevance and cross-modal consistency of terms. This solution quantifies term distribution differences through KL divergence, not only capturing vocabulary coverage but also identifying semantic drift, such as deviations in the distribution of terms in the commentary text that are unrelated to the video events. Furthermore, existing technologies are unable to effectively combine multimodal information to establish an evaluation benchmark. This solution, by extracting term distributions from video events, integrates visual, audio, and textual information into a unified evaluation framework, ensuring that the commentary content strictly corresponds to the actual scene.

[0241] Through the above technical solution, this application can establish a cross-modal evaluation benchmark for professional terminology in the sports field. By using statistical distribution differences to objectively measure the quality of generated commentary in terms of key information coverage, semantic consistency, and logical relevance, it addresses the problem that traditional general indicators are insufficiently sensitive to domain characteristics. For example, in a football match scenario, it can accurately detect whether the generated commentary accurately uses terms such as "outside the penalty area" and "long shot" and determine their spatiotemporal logical relationship with video events, thereby avoiding semantic bias or information omissions and improving the comprehensiveness and pertinence of the evaluation process.

[0242] See also Figure 3 , a sports video commentary generation system based on a multimodal large language model provided in this application, comprising:

[0243] The data set module 301 is used to obtain a multimodal data set, the data set including sports videos, and audio and commentary text corresponding to the sports videos, and the sports videos, audio and commentary text are calibrated based on timestamps;

[0244] A feature fusion module 302 is used to build a multimodal large language model, encode sports video, audio and commentary text, so that the corresponding video frames, audio waveforms and metadata are projected into a shared embedding space to determine a multimodal embedding vector;

[0245] A feature alignment module 303 is used to set up a multimodal clustering memory unit, group the multimodal embedding vectors, and optimize the feature alignment between modalities through contrastive learning and information entropy regularization;

[0246] The retrieval enhancement module 304 uses the sparse regularized distance metric to retrieve historical instances as reference inputs for the current input multimodal embedding vector based on the retrieval enhancement context learning mechanism;

[0247] The commentary generation module 305 is used to jointly input the current input multimodal embedding vector and the reference input into the multimodal large language model to obtain the sports video commentary.

[0248] In addition, it also includes: an evaluation module 306, which uses a KL divergence-based professional terminology distribution comparison evaluation method to comprehensively evaluate the relevance, consistency, fluency and logic of the content by quantifying the differences in the distribution of professional terminology in the interpretation text and video events.

[0249] Regarding the specific limitations of the sports video commentary generation system based on a multimodal large language model, please refer to the limitations of the sports video commentary generation method based on a multimodal large language model above, which will not be repeated here. The various modules in the above-mentioned sports video commentary generation device based on a multimodal large language model can be implemented in whole or in part by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the electronic device in the form of hardware, or can be stored in the memory of the electronic device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0250] In this embodiment, the sports video commentary generation system based on a multimodal large language model is essentially equipped with multiple modules to execute the sports video commentary generation method based on a multimodal large language model in any of the above embodiments. The specific functions and technical effects can be referred to the above embodiments and will not be repeated here.

[0251] In one embodiment, an electronic device is provided. The electronic device may be a server, and its internal structure may be as shown in FIG. Figure 6As shown. The electronic device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, the functions or steps on the server side of the above method are implemented.

[0252] In one embodiment, an electronic device is provided. The electronic device may be a client, and its internal structure diagram may be as follows: Figure 7 As shown. The electronic device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, the functions or steps of the client side of the above method are implemented.

[0253] In one embodiment, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0254] A multimodal dataset is obtained, which includes sports videos, as well as audio and commentary text corresponding to the sports videos; a multimodal large language model is constructed, and the sports videos, audio and commentary text are encoded so that the corresponding video frames, audio waveforms and metadata are projected into a shared embedding space to determine the multimodal embedding vector; a multimodal clustering memory unit is set to group the multimodal embedding vectors, and feature alignment between modalities is optimized through contrastive learning and information entropy regularization; based on the retrieval-enhanced contextual learning mechanism, historical instances are retrieved through sparse regularized distance metric as reference input for the current input multimodal embedding vector; the current multimodal embedding vector and the reference input are jointly input into the multimodal large language model to obtain the sports video commentary.

[0255] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0256] A multimodal dataset is obtained, which includes sports videos, as well as audio and commentary text corresponding to the sports videos; a multimodal large language model is constructed, and the sports videos, audio and commentary text are encoded so that the corresponding video frames, audio waveforms and metadata are projected into a shared embedding space to determine the multimodal embedding vector; a multimodal clustering memory unit is set to group the multimodal embedding vectors, and feature alignment between modalities is optimized through contrastive learning and information entropy regularization; based on the retrieval-enhanced contextual learning mechanism, historical instances are retrieved through sparse regularized distance metric as reference input for the current input multimodal embedding vector; the current multimodal embedding vector and the reference input are jointly input into the multimodal large language model to obtain the sports video commentary.

[0257] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or electronic device can be referred to the relevant descriptions on the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0258] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the above-mentioned computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0259] Those skilled in the art will clearly understand that for the sake of convenience and brevity in description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the above-mentioned device or system can be divided into different functional units or modules to complete all or part of the functions described above.

[0260] The embodiments provided above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for generating sports video commentary based on a multimodal large language model, characterized in that: include: Acquire a multimodal data set, the data set including a sports video, and audio and commentary text corresponding to the sports video, wherein the sports video, the audio, and the commentary text are calibrated based on timestamps; Constructing a multimodal large language model, encoding the sports video, the audio, and the commentary text so that corresponding video frames, audio waveforms, and metadata are projected into a shared embedding space, and determining a multimodal embedding vector; Setting a multimodal clustering memory unit to group the multimodal embedding vectors and optimizing feature alignment between modalities through contrastive learning and information entropy regularization; Based on the retrieval-enhanced contextual learning mechanism, historical instances are retrieved through a sparse regularized distance metric as reference inputs for the multimodal embedding vector currently input; The current input multimodal embedding vector and the reference input are jointly input into the multimodal large language model to obtain sports video commentary.

2. The method according to claim 1, characterized in that Obtain multimodal datasets, including: Preprocessing the data set, including cleaning, normalization, and enhancement; Dividing the sports video into events based on time stamps to determine video segments, wherein the events include goals, fouls, and timeouts; Performing action recognition on the video clip to determine classification labels, wherein the classification labels include shooting, saving and passing; The audio is transcribed word for word and aligned with the video clip according to the timestamp to ensure semantic consistency; The player information, venue information and event status in the commentary text are annotated so that the sports video, the audio and the commentary text in the data set maintain semantic consistency based on timestamps.

3. The method according to claim 1, characterized in that Encoding the sports video, the audio, and the commentary text so that corresponding video frames, audio waveforms, and metadata are projected into a shared embedding space, and determining a multimodal embedding vector includes: Encode the sports video, the audio, and the commentary text using a pre-trained encoder to generate a first embedding vector, a second embedding vector, and a third embedding vector, respectively; The first embedding vector, the second embedding vector, and the third embedding vector are projected into a shared embedding space through a fully connected layer to determine a joint multimodal embedding vector.

4. The method according to claim 1, wherein The multimodal clustering memory unit performs the following operations: Establish a cross-modal joint feature space to store multimodal embedding data; The similarity weights between features of different modalities are calculated through contrastive learning algorithms, and full batch comparisons are performed on positive and negative sample pairs to maximize training efficiency. Information entropy regularization constraints are imposed when dynamically updating cluster centers, and cluster centers are generated through semantic grouping to maintain semantic relationship consistency.

5. The method according to claim 4, characterized in that The multimodal clustering memory unit further performs the following operations: Each unimodal embedding is moved closer to its corresponding mixed-modal cluster center through contrastive loss to achieve unimodal alignment while moving away from other cluster centers; Each unimodal embedding is reconstructed and a reconstruction loss function is added to the loss function to constrain the reconstruction loss function.

6. The method according to claim 1, characterized in that The retrieval-enhanced context learning mechanism includes: Build a multimodal feature retrieval index library to store multimodal embeddings of historical scenes and corresponding explanatory texts; Design a similarity measurement function based on sparse coding to calculate the correlation between the current scene and historical instances; The explanation segments of the top K most relevant instances are selected as contextual prompts and input into the multimodal large language model.

7. The method according to claim 1, characterized in that The expression of the retrieval enhanced context learning mechanism includes: d = ||qk||1+β||k-μ||1, where q is the feature of the query instance, k is the feature of the multimodal instance in the database, μ is the mean vector of the multimodal feature distribution, which is used for centering; ||·||1 represents the L1 norm, which increases sparsity; and β is the weight that controls the sparsity regularization.

8. The method according to any one of claims 1 to 7, characterized in that: Also includes: The KL divergence-based comparative evaluation method for professional terminology distribution quantifies the differences in the distribution of professional terminology between interpretation text and video events, and comprehensively evaluates the relevance, consistency, fluency, and logic of the content.

9. The method according to claim 8, characterized in that A comparative evaluation method for terminology distribution based on KL divergence quantifies the differences in terminology distribution between text and video events, including: Build a dictionary of professional terms, covering the technical movements, tactical names and event-specific vocabulary of specific sports; Count the term distribution frequencies corresponding to key frames in video events; Calculate the KL divergence value of the generated commentary text and the video term distribution to quantify the degree of semantic matching between the two.

10. A sports video commentary generation system based on a multimodal large language model, characterized in that: include: A data set module is used to obtain a multimodal data set, wherein the data set includes a sports video, and audio and commentary text corresponding to the sports video, wherein the sports video, the audio and the commentary text are calibrated based on timestamps; a feature fusion module for constructing a multimodal large language model, encoding the sports video, the audio, and the commentary text so that corresponding video frames, audio waveforms, and metadata are projected into a shared embedding space, and determining a multimodal embedding vector; A feature alignment module is used to set up a multimodal cluster memory unit, group the multimodal embedding vectors, and optimize the feature alignment between modalities through contrastive learning and information entropy regularization; A retrieval enhancement module, based on a retrieval enhancement context learning mechanism, retrieves historical instances as reference inputs for the current input multimodal embedding vector via a sparse regularized distance metric; The commentary generation module is used to jointly input the current input multimodal embedding vector and the reference input into the multimodal large language model to obtain a sports video commentary.

Citation Information

Cited By

  • Marine ranch intelligent decision-making system based on self-playing chess

    CN121189166A

  • Multi-modal feature alignment method and device for heterogeneous data and medium

    CN121527790A

  • Multi-modal large model-based cross-device content synchronization and adaptive display method

    CN121680768A

  • Multi-modal large language model machine forgetting method based on causal orthogonal decoupling

    CN121809559A

  • Grain yield prediction method and device based on large model and time sequence retrieval and medium

    CN122019798A