Football commentary generation method and device based on visual reasoning and knowledge enhancement
By employing a two-stage processing framework of visual reasoning and knowledge enhancement, the problem of end-to-end models lacking fine-grained visual information and contextual knowledge in football commentary is solved, generating more accurate and information-rich football commentary and simulating the workflow of human commentators.
Patent Information
- Application Number
- CN202510964253.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-07-14
AI Technical Summary
Existing end-to-end deep learning models lack fine-grained visual information understanding and rich contextual knowledge in football commentary generation, resulting in commentary that cannot accurately reflect the specific details and background of the real-time game.
A two-stage processing framework based on visual reasoning and knowledge enhancement is adopted. First, the player and team references in the anonymized commentary are aligned through visual reasoning, and fine-grained visual information and game status information are used for detail enhancement. Then, knowledge enhancement is carried out by combining internal and external knowledge bases to generate more natural and information-rich commentary.
Improve the accuracy and information richness of commentary, enabling commentators to accurately reference entities, understand a broader context of the game, and generate insightful explanations and commentary similar to those of professional television commentators by incorporating historical statistics.
Smart Images

Figure CN120471178B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of multimodal artificial intelligence technology, and in particular to a method and apparatus for generating football commentary based on visual reasoning and knowledge enhancement. Background Technology
[0002] With the development of artificial intelligence technology, automated commentary generation technology has been widely used in many fields, especially in sports events.
[0003] Traditional automated football commentary generation methods often rely on end-to-end deep learning models, taking a football match video clip of about half a minute as input and outputting a natural language commentary for the video football clip.
[0004] However, such end-to-end models often lack fine-grained visual information understanding and rich contextual knowledge, resulting in generated commentary that cannot accurately reflect the specific details and background of the real-time match. Summary of the Invention
[0005] The purpose of this application is to provide a method and apparatus for generating football commentary based on visual reasoning and knowledge enhancement. Through a two-stage processing framework of entity alignment based on visual reasoning and commentary generation based on knowledge enhancement, it can better simulate the workflow of human commentators and generate more natural and information-rich commentary.
[0006] This application provides a method for generating football commentary based on visual reasoning and knowledge enhancement, including:
[0007] The process involves: acquiring anonymized, coarse-grained commentary information output by an end-to-end model based on a target video segment; extracting match status information and fine-grained visual information from the target video segment, and using the end-to-end model to obtain event location guidance information; the fine-grained visual information being used to represent detailed information in the match; performing visual reasoning based on the match status information, the fine-grained visual information, and the event location guidance information, and enhancing the anonymized, coarse-grained commentary information based on the visual reasoning results to obtain enhanced commentary information; the visual reasoning results including: identifying and corresponding player and team designations in the match, and matching the objects of each event; and performing knowledge enhancement on the enhanced commentary information based on an internal and external knowledge base to obtain knowledge-enhanced football commentary; wherein the external knowledge base includes: historical statistical data of different football matches; and the internal knowledge base includes: static background information and dynamically updated events of the current football match.
[0008] Optionally, the match status information includes: temporal team lineup information, key event timeline, and historical timeline; the extraction of match status information based on the target video clip includes: generating the temporal team lineup information, the key event timeline, and the historical timeline based on the target video clip and historical video clips; updating the temporal team lineup information based on the card-playing event and the player substitution event represented by the historical timeline; wherein, the temporal team lineup information is used to represent the player information of each team currently on the field; the key event timeline is used to represent the timestamps of goal events and card-playing events; the historical timeline consists of multiple recent commentaries and is used to represent the most recently occurring events.
[0009] Optionally, the fine-grained visual information includes: player information, jersey information, and team alias information; the extraction of fine-grained visual information based on the target video segment includes: classifying each shot in the target video segment using a visual language model to obtain shot classification information containing each shot category; the shot categories include: close-up shots, medium shots, long shots, and off-field shots; based on the player information contained in the temporal team lineup information, filtering out the facial images of each player from the database to obtain a set of facial images containing all player facial images, and using image recognition technology to identify facial images that match the facial images in the set of facial images from the medium-close shots of the target video segment to determine the players contained in the target video segment; and using optical character recognition technology to identify the jerseys appearing in the close shots of the target video segment to obtain the jersey information and team alias information of the players contained in the target video segment; wherein, the medium-close shots include: close-up shots and / or, medium shots.
[0010] Optionally, the step of obtaining event localization guidance information using the end-to-end model includes: obtaining the query of each cross-attention layer teammate of the end-to-end model, and calculating the average cross-attention between the query and the video frame to obtain the attention weight of each query; aggregating the attention weights of each query to obtain a frame-level importance vector, and obtaining the event localization guidance information based on the frame-level importance vector; wherein, the frame-level importance vector is used to characterize the importance weight of each video frame in the target video segment when generating anonymization and explanation information.
[0011] Optionally, the step of performing visual reasoning based on the match status information, the fine-grained visual information, and the event location guidance information, and then enhancing the details of the anonymized, coarse-grained commentary information according to the visual reasoning results to obtain enhanced commentary information, includes: performing visual reasoning based on the key event timeline, the historical timeline, the shot classification information, the facial image set, the players included in the target video clip, the jersey information, and the participating team alias information, and then enhancing the details of the anonymized, coarse-grained commentary information according to the visual reasoning results to obtain enhanced commentary information; wherein, the visual reasoning results are obtained based on a multimodal visual model.
[0012] Optionally, the step of enhancing the detailed commentary information based on internal and external knowledge bases to obtain knowledge-enhanced football commentary includes: using natural language questions as queries, based on the methods different commentators use to apply relevant knowledge during the commentary process, obtaining historical match data related to the current match from the external knowledge base through a language model, and obtaining real-time updated match status based on context learning from the internal knowledge base; supplementing and correcting the detailed commentary information based on the historical match data and the real-time updated match status to obtain the knowledge-enhanced football commentary.
[0013] This application also provides a football commentary generation device based on visual reasoning and knowledge enhancement, including:
[0014] The system comprises the following modules: an acquisition module for acquiring anonymized, coarse-grained commentary information output by an end-to-end model based on a target video segment; an information extraction module for extracting match status information and fine-grained visual information from the target video segment, and acquiring event location guidance information using the end-to-end model; the fine-grained visual information is used to represent detailed information in the match; a detail enhancement module for performing visual reasoning based on the match status information, the fine-grained visual information, and the event location guidance information, and enhancing the anonymized, coarse-grained commentary information based on the visual reasoning results to obtain enhanced commentary information; the visual reasoning results include: identifying and corresponding player and team designations in the match, and matching the objects of various events; and a knowledge enhancement module for enhancing the enhanced commentary information based on an internal and external knowledge base to obtain knowledge-enhanced football commentary; wherein the external knowledge base includes: historical statistical data of different football matches; and the internal knowledge base includes: static background information and dynamically updated events of the current football match.
[0015] Optionally, the match status information includes: temporal team lineup information, key event timeline, and historical timeline; the information extraction module is specifically used to generate the temporal team lineup information, key event timeline, and historical timeline based on the target video clip and historical video clips; the information extraction module is further used to update the temporal team lineup information based on the card-playing event and the player replacement event represented by the historical timeline; wherein, the temporal team lineup information is used to represent the player information of each team currently on the field; the key event timeline is used to represent the timestamps of goal events and card-playing events; the historical timeline consists of multiple recent commentaries and is used to represent the most recently occurred events.
[0016] Optionally, the fine-grained visual information includes: player information, jersey information, and participating team alias information; the information extraction module is specifically used to classify each shot in the target video segment using a visual language model to obtain shot classification information containing each shot category; the shot categories include: close-up shots, medium shots, long shots, and off-field shots; the information extraction module is further used to filter out the facial images of each player from the database based on the player information contained in the temporal team lineup information to obtain a set of facial images containing all player facial images, and to use image recognition technology to identify facial images that match the facial images in the set of facial images from the medium and close-up shots of the target video segment to determine the players included in the target video segment; the information extraction module is further used to identify the jerseys appearing in the close-up shots of the target video segment based on optical character recognition technology to obtain the jersey information and participating team alias information of the players included in the target video segment; wherein, the medium and close-up shots include: close-up shots and / or, medium shots.
[0017] Optionally, the information extraction module is specifically used to acquire the query of each cross-attention layer teammate of the end-to-end model, and calculate the average cross-attention between the query and the video frame to obtain the attention weight of each query; the information extraction module is further used to aggregate the attention weight of each query to obtain a frame-level importance vector, and obtain the event localization guidance information based on the frame-level importance vector; wherein, the frame-level importance vector is used to characterize the importance weight of each video frame in the target video segment when generating anonymization and explanation information.
[0018] Optionally, the detail enhancement module is specifically used to perform visual reasoning based on the key event timeline, the historical timeline, the shot classification information, the facial image set, the players contained in the target video clip, the jersey information, and the alias information of the participating teams, and to enhance the details of the anonymized, coarse-grained explanatory information according to the visual reasoning results, so as to obtain detailed-enhanced explanatory information; wherein, the visual reasoning results are obtained based on a multimodal visual model.
[0019] Optionally, the knowledge enhancement module is specifically used to, based on the methods by which different commentators apply relevant knowledge during the commentary process, use natural language questions as queries, obtain historical match data related to the current match from the external knowledge base through a language model, and obtain real-time updated match status based on context learning from the internal knowledge base; supplement and correct the detailed commentary information based on the historical match data and the real-time updated match status to obtain the knowledge-enhanced football commentary.
[0020] This application also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the football commentary generation method based on visual reasoning and knowledge enhancement as described above.
[0021] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the football commentary generation method based on visual reasoning and knowledge enhancement as described above.
[0022] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the football commentary generation method based on visual reasoning and knowledge enhancement as described above.
[0023] The football commentary generation method and apparatus based on visual reasoning and knowledge enhancement provided in this application first obtains anonymized, coarse-grained commentary information output by an end-to-end model based on a target video segment; then, it extracts match status information and fine-grained visual information based on the target video segment, and uses the end-to-end model to obtain event location guidance information; the fine-grained visual information is used to represent detailed information in the match; and visual reasoning is performed based on the match status information, the fine-grained visual information, and the event location guidance information, and the anonymized, coarse-grained commentary information is enhanced with details according to the visual reasoning results to obtain detailed-enhanced commentary information; the visual reasoning results include: identifying and corresponding player and team references in the match, and matching the objects of each event; finally, knowledge enhancement is performed on the detailed-enhanced commentary information based on an internal knowledge base and an external knowledge base to obtain knowledge-enhanced football commentary; wherein, the external knowledge base includes: historical statistical data of different football matches; the internal knowledge base includes: static background information and dynamically updated events of the current football match. Thus, by using a two-stage processing framework of entity alignment based on visual reasoning and narration generation based on knowledge enhancement, we can better simulate the workflow of human narrators and generate more natural and information-rich narrations. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is one of the flowcharts of the football commentary generation method based on visual reasoning and knowledge enhancement provided in this application;
[0026] Figure 2 This is the second flowchart of the football commentary generation method based on visual reasoning and knowledge enhancement provided in this application;
[0027] Figure 3 This is a schematic diagram of the structure of the football commentary generation device based on visual reasoning and knowledge enhancement provided in this application;
[0028] Figure 4 This is a schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0030] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0031] Professional football commentary possesses several key elements that attract viewers: ① Accurate referencing of entities: Commentators use specific details of players and teams seen on camera, appropriately connecting them to the unfolding events to help viewers understand what is happening. ② Understanding the broader context of the match: Commentators cannot focus solely on the current football event; they must seamlessly integrate background knowledge into the entire match, providing a comprehensive interpretation of the current state of the game. Commentary segments should be continuous, not fragmented, to maintain narrative coherence. ③ Utilizing historical statistics: Well-crafted commentary should include not only description but also explanation and analysis. Incorporating historical statistics forms the basis for these broader perspectives and deeper insights.
[0032] However, previous football commentary models failed to meet the above requirements, mainly due to the following problems: Problem 1: The generated dialogue used anonymous player and team tags as placeholders, failing to pinpoint specific players in the event; Problem 2: Relying on a series of NLP metrics for evaluation often led to context-dependent errors, such as incorrect score announcements after a goal, which is the most crucial information in the entire match; Problem 3: Traditional football commentary relied on live text communication as its data source, essentially focusing on descriptive commentary because the audience lacked access to video. Therefore, compared to live television commentary, these commentaries lacked the integration of statistical knowledge, while live television commentary incorporated visualizations to enhance viewer understanding.
[0033] To address the three key differences between end-to-end social ceramics interpretation models and human-centered television interpretation, embodiments of this application propose a GS system, such as... Figure 1 As shown, this is a two-stage model that treats football commentary generation as a knowledge-enhanced visual reasoning task. Unlike previous end-to-end models, GS is a two-stage system that can generate football commentary, treating it as a knowledge-enhanced visual reasoning task. It achieves entity-aligned visual reasoning in the first stage, and uses knowledge to correct and revise anonymous comments in the second stage.
[0034] like Figure 1 As shown, in the first stage, GS aligns the anonymous commentary output by the end-to-end commentary model with specific player and team names. Pre-experiments of humans recognizing players while watching live sports show that humans rely on various cues to understand details on the football field. In long-view scenes, only 15.5% of the players relevant to the event can be identified solely through player tracking, while the others can be identified through rich context and visual cues. Inspired by this, embodiments of this application establish internal game context construction and novel granular shot analysis modules to reconstruct contextual information in specific game states and uncover visual details that the end-to-end communication generation model cannot capture. GS introduces supervised tuning with answer-guided reasoning strategies and grouping relative strategy optimization on the video large language model, improving the entity alignment accuracy on the SN-Caption-test-align benchmark to 71.1%, surpassing Gemini 2.0 Flash-Lite's 44.3%. It effectively links coarse anonymous commentary with players involved in specific actions, thereby advancing to realistic commentary with precise entity references. In the second phase, GS enriches entity-aligned commentary by combining external and internal football match knowledge, moving beyond mere descriptive content and thus providing a solution to problems 2 and 3 mentioned above. Analysis of television commentator practices reveals their frequent and flexible application of relevant knowledge. Based on this behavior, embodiments of this application develop a Football Knowledge Augmentation Generation (KAG) system, enabling the model to query external historical statistical data and an iteratively updated internal database of match contextual knowledge, reviewing highlights of the current match through contextual learning. These modules allow GS to not only provide accurate, context-sensitive commentary but also insightful explanations and commentary consistent with professional television commentary.
[0035] The following description, in conjunction with the accompanying drawings, details the football commentary generation method based on visual reasoning and knowledge enhancement provided in this application through specific embodiments and application scenarios.
[0036] like Figure 2As shown in the embodiment of this application, a football commentary generation method based on visual reasoning and knowledge enhancement is provided. This method may include the following steps 201 to 203:
[0037] Step 201: Obtain the anonymized, coarse-grained explanatory information output by the end-to-end model based on the target video segment.
[0038] For example, the football commentary generation method based on visual reasoning and knowledge enhancement provided in this application is an anonymized, coarse-grained commentary synthesized based on the existing end-to-end football model (i.e., the aforementioned end-to-end model). That is, based on the anonymized, coarse-grained commentary generated by the end-to-end model, visual reasoning technology is used to generate entity-aligned commentaries with accurate player and team references. Furthermore, based on the analysis of entity states and football events, combined with information retrieval from internal and external knowledge bases such as dynamic match context information and historical statistical data, knowledge-enhanced commentaries are generated.
[0039] Step 202: Extract competition status information and fine-grained visual information based on the target video segment, and use the end-to-end model to obtain event localization guidance information.
[0040] The fine-grained visual information is used to characterize the details in the competition.
[0041] For example, after obtaining the anonymized, coarse-grained commentary generated by the end-to-end model, the first stage of visual reasoning and entity alignment can be carried out. In the first stage, by building a refined game state information database and video clue analysis system, the visual big model is trained to perform anonymous entity reasoning in order to identify and correspond the players and teams referred to in the game, realize the specific objects of the commentary events, and make the commentary have practical meaning.
[0042] For example, the first phase of GS primarily leverages various cues to align anonymous entities (including players and teams) in sports match footage. Alignment is achieved through visual reasoning, which includes contextual cues derived from reconstructing the context of the match state, visual cues from fine-grained shot analysis, and prior knowledge guidance extracted from an end-to-end football foundational model.
[0043] For example, before performing visual reasoning, it is also necessary to reconstruct the match state information, which includes: temporal team lineup information, key event timeline, and historical timeline.
[0044] Specifically, the step 202 above, which involves extracting the match status information based on the target video segment, may further include the following steps 202a1 and 202a2:
[0045] Step 202a1: Generate the temporal team lineup information, the key event timeline, and the historical timeline based on the target video clip and historical video clips.
[0046] Step 202a2: Update the temporal team lineup information based on the card-playing events and the player replacement events represented by the historical timeline.
[0047] The temporal team lineup information is used to represent the player information of each team on the field at present; the key event timeline is used to represent the timestamps of goal events and card-playing events; the historical timeline consists of multiple recent commentaries and is used to represent the most recent events.
[0048] For example, reconstructing the match state based on event timestamps and match progress helps to accurately identify the range of players involved in the event and the current match situation. Taking a full 90-minute football match (i.e., the target video segment mentioned above) as the object, key events affecting the match state, such as substitutions, goals, and refereeing decisions, are identified. Based on timestamps and event sequence, an internal information database of the match state is constructed and dynamically maintained and updated to obtain the on-field lineup and key situations at the time of the segment to be narrated, thus narrowing the scope of inference.
[0049] For example, the contextual game state is updated repeatedly throughout the game. It mainly includes temporal team roster information, a timeline of key events, and a historical timeline. The temporal team roster information records the home team currently active on the field (…). ) and the away team ( The player's identity, as well as detailed information such as the player's name, jersey number, and position, can be represented using the following formulas one and two:
[0050] (Formula 1)
[0051] (Formula 2)
[0052] in, Team Indicates the team a , h For the two teams playing on the field, Player This refers to the team members. As a coach, The position of the players (e.g., forward, center forward, defender, goalkeeper, etc.).
[0053] For example, the key event timeline records the timestamps and details of goals and cards played at the current moment of the match. This is because they are the most important match information for tracking match highlights. On the other hand, the historical timeline consists of the most recent... k The commentary section comprises articles reflecting recent events, which can be represented by the following formula three:
[0054] ,
[0055] (Formula 3)
[0056] in, As a key event, for t Historical events at that moment, For specific events, such as Playing cards (Including red and yellow cards), etc. Throughout the match, the constantly changing score and dynamic team lineups are key factors in determining the outcome. The score changes with each goal, denoted as:
[0057] ,
[0058]
[0059] in, P For the players, Score To score points, An own goal. Each time a player is substituted. The team's lineup will be updated, denoted as:
[0060] ,
[0061]
[0062]
[0063] Among them, player position attributes during substitution Position Changes may occur, and can be represented using P->P' or Q->Q'.
[0064] For example, when a player on a team is replaced, the team's roster is updated.
[0065] For example, information obtained solely from wide shots of the entire football field is often vague and limited. Therefore, football broadcasts often supplement the audience's understanding of football events through visual language, such as replays and close-up shots of players. This visual information serves as a key clue linking entities to football events. This application's embodiments design a fine-grained visual information extraction framework. First, video frame analysis is performed based on shot classification (wide shot, medium shot, close-up, off-field, etc.) to identify fine-grained visual information related to match events in medium and close-up shots. Second, a football player photo database is constructed, and facial recognition and jersey number recognition technologies are used to accurately identify players appearing in different video clips. Therefore, before performing visual reasoning, fine-grained visual information extraction is necessary, including player information, jersey information, and team alias information.
[0066] Specifically, step 202 above, the step of extracting fine-grained visual information, may further include steps 202b1 to 202b3:
[0067] Step 202b1: Use a visual language model to classify each shot in the target video segment to obtain shot classification information containing the category of each shot.
[0068] The lens categories include: close-up lenses, medium-range lenses, long-range lenses, and off-site lenses.
[0069] Understandably, camera shots in football videos are typically categorized into four types: close-up shots, etc. Medium shot Long shot and off-site cameras In football match broadcasts, players appearing in close-up shots are more likely to be associated with recent events. Therefore, they can provide potentially identifiable information to identify anonymous entities.
[0070] For example, in the embodiments of this application, fine-grained visual information in video clips can be captured through the following steps: shot boundary detection and classification, player face recognition, player jersey recognition, and team alias detection.
[0071] For example, regarding shot boundary detection and classification, embodiments of this application can segment video clips into a series of shots, represented as follows: Then, a visual language model (VLM) is used to classify each shot to obtain the category to which each shot belongs:
[0072]
[0073] Step 202b2: Based on the player information contained in the temporal team lineup information, filter out the facial images of each player from the database to obtain a set of facial images containing all players' facial images, and use image recognition technology to identify facial images that match the facial images in the set of facial images from the medium and close-up shots of the target video clip to determine the players contained in the target video clip.
[0074] The medium-close-up lens includes: a close-up lens, and / or a medium-range lens.
[0075] For example, regarding player face recognition, this application embodiment crawled facial images of 3213 relevant players from the Soccernet-v2 sports statistics website, creating an image database for close-up face registration. For each medium-close-up shot in the video clip... First, the current lineup is obtained based on the steps of reconstructing the game state context described above. and Then, facial images of players were selected from the database as 22 on-field candidates, denoted as... .
[0076] Afterwards, a professional facial recognition model can be used ( For video frame sequences The system detects faces appearing in the database and compares them with candidate faces. A comparison is performed. This process involves dynamically selecting three keyframes. This process is repeated to mitigate issues such as image blurring and faces going out of bounds caused by dynamic camera switching. The final recognition result is labeled as... .
[0077] For example, the facial recognition process can be formally described as follows:
[0078] ,
[0079] .
[0080] Step 202b3: Based on optical character recognition technology, identify the jerseys appearing in the close-up shots of the target video clip to obtain the jersey information of the players and the alias information of the participating teams contained in the target video clip.
[0081] For example, for player jersey recognition, optical character recognition (OCR) technology based on a visual language model can be used. However, because the details of jersey names and player numbers are small, blurring and overlapping are prone to occur. Therefore, in this embodiment, only a visual language model is used for close-up shots. The results of optical character recognition can be represented in the following form:
[0082]
[0083] in, Used to indicate the player's current action, such as running, dribbling, standing, etc.
[0084] Finally, fine-grained visual information is extracted from the target video segment. This can be expressed using the following formula four:
[0085] (Formula 4)
[0086] For example, team alias detection refers to associating a team with the color of the jerseys worn during matches. To achieve accurate identification, this application employs a triple verification mechanism: First, a Visual Language Model (VLM) is used in conjunction with traditional team color knowledge for cross-validation of video frames; second, a Large Language Model (LLM) is used to integrate similar color descriptions (e.g., classifying "red and blue stripes" and "blue and red stripes") and filter out clear cases. When color ambiguity arises (e.g., distinguishing between "light blue" and "dark blue" teams), the team is identified by the unique jersey numbers captured in the video, which are unique within the current lineup. Finally, the home team color is determined through manual verification. away team .
[0087] For example, regarding event location guidance information, a 30-second video clip contains far more events than described in the commentary because the input video is not precisely cut at the event level, but rather a fixed-length segment centered on a timestamp. Therefore, retrieving key football events mentioned in the commentary within this fixed-length segment is crucial. Thus, this embodiment of the application cleverly utilizes a pre-trained end-to-end football commentary model to acquire prior knowledge of the connection between video and commentary events. Since accurate end-to-end football commentary requires an inherent ability to locate key football events from video clips, this embodiment extracts the cross-attention layer of the end-to-end football commentary model and performs a series of vector transformations to obtain the model's attention coefficients for the input video in seconds, serving as clues for locating key football events.
[0088] Specifically, step 202 above, the step of obtaining event location guidance information using the end-to-end model, may further include the following steps 202c1 and 202c2:
[0089] Step 202c1: Obtain the query of each cross-attention layer teammate of the end-to-end model, and calculate the average cross-attention between the query and the video frame to obtain the attention weight of each query.
[0090] Step 202c2: After aggregating the attention weights of each query, a frame-level importance vector is obtained, and the event localization guidance information is obtained based on the frame-level importance vector.
[0091] The frame-level importance vector is used to characterize the importance weight of each video frame in the target video segment when generating anonymization and dissemination information.
[0092] For example, in order to obtain event localization guidance information, the average cross-attention between queries and video frames is first calculated. Let... Indicates the first Layer and first h Within the individual, search q With frames n Attention weights between all [parts / entities]. L Taking the average of each layer and all heads, we can obtain:
[0093]
[0094] Then, use the query output. The L2 criterion is used to calculate query importance, where Q is the total number of queries:
[0095]
[0096] The final frame-level importance vector is:
[0097]
[0098] For example, based on the above steps, a vector divided by frame can be obtained. W It represents the relevance of each video frame to the generated narration. It can be seen as the arousal level of each frame when generating the narration text, serving as a basic guide for the event.
[0099] Step 203: Perform visual reasoning based on the match status information, the fine-grained visual information, and the event location guidance information, and enhance the details of the anonymized, coarse-grained commentary information according to the visual reasoning results to obtain enhanced commentary information.
[0100] The visual reasoning results include: identifying and matching player and team references in the game, and matching the objects of each event.
[0101] For example, the final contextual clue structure of visual reasoning consists of the following information:
[0102]
[0103] Specifically, in step 203 above, the step of enhancing the details of the anonymized, coarse-grained explanatory information based on the visual reasoning result to obtain the enhanced-detail explanatory information may further include the following step 203a:
[0104] Step 203a: Perform visual reasoning based on the key event timeline, the historical timeline, the shot classification information, the facial image set, the players contained in the target video clip, the jersey information, and the alias information of the participating teams, and enhance the details of the anonymized, coarse-grained commentary information according to the visual reasoning results to obtain enhanced commentary information.
[0105] The visual reasoning result is obtained based on a multimodal visual model.
[0106] For example, this application embodiment first employs the advanced multimodal visual model Qwen-2.5VL to design inference steps, reasoning backwards on how to infer the correct player using the collected fine-grained clues (current on-field lineup, jersey team affiliation, player faces and jerseys appearing in close-up shots in the video, and key video event location guidance). Training data is constructed using this answer-guided visual reasoning method, and the inference ability of the multimodal visual model on this task is further optimized using two fine-tuning strategies: supervised fine-tuning and group relative strategy optimization. In this application embodiment, with 30 seconds of video and contextual cues, the visual understanding and reasoning capabilities of the visual language model can be used to infer the team members. Player and team name Team This allows for the alignment of entities.
[0107] Step 204: Perform knowledge enhancement on the detailed commentary information based on the internal and external knowledge bases to obtain knowledge-enhanced football commentary.
[0108] The external knowledge base includes historical statistical data of different football matches; the internal knowledge base includes static background information and dynamic update events of the current football match.
[0109] For example, the second phase of GS aims to enrich and refine the generated commentary using external historical data and dynamic match data.
[0110] To address the acquisition of external knowledge, this application's embodiments construct an external knowledge database integrating historical statistical data, such as match statistics, player performance, and team results from several past seasons. Through a knowledge-enhanced retrieval generation module (KAG), using natural language questions as queries, various data types frequently cited by commentators can be retrieved from the external database to obtain corresponding professional historical match data.
[0111] Regarding the acquisition of internal knowledge, this application embodiment utilizes real-time maintained fine-grained match data, such as match progress and score changes, and match information, such as the number of spectators, venue, and player personal data, as an internal knowledge base to obtain real-time and dynamic match status information, supporting commentary based on an understanding of the match background and context.
[0112] Then, by using an advanced large language model, the system spontaneously generates accurate questions needed to refine the current commentary event. Combined with historical match data from an external database (using a knowledge-enhanced retrieval and generation module) and real-time updated match status (using contextual learning), it generates more accurate and in-depth analytical commentary and more realistic and empathetic commentary.
[0113] Specifically, step 204 above may also include step 204a:
[0114] Step 204a: Based on the different commentators' methods of applying relevant knowledge during the commentary process, using natural language questions as queries, the system retrieves historical match data related to this match from the external knowledge base through a language model, and retrieves the real-time updated match status based on context learning from the internal knowledge base.
[0115] Based on the historical match data and the real-time updated match status, the detailed commentary information is supplemented and corrected to obtain the knowledge-enhanced football commentary.
[0116] For example, live television commentary exhibits a variety of styles due to factors such as the commentator's stance, speaking habits, personality, and cultural background. While comparing these different commentary styles within a single evaluation framework is challenging, human commentators often have a natural preference for incorporating contextual knowledge to enhance the accuracy and depth of their insights and to interpret and comment on visual scenes. This application embodiment simultaneously analyzes real-time text commentary and English-to-English automatic speech recognition (ASR) audio transcriptions surrounding each real-time text commentary timestamp. The large language model extracts knowledge through paired explicit query representations; for example, in this application embodiment, it prompts GPT-4o to identify pre-relevant knowledge implied in the commentary. Furthermore, this application embodiment employs a questioning mechanism to generate constraint questions based on the extracted knowledge.
[0117] The results showed that 15.02% of live TV commentary referenced various knowledge points, while only 3.16% of text-based live commentary contained knowledge, mostly in the form of score reports after goals or at the end of the match. To examine the distribution of knowledge references in human commentary, we used the same hints as MatchVision for transcriptions containing knowledge, summarizing the commentary into specific types of standard football events. The main events in which human commentators might reference knowledge included goals, corner kicks, and cards. Furthermore, based on the information source, knowledge could be divided into external knowledge and internal knowledge; the former relies on historical statistics from other matches, while the latter refers to static background information and dynamic updates from the current match. To ensure that the commentary generation model possesses sufficient knowledge as human commentators, we built a football KAG system for acquiring external statistics and a background match knowledge base for tracking internal information.
[0118] External Football KAG System Modification: The SoccerNet dataset was used to reconstruct well-organized data, and SoccerRAG was used to interpret the structured information into a database format. Building upon these excellent pioneering works, we now possess a comprehensive database covering six seasons of football matches across three major leagues. To better align with the retrieval habits of human commentators, this application's embodiments improve the football KAG system by incorporating more fine-grained details, increasing the number of query examples, improving SQL query construction, and validating the execution results. Specifically, the main differences between this application's technical solution and SoccerRAG include: 1. Adding more query examples to the relevant knowledge used by high-frequency commentators to improve the accuracy of natural language problem description -> SQL query construction. 2. Improving the data structure to increase the dataset's support for fine-grained event-level queries, such as individual goals and cards for players and teams. 3. Adding strict time constraints to the query pattern to protect the database from the influence of future matches when responding to the current specified match, and carefully checking the answer time range using LLM.
[0119] Internal match context construction: Unlike the player detection task in the scenario description, where we provided jersey numbers for all players in the lineup, here we provide more detailed information about the mentioned players, such as their nationality. To improve the accuracy of ICL and reduce confusion, we individually labeled the key timelines for each event type. Furthermore, we captured fine-grained information such as the goal scorer, the assisting player, and the specific type of goal (e.g., penalty, header, or own goal). Other relevant details, including tactics, audience numbers, and the match venue, were also incorporated to provide comprehensive background information.
[0120] Commentary Integration with a Large Language Model: Starting with the first phase of commentary describing the current event, we use GPT-4o to generate questions relevant to external statistics. The responses from the KAG system are then double-checked to eliminate invalid answers. Furthermore, the model is explicitly prompted to reference internal game contextual knowledge, such as key timelines of the mentioned events and detailed player information. This process improves the accuracy and depth of the commentary, making it more consistent with live television commentary.
[0121] The football commentary generation method based on visual reasoning and knowledge enhancement provided in this application first obtains anonymized, coarse-grained commentary information output by an end-to-end model based on a target video segment. Then, it extracts match status information and fine-grained visual information based on the target video segment and uses the end-to-end model to obtain event location guidance information. The fine-grained visual information is used to represent detailed information in the match. Visual reasoning is then performed based on the match status information, the fine-grained visual information, and the event location guidance information. Based on the visual reasoning results, the anonymized, coarse-grained commentary information is enhanced to obtain detailed commentary information. The visual reasoning results include: identifying and corresponding player and team designations in the match, and matching the objects of various events. Finally, knowledge enhancement is performed on the detailed commentary information based on an internal knowledge base and an external knowledge base to obtain knowledge-enhanced football commentary. The external knowledge base includes historical statistical data from different football matches; the internal knowledge base includes static background information and dynamically updated events for the current football match. Thus, by using a two-stage processing framework of entity alignment based on visual reasoning and narration generation based on knowledge enhancement, we can better simulate the workflow of human narrators and generate more natural and information-rich narrations.
[0122] It should be noted that the football commentary generation method based on visual reasoning and knowledge enhancement provided in this application embodiment can be executed by a football commentary generation device based on visual reasoning and knowledge enhancement, or by a control module within that device for executing the football commentary generation method based on visual reasoning and knowledge enhancement. This application embodiment uses the execution of the football commentary generation method based on visual reasoning and knowledge enhancement by a football commentary generation device as an example to illustrate the football commentary generation device based on visual reasoning and knowledge enhancement provided in this application embodiment.
[0123] It should be noted that, in the embodiments of this application, the accompanying drawings of the various methods described above illustrate the football commentary generation methods based on visual reasoning and knowledge enhancement, all of which are exemplified by referring to one of the accompanying drawings in the embodiments of this application. In specific implementation, the football commentary generation methods based on visual reasoning and knowledge enhancement shown in the accompanying drawings of the various methods described above can also be implemented in conjunction with any other accompanying drawings that can be combined as illustrated in the above embodiments, which will not be elaborated here.
[0124] The football commentary generation device based on visual reasoning and knowledge enhancement provided in this application is described below. The football commentary generation method based on visual reasoning and knowledge enhancement described below can be referred to in correspondence with the football commentary generation method described above.
[0125] Figure 3 A schematic diagram of the structure of the football commentary generation device based on visual reasoning and knowledge enhancement provided in the embodiments of this application is shown below. Figure 3 As shown, it specifically includes:
[0126] The acquisition module 301 is used to acquire anonymized, coarse-grained commentary information output by the end-to-end model based on the target video segment; the information extraction module 302 is used to extract match status information and fine-grained visual information based on the target video segment, and to acquire event location guidance information using the end-to-end model; the fine-grained visual information is used to represent detailed information in the match; the detail enhancement module 303 is used to perform visual reasoning based on the match status information, the fine-grained visual information, and the event location guidance information, and to enhance the anonymized, coarse-grained commentary information based on the visual reasoning results to obtain detailed-enhanced commentary information; the visual reasoning results include: identifying and corresponding player and team designations in the match, and matching the objects of each event; the knowledge enhancement module 304 is used to enhance the detailed-enhanced commentary information based on an internal knowledge base and an external knowledge base to obtain knowledge-enhanced football commentary; wherein, the external knowledge base includes: historical statistical data of different football matches; the internal knowledge base includes: static background information and dynamically updated events of the current football match.
[0127] Optionally, the match status information includes: temporal team lineup information, key event timeline, and historical timeline; the information extraction module 302 is specifically used to generate the temporal team lineup information, the key event timeline, and the historical timeline based on the target video clip and historical video clips; the information extraction module 302 is further used to update the temporal team lineup information based on the card-playing event and the player replacement event represented by the historical timeline; wherein, the temporal team lineup information is used to represent the player information of each team on the field at present; the key event timeline is used to represent the timestamps of goal events and card-playing events; the historical timeline consists of multiple recent commentaries and is used to represent the most recently occurred events.
[0128] Optionally, the fine-grained visual information includes: player information, jersey information, and participating team alias information; the information extraction module 302 is specifically used to classify each shot in the target video segment using a visual language model to obtain shot classification information containing each shot category; the shot categories include: close-up shots, medium shots, long shots, and off-field shots; the information extraction module 302 is further used to filter out the facial images of each player from the database based on the player information contained in the temporal team lineup information to obtain a set of facial images containing all player facial images, and to use image recognition technology to identify facial images that match the facial images in the set of facial images from the medium and close-up shots of the target video segment to determine the players contained in the target video segment; the information extraction module 302 is further used to identify the jerseys appearing in the close-up shots of the target video segment based on optical character recognition technology to obtain the jersey information and participating team alias information of the players contained in the target video segment; wherein, the medium and close-up shots include: close-up shots and / or, medium shots.
[0129] Optionally, the information extraction module 302 is specifically used to obtain the query of each cross-attention layer teammate of the end-to-end model, and calculate the average cross-attention between the query and the video frame to obtain the attention weight of each query; the information extraction module 302 is also specifically used to aggregate the attention weight of each query to obtain a frame-level importance vector, and obtain the event localization guidance information based on the frame-level importance vector; wherein, the frame-level importance vector is used to characterize the importance weight of each video frame in the target video segment when generating anonymization and explanation information.
[0130] Optionally, the detail enhancement module 303 is specifically used to perform visual reasoning based on the key event timeline, the historical timeline, the shot classification information, the facial image set, the players contained in the target video clip, the jersey information, and the alias information of the participating teams, and to enhance the details of the anonymized, coarse-grained explanatory information according to the visual reasoning results, so as to obtain detailed-enhanced explanatory information; wherein, the visual reasoning results are obtained based on a multimodal visual model.
[0131] Optionally, the knowledge enhancement module 304 is specifically used to, based on the methods by which different commentators apply relevant knowledge during the commentary process, use natural language questions as queries, obtain historical match data related to the current match from the external knowledge base through a language model, and obtain the real-time updated match status based on context learning from the internal knowledge base; supplement and correct the detailed enhanced commentary information based on the historical match data and the real-time updated match status to obtain the knowledge-enhanced football commentary.
[0132] The football commentary generation device based on visual reasoning and knowledge enhancement provided in this application first obtains anonymized, coarse-grained commentary information output by an end-to-end model based on a target video segment; then, it extracts match status information and fine-grained visual information based on the target video segment, and uses the end-to-end model to obtain event location guidance information; the fine-grained visual information is used to represent detailed information in the match; and visual reasoning is performed based on the match status information, the fine-grained visual information, and the event location guidance information, and the anonymized, coarse-grained commentary information is enhanced with details according to the visual reasoning results to obtain detailed-enhanced commentary information; the visual reasoning results include: identifying and corresponding player and team references in the match, and matching the objects of each event; finally, knowledge enhancement is performed on the detailed-enhanced commentary information based on an internal knowledge base and an external knowledge base to obtain knowledge-enhanced football commentary; wherein, the external knowledge base includes: historical statistical data of different football matches; the internal knowledge base includes: static background information and dynamically updated events of the current football match. Thus, by using a two-stage processing framework of entity alignment based on visual reasoning and narration generation based on knowledge enhancement, we can better simulate the workflow of human narrators and generate more natural and information-rich narrations.
[0133] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communications bus 440, wherein the processor 410, the communications interface 420, and the memory 430 communicate with each other through the communications bus 440. The processor 410 can call logical instructions in the memory 430 to execute a football commentary generation method based on visual reasoning and knowledge enhancement. This method includes: first, acquiring anonymized, coarse-grained commentary information output by an end-to-end model based on a target video segment; then, extracting match status information and fine-grained visual information based on the target video segment, and using the end-to-end model to acquire event location guidance information; the fine-grained visual information is used to represent detailed information in the match; and performing visual reasoning based on the match status information, the fine-grained visual information, and the event location guidance information, and enhancing the anonymized, coarse-grained commentary information based on the visual reasoning results to obtain enhanced commentary information; the visual reasoning results include: identifying and corresponding player and team designations in the match, and matching the objects of various events; finally, performing knowledge enhancement on the enhanced commentary information based on an internal knowledge base and an external knowledge base to obtain knowledge-enhanced football commentary; wherein the external knowledge base includes: historical statistical data of different football matches; the internal knowledge base includes: static background information and dynamically updated events of the current football match. Thus, by using a two-stage processing framework of entity alignment based on visual reasoning and narration generation based on knowledge enhancement, we can better simulate the workflow of human narrators and generate more natural and information-rich narrations.
[0134] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0135] On the other hand, this application also provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer can execute the football commentary generation method based on visual reasoning and knowledge enhancement provided by the above methods. This method includes: first, acquiring anonymized, coarse-grained commentary information output by an end-to-end model based on a target video segment; then, extracting match status information and fine-grained visual information based on the target video segment, and using the end-to-end model to acquire event location guidance information; the fine-grained visual information is used to characterize detailed information in the match; and Visual reasoning is performed based on the match status information, fine-grained visual information, and event location guidance information. The anonymized, coarse-grained commentary information is then enhanced with detail based on the visual reasoning results, resulting in enhanced commentary. The visual reasoning results include identifying and matching player and team designations in the match, and matching the objects of various events. Finally, knowledge enhancement is performed on the enhanced commentary information based on an internal and external knowledge base, resulting in knowledge-enhanced football commentary. The external knowledge base includes historical statistical data from different football matches, and the internal knowledge base includes static background information and dynamically updated events for the current football match. Thus, through a two-stage processing framework of entity alignment based on visual reasoning and commentary generation based on knowledge enhancement, the workflow of a human commentator can be better simulated, generating more natural and information-rich commentary.
[0136] Furthermore, this application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the aforementioned methods for generating football commentary based on visual reasoning and knowledge enhancement. This method includes: first, acquiring anonymized, coarse-grained commentary information output by an end-to-end model based on a target video segment; then, extracting match status information and fine-grained visual information based on the target video segment, and using the end-to-end model to acquire event location guidance information; the fine-grained visual information is used to characterize detailed information during the match; and based on the match status information and the fine-grained visual information... The system performs visual reasoning on the information and event location guidance information, and enhances the details of the anonymized, coarse-grained commentary information based on the visual reasoning results to obtain enhanced commentary information. The visual reasoning results include: identifying and corresponding player and team references in the match, and matching the objects of each event. Finally, the enhanced commentary information is augmented with knowledge based on an internal and external knowledge base to obtain knowledge-enhanced football commentary. The external knowledge base includes historical statistical data of different football matches; the internal knowledge base includes static background information and dynamically updated events of the current football match. Thus, through a two-stage processing framework of entity alignment based on visual reasoning and commentary generation based on knowledge enhancement, the workflow of a human commentator can be better simulated, generating more natural and information-rich commentary.
[0137] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0138] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0139] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for generating football commentary based on visual reasoning and knowledge enhancement, characterized in that, include: Obtain anonymized, coarse-grained explanatory information from the end-to-end model output based on the target video segment; Based on the target video segment, extract the competition status information and fine-grained visual information, and use the end-to-end model to obtain event localization guidance information; The fine-grained visual information is used to characterize the details in the competition; Visual reasoning is performed based on the match status information, the fine-grained visual information, and the event location guidance information. Based on the visual reasoning results, the anonymized, coarse-grained commentary information is enhanced in detail to obtain the enhanced commentary information. The visual reasoning results include: identifying and matching player and team references in the game, and matching the objects of each event. The detailed commentary information is enhanced based on internal and external knowledge bases to obtain knowledge-enhanced football commentary. The external knowledge base includes historical statistical data of different football matches; the internal knowledge base includes static background information and dynamic update events of the current football match. The step of obtaining event localization guidance information using the end-to-end model includes: Obtain the query corresponding to each cross-attention layer of the end-to-end model, and calculate the average cross-attention between the query and the video frame to obtain the attention weight of each query; After aggregating the attention weights of each query, a frame-level importance vector is obtained, and the event localization guidance information is obtained based on the frame-level importance vector. The frame-level importance vector is used to characterize the importance weight of each video frame in the target video segment when generating anonymization and disclaimer information; The process of enhancing the detailed commentary information based on internal and external knowledge bases to obtain knowledge-enhanced football commentary includes: Based on the different commentators' methods of applying relevant knowledge during the commentary process, using natural language questions as queries, the system retrieves historical match data related to this match from the external knowledge base through a language model, and retrieves the real-time updated match status based on context learning from the internal knowledge base. Based on the historical match data and the real-time updated match status, the detailed commentary information is supplemented and corrected to obtain the knowledge-enhanced football commentary.
2. The football commentary generation method based on visual reasoning and knowledge enhancement according to claim 1, characterized in that, The match status information includes: temporal team lineup information, key event timeline, and historical timeline; the temporal team lineup information is used to represent the player information of each team currently on the field; the key event timeline is used to represent the timestamps of goal events and card-playing events; the historical timeline consists of multiple recent commentaries and is used to represent the most recent events. The step of extracting match status information based on the target video segment includes: The temporal team lineup information, the key event timeline, and the historical timeline are generated based on the target video clip and historical video clips. The temporal team roster information is updated based on the card-playing events and the player replacement events represented by the historical timeline.
3. The football commentary generation method based on visual reasoning and knowledge enhancement according to claim 2, characterized in that, The fine-grained visual information includes: player information, jersey information, and alias information of participating teams; The extraction of fine-grained visual information based on the target video segment includes: Each shot in the target video segment is classified using a visual language model to obtain shot classification information containing each shot category; the shot categories include: close-up shot, medium shot, long shot, and off-camera shot; Based on the player information contained in the temporal team lineup information, facial images of each player are filtered out from the database to obtain a set of facial images containing all players' facial images. Then, using image recognition technology, facial images that match the facial images in the set of facial images are identified from the medium and close-up shots of the target video clip to determine the players contained in the target video clip. Based on optical character recognition technology, the jerseys appearing in close-up shots of the target video clip are identified to obtain the jersey information of the players and the alias information of the participating teams contained in the target video clip; The medium-close-up lens includes: a close-up lens, and / or a medium-range lens.
4. The football commentary generation method based on visual reasoning and knowledge enhancement according to claim 3, characterized in that, The process involves performing visual reasoning based on the match status information, the fine-grained visual information, and the event location guidance information, and then enhancing the details of the anonymized, coarse-grained commentary information based on the visual reasoning results to obtain enhanced commentary information, including: Visual reasoning is performed based on the key event timeline, the historical timeline, the shot classification information, the facial image set, the players contained in the target video clip, the jersey information, and the alias information of the participating teams. Based on the visual reasoning results, the anonymized, coarse-grained explanatory information is enhanced in detail to obtain enhanced explanatory information. The visual reasoning result is obtained based on a multimodal visual model.
5. A football commentary generation device based on visual reasoning and knowledge enhancement, characterized in that, The device includes: The acquisition module is used to acquire anonymized, coarse-grained explanatory information output by the end-to-end model based on the target video segment; The information extraction module is used to extract competition status information and fine-grained visual information based on the target video segment, and to obtain event localization guidance information using the end-to-end model; the fine-grained visual information is used to characterize the details in the competition. The detail enhancement module is used to perform visual reasoning based on the match status information, the fine-grained visual information, and the event location guidance information, and to enhance the details of the anonymized, coarse-grained commentary information according to the visual reasoning results, so as to obtain the detailed-enhanced commentary information; the visual reasoning results include: identifying and corresponding player references and team references in the match, and matching the objects of each event. The knowledge enhancement module is used to enhance the detailed commentary information based on the internal and external knowledge bases to obtain knowledge-enhanced football commentary. The external knowledge base includes historical statistical data of different football matches; the internal knowledge base includes static background information and dynamic update events of the current football match. The information extraction module is specifically used to obtain the query corresponding to each cross-attention layer of the end-to-end model, and calculate the average cross-attention between the query and the video frame to obtain the attention weight of each query; the information extraction module is also specifically used to aggregate the attention weight of each query to obtain a frame-level importance vector, and obtain the event localization guidance information based on the frame-level importance vector; wherein, the frame-level importance vector is used to characterize the importance weight of each video frame in the target video segment when generating anonymization explanation information; The knowledge enhancement module is specifically used to retrieve historical match data related to the current match from the external knowledge base based on the methods by which different commentators apply relevant knowledge during the commentary process, using natural language questions as queries, and retrieving the match status updated in real time based on context learning from the internal knowledge base. The knowledge enhancement module is further used to supplement and correct the detailed commentary information based on the historical match data and the real-time updated match status, so as to obtain the knowledge-enhanced football commentary.
6. The football commentary generation device based on visual reasoning and knowledge enhancement according to claim 5, characterized in that, The match status information includes: temporal team lineup information, key event timeline, and historical timeline; the temporal team lineup information is used to represent the player information of each team currently on the field; the key event timeline is used to represent the timestamps of goal events and card-playing events; the historical timeline consists of multiple recent commentaries and is used to represent the most recent events. The information extraction module is specifically used to generate the temporal team lineup information, the key event timeline, and the historical timeline based on the target video clip and historical video clips. The information extraction module is further used to update the temporal team lineup information based on the card-playing event and the player replacement event represented by the historical timeline.
7. An electronic device, characterized in that, The system includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the steps of the football commentary generation method based on visual reasoning and knowledge enhancement as described in any one of claims 1 to 4.
8. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the football commentary generation method based on visual reasoning and knowledge enhancement as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Explanation method and device for sports competition
CN110826361A
Automatic commentary device and method for football match
CN113268515A